freecoding.school100% FREE · NO SIGNUP
Tensor TownISSUE #117 of 120

latency · ttft · tps · streaming

NeuraVSThe Overfit Ogre
Neura saysLatency matters: time-to-first-token, tokens-per-second, and streaming shape UX.

Two metrics define LLM responsiveness: TTFT (time to first token — how long until something appears) and TPS (tokens per second — how fast it streams after). Streaming the response token-by-token makes long answers feel fast because the user reads as it generates. For UX, prioritize low TTFT and stream; for throughput/cost, batch. Model size, context length, and load all affect both.

Power-ups you unlock

The Overfit Ogre attacks — common mistakes

Boss battleExplain why streaming improves perceived speed even at the same TPS.

Example code

<!doctype html><html><head><meta charset="utf-8"></head>
<body style="background:#06040d;color:#e6e0ff;font-family:monospace;padding:20px"><pre>TTFT: how soon the first token appears
TPS: how fast it streams after
stream → user reads as it generates → feels fast</pre></body></html>
▶ Open the interactive comic issue
‹ Cost · Per-Token EconomicsCarbon · The Energy Footprint ›