← all articles

Time to first token vs total latency: what you should actually measure

The number everyone quotes and the number that actually matters

Every model provider’s status page shows time to first token (TTFT). It’s the easiest latency number to make look good, and it’s also the one most people misread. TTFT tells you how long a user stares at a blank screen before anything happens. It does not tell you how long they wait for the answer they actually asked for. If you’re building a chat interface, those are two very different products.

I’ve shipped both streaming and non-streaming endpoints in production apps, and the gap between “feels fast” and “is fast” comes down to conflating these two numbers. Let’s separate them properly.

What time to first token actually measures

TTFT is the interval between when you send a request and when the first token of the response arrives. For an API call, that interval includes:

  • Network round trip to the provider’s edge
  • Queue time if the provider is under load
  • Prompt processing (the model has to run your entire input through its layers before it can emit anything, so a long prompt costs you here even though no output has appeared yet)
  • Time to generate that first output token

That last point trips people up. TTFT is not free of compute cost. A 20,000 token context window with a short question at the end still has to be processed in full before token one comes out. If you’ve ever noticed that a call with a huge system prompt “hangs” longer before streaming starts, that’s prompt processing, not network lag.

TTFT is the right metric when your product is a conversational surface: chatbots, coding assistants inline in an editor, anything where the user is watching text appear and forming their expectations from the first visible sign of life. A low TTFT with a slow tail still feels responsive, because the user has something to read while the rest generates.

What total latency actually measures

Total latency is the time from request to the last token, full stop. It’s TTFT plus generation time for the entire output, plus any time your own code spends before it can act on the response.

This is the number that matters when the output isn’t meant to be read as it streams. If you’re calling an LLM to produce a JSON object that your code parses, or a batch summarization job that writes to a database, nobody is watching the tokens arrive. Streaming a partial JSON blob doesn’t help you if you can’t do anything with it until it’s syntactically complete. In that case TTFT is close to irrelevant and total latency is the only number your users (or your downstream service) will ever feel.

I’ve seen teams optimize TTFT on a structured-output extraction pipeline and wonder why the end-to-end job didn’t get faster. The pipeline was never bottlenecked on time-to-first-token: it was bottlenecked on total tokens generated, because nothing downstream could consume a partial object.

Why generation speed is the variable that connects them

Total latency roughly decomposes as:

total latency ≈ TTFT + (output tokens / tokens-per-second)

That’s a simplification, since generation speed isn’t perfectly constant across a response, but it’s close enough to reason about. It means two calls with identical TTFT can have wildly different total latency if one produces 50 output tokens and the other produces 2,000. A model that streams its first token quickly but generates slowly per-token will still leave you waiting a long time for a long answer.

This is also why “this model feels fast” impressions from casual use are unreliable for judging batch or structured-output workloads. A snappy first token on a short chat reply says nothing about how that same model performs generating a 3,000 word document or a large JSON payload. Different output lengths exercise different parts of the latency budget.

Where reasoning and thinking tokens complicate this

Models that do extended reasoning before answering (chain-of-thought or explicit “thinking” tokens that get generated but not always shown to the user) push a chunk of the work into a phase that may or may not count toward what you perceive as TTFT, depending on whether the provider streams those intermediate tokens or holds them back until the visible answer starts. If a provider buffers reasoning tokens and only streams the final answer, your measured TTFT will look worse than the model’s raw first-token speed, because you’re waiting through an invisible generation phase before anything appears. If you’re benchmarking TTFT across providers or model tiers, check whether reasoning tokens are streamed, suppressed, or billed separately, because that changes what the number represents even when the label on the dashboard says the same thing.

Measuring it yourself

Don’t trust a vendor dashboard’s aggregate number for your own workload, because it’s averaged across traffic patterns that aren’t yours. If you’re using a streaming API, the practical way to measure TTFT is to record a timestamp right before you send the request and a second timestamp the moment your client receives the first chunk from the stream. For total latency, record the timestamp when the stream closes or the non-streaming response returns.

A few things that will quietly wreck your numbers if you don’t account for them:

  • Client-side buffering. Some HTTP client libraries buffer response chunks before handing them to your code, which inflates your measured TTFT beyond what the provider actually sent.
  • Cold starts on self-hosted or serverless inference. The first call after an idle period pays a model-loading tax that has nothing to do with steady-state performance.
  • Measuring from your own server versus measuring from the end user’s browser. A latency number captured server-side in the same region as the provider’s API will always look better than what your actual user experiences, because you’ve excluded the last-mile network hop.
  • Prompt length changing between test runs. If you’re comparing TTFT across two configurations, keep the input length roughly constant, or you’re really just measuring prompt processing time and calling it something else.

Run your own numbers on your own prompts and your own network path. A benchmark you didn’t run, on hardware you don’t control, with a prompt shape that doesn’t match yours, is not a number you should build a decision on.

Picking the number to optimize

If your product streams text into a UI the user is actively watching, optimize TTFT first and generation speed second, because the perceived experience is dominated by how quickly something appears on screen. If your product produces a complete artifact before anything downstream can use it (structured output, function calls, batch jobs, anything parsed as a whole), optimize total latency and treat TTFT as a secondary curiosity. And if you’re paying per-token for a reasoning model, remember that a fast TTFT on a call that burns thousands of hidden reasoning tokens before the visible answer starts can still leave you with both a slow total latency and a surprising bill.

Neither number alone tells you whether your app is fast. Report both, on your own traffic, and you’ll know which one to spend your engineering time on.

For more explainers on evaluating AI tools without the vendor spin, head to the AI Tool Gazette home page.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →