Why your LLM app feels slow when the model is fast
The tagging step in my video pipeline took eleven seconds per script. I spent an evening with benchmark tables open, comparing tokens per second across four models, working out which one to switch to.
Then I printed the prompt it was actually sending. A style guide, six examples I had pasted in during a debugging session in March and never removed, and the full script included twice because two different functions each appended it.
The model was never the slow part. It rarely is.
Latency hides in four places and the model is the fourth. Here they are in the order I check them.
Two numbers, not one
Most apps log a single latency figure per request. That figure hides everything.
Split it. How long until the first word appears, and how quickly the words arrive after that.
Those two have almost nothing to do with each other. The first is dominated by how much text you sent and whatever setup runs before generation begins. The second is dominated by the hardware and by how many other requests are sharing it.
The split matters because people tolerate them differently. A twelve second answer that started streaming at 400ms feels fine. The same twelve seconds behind a spinner feels broken, and users will tell you the app is slow while your dashboard shows an unchanged average.
With one latency number you cannot diagnose anything. Add the second before you touch anything else on this list.
Streaming first, because it is free
If you are holding the response until it is complete, that is the largest perceived speed problem you have, and no model change will touch it.
Stream the output. Total duration stays identical. The wait turns into visible progress.
I have watched people skip this because it complicates the client, then spend three weeks shaving 800ms off a retrieval step that streaming would have covered for nothing.
Where you genuinely cannot stream, fake the progress honestly. My TTS step hands back a finished mp3 with nothing to stream mid flight, so the pipeline prints paragraph counts as it works through the script. Cruder signal, same effect, still better than a frozen bar.
Print the prompt you are actually sending
Not the template. The assembled string, every character of it, dumped to a terminal where you can read it.
All of that has to be processed before a single output token exists. A system prompt that grew for eight months, plus ten retrieved chunks, plus the whole conversation history, is a large block of text being read from scratch on every request.
Mine was around 4,100 tokens. Maybe 900 of them were doing work.
The usual culprits, roughly in the order I find them:
- examples added while debugging and never deleted
- the same instruction written three ways because nobody trusted the first version
- context from a retriever configured to always return ten results whether ten are relevant or not
- template fields inserted unconditionally, half of them empty
On retrieval specifically, cutting from ten chunks to four made my extraction step faster and the outputs better, because the useful passage stopped competing for attention with six irrelevant ones. Fewer, better chunks wins in both directions.
Cache the front, vary the back
If a large slice of your prompt is identical on every call, and it usually is, you are paying to process the same text over and over.
Every major provider lets you mark the stable portion so it gets reused between calls. The effect on time to first token is large, and it cuts the bill at the same time.
The catch is ordering. The reusable prefix has to come first and stay byte identical. System instructions and fixed reference material at the top. User message, retrieved context, anything that varies at all, at the bottom.
I had a brand style guide sitting after the article body, purely because that was the order I wrote the function in months earlier. Moving it above the body took five minutes.
Then I broke it myself. The system prompt included the current date for freshness, and I had formatted that date with the time in it, so every prefix was unique and every single call missed. Cache hit rate sat at zero for a fortnight before I looked. If yours reads zero, something in your prefix is changing and you almost certainly put it there deliberately.
Your own call graph
This is the part that is entirely your fault, and where my worst offender lived.
Plenty of apps make several model calls per user action. A classification step, then retrieval, then generation, then a check on the output. Run those one after another and your latency is the sum of all of them.
Most of them do not depend on each other. Run them concurrently and your latency becomes the slowest one instead of the total.
My tagger called the model once per paragraph. A script has fifty odd paragraphs, so that was fifty sequential calls at roughly 300ms each, none of which needed the answer from any of the others.
Then ask whether each call needs to exist at all. Mine picked one of about twenty b-roll categories from a paragraph of narration. A keyword table gets that right most of the time, and I eyeball the tags before rendering anyway, so a wrong one costs me a glance.
People argue with me about this, so I will state it plainly. A lookup table that is right 85% of the time with a human reading the output beats a model call that is right 97% of the time and costs 300ms plus a token line. Accuracy you never inspect is worth less than speed you can feel.
The fastest call is the one you deleted. Second fastest is the one running alongside another.
The queue you cannot see
Part of your wait is your request sitting behind other people’s requests on a shared endpoint. You cannot see it, cannot control it, and it drifts with the time of day.
So when your latency has a wide spread, where the median looks fine and the slow requests are dramatically worse, look here before you go hunting through your own code. Averages bury exactly this behaviour. Everybody having a bad time is in the tail.
Three things you can do about it. Set a timeout and retry, on the theory that a fresh request lands somewhere quieter. Run the model yourself. Or design the interface so a slow response is survivable, which walks you straight back to streaming.
I keep my own hardware for other parts of the business, a rack of Android phones and a pile of modems, so I have a reasonable sense of what self hosting costs once you own it. A GPU idle for most of the day to fix your 95th percentile is a bigger bill than the 95th percentile was. Unless that card is already busy with something else, the trade is bad.
Some of your latency is simply not yours to fix, and knowing which part stops you optimising code that was never the problem.
Output length is a latency decision
Generation is sequential. Each token waits for the one in front of it. Twice the output, twice the wait, and no infrastructure change alters that arithmetic.
Which makes response length something you chose in your prompt, probably without thinking of it as a performance setting. Ask for a detailed and thorough explanation and you have asked for a slow response.
Be specific about length. Three sentences, when three sentences will do.
In anything conversational this compounds badly. Every long answer becomes history that gets resent on the next turn, so a verbose model makes every subsequent request slower and more expensive by inflating what you have to send again. Trimming output pays twice.
Reasoning models break most of the above
A model that thinks before answering spends tokens working through the problem privately, and you wait for all of it before the visible answer starts. The gap before the first word can be very long, and it swings with how hard the model decides the problem is.
That kills streaming as a comfort blanket, because there is nothing to stream during the thinking. It also makes latency wildly unpredictable. Two requests that look identical to you can differ by a factor of five.
Give them their own tier and their own interface treatment. Tell the user something is being worked through. Do not put one behind a button that is meant to feel instant.
And check whether the request in front of you needs one. Mostly it does not, and routing only the genuinely hard cases keeps the common path quick.
Testing with hello tells you nothing
People measure latency by firing off a short request and timing it. That number is meaningless, because it exercises none of what makes the application slow. No long prompt. No retrieval. No history. No realistic output length.
Capture a hundred genuine requests from your own traffic, replay them, and read the median and the 95th percentile side by side. That is the only measurement that predicts what your users experience, and it will usually disagree with your intuition about which part is slow.
The order I would check it in
Stream, if you are not already. Largest perceived gain, independent of everything else on the list.
Print the assembled prompt and cut it.
Reorder so the stable part sits at the front, turn caching on, then go and read the hit rate rather than assuming it worked.
Look for sequential calls that could run together, and for steps that do not need a model.
Only then evaluate a different model. By that point the question has usually answered itself, because the model was contributing a fraction of what the four things above were.
Current pricing and the model comparison tables I keep updated are here.