← all articles

The gap between an AI demo and something you can run

The five-minute demo lied to you

You’ve seen the video. Someone pastes a prompt into a chat window, the model spits out working code or a coherent answer, and the caption says “built in an afternoon.” That part is true. What the caption doesn’t say is that the afternoon build is a single happy-path run, on a fast connection, with no other users, no rate limits, and no one feeding it a weird PDF or a question in a language the system prompt didn’t anticipate.

Shipping the same thing to actual users means it has to survive concurrency, bad input, partial failures, and a bill that scales with usage. That’s a different engineering problem than getting one good output on one try. This piece is about where that gap actually shows up, based on what breaks when you take a demo and try to keep it running.

Latency you can hide once, not forever

In a demo you can narrate over a 4-second wait. “And now it’s thinking…” In production, a 4-second response on every request is a UX problem, and worse, it’s a queueing problem. If your app calls an LLM synchronously inside a web request, every one of those requests holds a thread or a connection open for the full generation time. Add a second LLM call to check or reformat the first one’s output, and you’ve doubled that hold time.

The usual fixes are streaming the response token by token so the user sees progress, or moving the call off the request path entirely into a background job with a status check. Both add real code: a queue, a way to poll or push status, error states for the in-between. None of that exists in a demo because a demo has one user and no queue.

Concurrency and rate limits are where the API contract gets tested

A provider’s API has a rate limit measured in requests per minute and tokens per minute, and those limits are usually tied to your account tier, not your app’s actual usage pattern. A demo makes one call. A product with fifty concurrent users making requests can burst past a limit in seconds, and the failure mode is a 429 response, not a slow response.

Handling that means retry logic with backoff, and deciding what happens to the user’s request while you retry: do they wait, do they see a spinner, do you silently queue it. It also means understanding that a naive “retry on any error” loop can make a rate-limit problem worse by resending the same burst of requests. None of this is visible in a single-shot demo because there’s nothing else competing for the same limit.

The context window is a budget, not a courtesy

A demo prompt is short. A production prompt often isn’t, because you’re stuffing in retrieved documents, conversation history, system instructions, and maybe a few examples. Every one of those consumes tokens against a fixed window, and tokens cost money on both the input and output side.

This is where RAG pipelines specifically get expensive to reason about. If your retrieval step returns five chunks of 500 tokens each to “be safe,” you’re paying for 2,500 tokens of context on every single query whether or not the chunks were relevant. Multiply that by call volume and it’s a real line item, not a rounding error. The demo version of a RAG app usually retrieves against a small, clean test corpus and never has to answer the question “what do we do when retrieval returns garbage.”

Structured output breaks in ways a demo never surfaces

Asking a model to return JSON works fine when you try it three times in a notebook. It stops working fine when you run it ten thousand times and a small percentage come back with a trailing comma, an extra explanation before the JSON, or a field that’s supposed to be a number but comes back as the string “N/A.” A demo doesn’t have a ten-thousand-run tail. Production does.

The practical response is validation on every response, not just a try/except around a JSON parse. That means schema validation, a defined fallback when parsing fails, and a decision about whether to retry the call, ask the model to fix its own output, or hand the user a degraded response. If you’re using a provider’s structured-output or function-calling mode, that reduces the failure rate but doesn’t take it to zero, and you still need code for the cases it doesn’t catch.

Cost has a shape, and the shape matters more than the average

A demo costs a few cents. A production system has a cost curve, and the curve is usually driven by a small number of expensive requests, not the average request. A user who pastes in a 50-page document, or a conversation that’s been going for forty turns and is now dragging its entire history into every new call, costs far more than the median interaction. If you price your product against the average cost per request and never look at the tail, you can lose money on your heaviest users without noticing until the monthly invoice arrives.

Managing this means tracking cost per request as a real metric, not an afterthought, and making explicit decisions: do you cap conversation history length, do you summarize old turns instead of resending them, do you charge differently for heavier usage. A demo never has a forty-turn conversation, so this problem simply doesn’t exist until you have real users who keep talking to the thing.

Evaluation without a benchmark you actually ran

It’s tempting to point at a public leaderboard number and call the model choice settled. That number was measured on someone else’s tasks with someone else’s prompts. It tells you very little about whether the model handles your specific documents, your specific output format, or your specific edge cases well. The honest version of evaluation is building a small set of real examples from your own use case, running your actual prompt and pipeline against them, and looking at the outputs yourself before you trust the system with a stranger’s request. That’s slower than quoting someone else’s benchmark, and it’s the only version that tells you anything true about your app.

This also means being honest about what you haven’t tested. If you’ve only tried a pipeline against English text and a user sends it something in another language, or a scanned document instead of clean text, you don’t know what happens. Saying so, and building the guardrail before you find out from a support ticket, is cheaper than finding out after.

Failure has to be a first-class case, not a footnote

APIs time out. Providers have outages. A model occasionally refuses a reasonable request or gets stuck in a repetition loop. In a demo, if any of this happens, you just run it again and cut that part from the recording. In production, every one of those failure modes needs a defined behavior: a timeout value, a fallback message, a way to tell the difference between “the model is down” and “the model gave a bad answer,” because those need different responses. Logging matters here too. When something goes wrong at 2am, you want to know which request, which prompt, which model version, and what came back, not just that an error occurred.

What actually closes the gap

None of the above is exotic engineering. It’s the same discipline that applies to any external dependency: handle the failure modes, respect the rate limits, watch the cost, validate what comes back, and test against your own data before you trust it with someone else’s. The reason it feels new with LLMs is that the demo is so convincing that it’s easy to mistake “the model produced a good answer once” for “the system is ready.” The model was never the hard part. The system around it is.

If you’re building with these tools and want the practical side of it, without vendor pitches or numbers we didn’t measure ourselves, that’s what we cover here.

Visit AI Tool Gazette for more on shipping AI tools, RAG pipelines, and coding assistants that hold up outside a demo.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →