← all articles

How to add streaming responses to an LLM app

A long answer from a big model can take tens of seconds to finish. If your app waits for all of it before showing anything, users stare at a spinner and assume the thing is broken. Streaming pushes tokens to the screen as the model produces them. Total time barely changes, but the first words show up almost immediately, and that’s the part people feel.

This is for anyone with a working non-streaming LLM call and a small web frontend: a support widget, an internal tool, a chat box on a content site. I’m assuming Python on the server and plain JavaScript in the browser. By the end you’ll have a FastAPI endpoint that streams server-sent events, a browser reader that renders them, a working stop button, and the proxy settings that stop it all arriving in one lump.

One limit up front. I wrote the server side against Anthropic’s Python SDK because it’s the one I know best. Other providers stream over server-sent events too, so steps 3 to 9 shouldn’t change if you use one of them. Only the call in step 2 does.

what you need

  • python 3.10 or newer, then pip install anthropic fastapi "uvicorn[standard]"
  • an Anthropic API key exported as ANTHROPIC_API_KEY. It’s pay as you go, and a handful of test prompts costs cents. Streaming doesn’t change per-token pricing, but output tokens cost more than input tokens, so read reading an LLM provider pricing page properly before you budget.
  • one HTML file you can edit for the frontend
  • curl. In Windows PowerShell, plain curl can be an alias for Invoke-WebRequest, so type curl.exe
  • nginx or another reverse proxy, only if you deploy behind one. Get it working locally first.

step by step

1. Get the plain call working first

Confirm the non-streaming version works before you change anything. Put this in app.py. I read the model name from an env var so I can swap models without a deploy; the default here is claude-opus-5.

import os
from anthropic import AsyncAnthropic

client = AsyncAnthropic()  # reads ANTHROPIC_API_KEY
MODEL = os.getenv("LLM_MODEL", "claude-opus-5")

async def answer_once(prompt: str) -> str:
    msg = await client.messages.create(
        model=MODEL,
        max_tokens=16000,
        messages=[{"role": "user", "content": prompt}],
    )
    return next(b.text for b in msg.content if b.type == "text")

Expected output: one complete string, after a pause that grows with the answer length. If it breaks: a 401 means the key isn’t set in this shell, and a NotFoundError means the model id is wrong.

2. Switch to the SDK’s stream helper

The raw API sends server-sent events such as message_start and content_block_delta (all listed in Anthropic’s streaming docs). The SDK parses them for you. Streaming also avoids HTTP timeouts on long outputs.

async def answer_stream(prompt: str):
    async with client.messages.stream(
        model=MODEL,
        max_tokens=16000,
        messages=[{"role": "user", "content": prompt}],
    ) as stream:
        async for text in stream.text_stream:
            yield text

Expected output: call it from a scratch script with print(piece, end="", flush=True) and the words appear progressively. If it breaks: everything arriving at once means you’re still on create, or your terminal is buffering.

3. Wrap the chunks as server-sent events

An SSE frame is one or more data: lines ended by a blank line. I put JSON in each frame so newlines inside the model’s text can’t break the framing. Errors need care: once the first byte is out, the status is already 200, so a failure has to travel as an event. I use SSE, not WebSockets: chat is one request out and one stream back, and SSE works through ordinary HTTP proxies.

import json
import anthropic

def sse(payload: dict, event: str | None = None) -> str:
    head = f"event: {event}\n" if event else ""
    return f"{head}data: {json.dumps(payload)}\n\n"

async def sse_events(prompt: str):
    try:
        async with client.messages.stream(
            model=MODEL,
            max_tokens=16000,
            messages=[{"role": "user", "content": prompt}],
        ) as stream:
            async for text in stream.text_stream:
                yield sse({"text": text})
            final = await stream.get_final_message()
        yield sse({"stop_reason": final.stop_reason,
                   "output_tokens": final.usage.output_tokens}, event="done")
    except anthropic.APIError as e:
        yield sse({"message": type(e).__name__}, event="error")

Expected output: nothing visible yet. If it breaks: an import error on anthropic means it isn’t installed in this environment.

4. Expose it with FastAPI

FastAPI’s StreamingResponse takes an async generator directly. Two headers matter: Cache-Control: no-cache, and X-Accel-Buffering: no, which tells nginx not to buffer this response.

from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from pydantic import BaseModel

app = FastAPI()

class ChatIn(BaseModel):
    prompt: str

@app.post("/chat")
async def chat(body: ChatIn):
    return StreamingResponse(
        sse_events(body.prompt),
        media_type="text/event-stream",
        headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
    )

Run uvicorn app:app --port 8000. Expected output: Uvicorn running on http://127.0.0.1:8000. If it breaks: a 422 from the endpoint means the body isn’t JSON with a prompt field.

5. Test with curl before touching the browser

Save {"prompt": "Explain SSE in two sentences."} as body.json, because quoting JSON on the Windows command line is miserable, and run:

curl.exe -N -X POST http://127.0.0.1:8000/chat -H "content-type: application/json" -d "@body.json"

Expected output, illustrative since your token boundaries will differ:

data: {"text": "Server-sent"}

data: {"text": " events stream"}

event: done
data: {"stop_reason": "end_turn", "output_tokens": 41}

If it breaks: one big lump at the end means something is buffering. Check you passed -N, then look for a proxy in the path. Sort it out now, because the browser will only hide the problem.

6. Read the stream in the browser

EventSource only does GET and can’t send a body (MDN’s guide covers its limits), so a chat POST needs fetch and a manual reader. Use TextDecoderStream: it handles multi-byte characters split across chunks, which you’ll hit the first time someone chats in Chinese.

async function ask(prompt, onText, signal) {
  const res = await fetch("/chat", {
    method: "POST",
    headers: { "content-type": "application/json" },
    body: JSON.stringify({ prompt }),
    signal,
  });
  if (!res.ok) throw new Error("http " + res.status);
  const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();
  let buf = "", done = null;
  for (;;) {
    const { value, done: eof } = await reader.read();
    if (eof) return done;
    buf += value;
    let i;
    while ((i = buf.indexOf("\n\n")) >= 0) {
      const frame = buf.slice(0, i);
      buf = buf.slice(i + 2);
      const event = /^event: (.+)$/m.exec(frame)?.[1];
      const data = JSON.parse(/^data: (.+)$/m.exec(frame)[1]);
      if (event === "error") throw new Error(data.message);
      if (event === "done") done = data;
      else onText(data.text);
    }
  }
}

Expected output: text appears word by word in your page. If it breaks: garbled characters mean you’re decoding chunks by hand, and TypeError: Failed to fetch usually means CORS or a wrong URL.

7. Add a stop button and handle the end state

Pass an AbortController signal into ask. When the browser disconnects, Starlette cancels the generator, the async with exits and the upstream stream closes, so you stop generating tokens nobody reads. I haven’t verified how every provider bills a cancelled generation, so check yours.

const out = document.querySelector("#out"), stopBtn = document.querySelector("#stop");
const ctrl = new AbortController();
stopBtn.onclick = () => ctrl.abort();
ask(prompt, t => (out.textContent += t), ctrl.signal)
  .then(d => { if (d?.stop_reason === "max_tokens") out.append(" [cut off]"); })
  .catch(e => { if (e.name !== "AbortError") out.append(" [error: " + e.message + "]"); });

The done payload carries stop_reason. max_tokens means the answer was cut off and refusal means the model declined, so tell the user in both cases instead of showing a silently short answer. If it breaks: put a print in a finally block inside sse_events. It should fire when you press stop. If it doesn’t, something upstream is holding the connection open.

8. Log time to first token

The payoff from streaming is time to first token (ttft), so measure it. Add a timer around the loop from step 3:

import time  # top of app.py

started = time.perf_counter()  # first line of sse_events
first = None
# inside the async for loop, before the yield:
if first is None:
    first = time.perf_counter() - started
# after get_final_message():
print(f"ttft={(first or 0):.2f}s total={time.perf_counter() - started:.2f}s")

Expected output: one log line per request. Total should look like the non-streaming version and ttft should be far smaller. If ttft is long, see the thinking pause in the pitfalls below.

9. Turn off buffering at every hop

nginx buffers proxied responses by default. The proxy_buffering directive turns that off, and the header from step 4 does it per response.

location /chat {
    proxy_pass http://127.0.0.1:8000;
    proxy_http_version 1.1;
    proxy_buffering off;
    gzip off;
    proxy_read_timeout 300s;
}

Turn gzip off for that location as well, since compression waits to fill a block. Any CDN, load balancer or AI gateway in front is another hop that can buffer. Test with curl.exe -N against the real hostname, not localhost. If it breaks: compare response headers on localhost and production.

common pitfalls

  • Testing only on localhost. Buffering shows up in production, behind nginx, a CDN or a gateway, never on your laptop. Run step 5 against the deployed hostname every time.
  • Ignoring the thinking pause. On claude-opus-5, thinking runs by default and its text is hidden by default, so the stream can sit silent before the first word. Show a “thinking” indicator, or lower the effort on chat routes with output_config={"effort": "low"}.
  • Streaming JSON you plan to parse. Half an object doesn’t parse. Stream prose to the user, and for machine-readable output wait for the final message and validate it, as in getting structured output that validates.
  • Re-rendering markdown on every token. A full parse and layout per token makes long answers janky on a mid-range phone. Buffer the text and flush once per requestAnimationFrame.
  • Retrying blindly mid-stream. The SDK retries requests that fail before the stream starts. After the first token, a dropped connection leaves partial text, and a silent retry would print the answer twice. I keep what arrived and show a retry button.

scaling this

  • 10x: nothing to do. An async server handles many streams in one process because they spend nearly all their time waiting on the network. Watch ttft from step 8 and leave the rest alone.
  • 100x: think in open connections. Concurrent streams equal new chats per second times average duration, so 10 chats a second at 20 seconds each is 200 sockets held open at once (invented numbers, use your own). Each also holds an upstream connection to the provider. Add uvicorn workers, raise the file descriptor limit, and check your proxy’s connection and read-timeout limits. Provider rate limits arrive as 429s before the stream starts, and the SDK retries those twice with backoff by default.
  • 1000x: split the streaming tier from the rest of the app so a slow model can’t starve your normal endpoints. Cap concurrent streams per user and decide what happens when you’re full, a queue or a plain “busy” message. A gateway starts earning its keep for keys, limits and fallbacks, and the article linked in step 9 covers what it costs. If you self-host, vLLM vs TGI both stream, and only step 2 changes, because the browser reads your event format from step 3, not the provider’s.
  • at any scale: decide early whether you log full transcripts. They can hold personal data, and The Privacy Wire covers that side. This is not legal advice.

where to go next

More on the blog index.

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-21.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →