← all articles

What an AI agent actually costs to run

ai-agents llm-costs prompt-caching context-window

The first thing I built to fix an agent’s bill was a summariser. Every few steps it took the conversation so far, compressed it, and put the short version back in place of the long one.

The bill went up.

Two reasons, and I had thought about neither. The summariser read the entire history every time it ran, so I had added a large model call in order to shrink other large model calls. And rewriting the history invalidated the prompt cache from the rewrite point onward, so a prefix that had been costing me a fraction started costing full price again.

That is the shape of most agent cost problems. You reach for a control that looks like it should help, because you are still picturing the thing as a series of messages.

The loop resends everything

A chat completion is one transaction. Text in, text out, both billed, done.

An agent is a loop. The model asks for a tool. Your code runs it. The result goes back. The model asks for the next thing. That repeats until it decides it has finished or until you stop it.

Nothing is held on the provider’s side between those calls. No session, no memory, no warm state. Each step is a standalone request that has to carry the whole history or the model has no idea what happened three steps ago.

So the tenth step carries nine steps of history on its back. It is the most expensive step in the run so far, and the eleventh will beat it.

Riding along on every one of those requests is a block you probably never think about: the system prompt and the full tool definitions. Thirty tools with real JSON schemas and useful descriptions can be two or three thousand tokens. In a twenty step run where the agent touches two of them, you have paid for all thirty, twenty times.

Twenty steps, with made up numbers

Per token prices move and I would rather you kept the shape than a figure that expires. So these are round and invented, and the ratio is the point.

Fixed overhead of 2,000 tokens on every call. Each step appends about 1,000 more, between the model’s own output and whatever the tool returned.

Step one bills 3,000 input tokens. Step five bills 7,000. Step twenty bills 22,000.

Total across the run is around 270,000 input tokens. The estimate most people write down, twenty steps times a 3,000 token request, is 60,000.

That is a factor of four and a half from arithmetic alone, before anything goes wrong. It scales with roughly the square of the step count, which is why an agent that behaves impeccably on a five step task can produce an alarming number on a forty step one.

Output tokens hardly matter in that total. A couple of hundred written per step against tens of thousands read. But output does not vanish once you have paid for it. It becomes input on the next step and on every step after that.

The parts nobody puts in the estimate

Tool output is the largest one, and it never looks like a cost decision when you write it. It looks like plumbing.

My research agent had a fetch step I wrote in about four minutes. It returned the response body, complete. A modern pricing page is a hundred kilobytes of HTML before you read a word: inline scripts, an image or two pasted in as text, several hundred lines of tracking. The agent wanted one table. It got the document, then carried the document through every remaining step.

Nothing in your code punishes a tool that returns too much. It works. The tests pass. It surfaces on an invoice weeks later, blended in with everything else.

Retries are the second one. A timeout, malformed JSON, a schema check that fails and gets handed back. That retry pays the full accumulated context again for a step that produced nothing, and the failed attempt usually stays in the history because you want the model to see what went wrong. Failures late in a long run are dramatically more expensive than the same failures early, since what gets resent is so much bigger by then.

Third is the run that will not stop. An agent that cannot tell it is stuck rephrases and tries again, each attempt a full step at the current context size, with the context still growing. Mine spent nine minutes on a supplier page that had been taken down. Nothing in my monitoring would have described that as broken.

A step limit is the guard everybody adds and it is the wrong unit. Steps are not equal, so a cap of fifty steps caps a number you do not care about. Count tokens as the run proceeds and kill it at your threshold. It is a dozen lines of code.

Fan out deserves its own line. If your framework spawns three subagents to work in parallel, each one starts with its own copy of the system prompt and grows its own history from there, and the parent then pulls all three results into the context it was already carrying. Three helpers do not cost you three extra steps. They cost three extra runs, plus the room their answers take up in the run that called them. Often still worth it for wall clock time. Never the cheap option, whatever the architecture diagram implies.

Last, reasoning tokens. A model that thinks before answering spends output priced tokens you never read, and the amount scales with how hard it judges the problem to be, which you cannot predict from outside. On a single call that is a known trade. Inside a loop you are making that trade on every step.

I do not know the current billing semantics for that across every provider well enough to tell you how prior thinking is treated on subsequent steps. It differs, and it has changed more than once while I have been watching. Read the provider docs, then read your own token counts, and trust the counts.

The order that actually works

Not calling the model at all is worth more than everything below it combined. Every step you delete removes a full sized request, and the steps you delete are usually near the end where they cost the most. My paragraph tagger in the video pipeline was one model call per paragraph until I worked out that a keyword table got most of them right and I was checking the output before rendering anyway.

Decide in code whatever code can decide. If the value is in a config file, read the config file.

Next, keep tool responses small at the source. Inside the function, before the text ever reaches the context. Return the status code and the parsed table instead of the page. Return ten rows instead of the result set. Truncate, and say so in the tool description so the model knows to ask for more. Check for duplicate reads too: mine fetched the same page twice across a run, four steps apart, and then carried both copies.

Then prompt caching, which fits agent loops unusually well. The prefix grows by appending and rarely gets edited, which is exactly the access pattern caching rewards. Instructions and tool definitions at the top and byte identical. Anything that varies, a timestamp, a session id, a counter, at the bottom or removed. A single changing character near the front costs the whole cache. Go and read your hit rate afterwards rather than assuming it took.

Cheap model routing comes last. Choosing between two tools is not difficult work, and neither is judging whether a fetched page contained a pricing table. Those go small. The step that has to reason across everything the run gathered goes large.

Everybody reaches for that last one first. It works, briefly, which is the problem.

Sitting outside the ranking because it is a one time edit rather than a habit: send only the tools a task could plausibly use. My mail agent had inherited the full tool list from the research agent, including four things it had no business touching. Trimming that list on the mail code path took ten minutes and removed a couple of thousand tokens from every step of every run it has done since.

Most of this is context management with a price tag on it

The bill is a readout of how much text you move through the loop and how many times you move it. Swapping in a cheaper model divides the readout by a constant and leaves the loop exactly as you built it, so the relief lasts until the agent gets a slightly longer task and surprises you again.

I will happily argue this one. Before you change models, log two numbers per run: total tokens, and the size of the largest single tool response. I have never looked at that pair and concluded the model was the expensive part.

Where I am less confident is scale. My agents run ten to forty steps and finish in minutes. Anyone running something for hours, across hundreds of steps with a real memory layer underneath, is solving a different problem and I would not tell them what their answer looks like.

The version of my research agent that finally behaved was duller than the summariser I was proud of. A fetch tool that returns extracted text with a cap, and a shorter tool list per task type. Nothing clever.

Current per token prices and the cost notes I keep on each agent framework are here.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →