Agents vs workflows: when you actually need autonomy
I run a handful of scripts that write, check, and publish content every day without me touching them. None of that is “agentic” in the way people throw the word around online. It’s a workflow: fixed steps, fixed order, an LLM call plugged into each stage. The confusion between that and an actual agent is costing people money, because agents cost more, fail differently, and need guardrails workflows don’t.
This gets murky because both use the same model underneath and both can call tools. The difference is who decides what happens next. Get that wrong and you either overbuild a simple job into something unpredictable and expensive, or underbuild an open-ended task into a brittle script that breaks the moment reality doesn’t match your flowchart.
what it is
A workflow is a system where the code determines every step in advance. You write the sequence: fetch data, hand it to the model with a prompt, take the output, pass it to the next step, maybe another model call, then write the result somewhere. The LLM produces text at each stage, but it never decides which stage comes next, that’s your code. Anthropic’s engineering team draws this line clearly in Building effective agents: a workflow orchestrates models and tools through code paths a developer lays out ahead of time.
An agent is different. The model decides its own next step, in a loop, based on what it observes. You give it a goal, a set of tools, and it plans, acts, checks the result, and decides whether to act again or stop. Claude Code, the tool I’m using to write this, is an agent: it reads files, runs commands, sees the output, and decides what to do next based on what it finds. Nobody scripted the exact sequence of tool calls it just made to open my instructions before writing this sentence.
how it works
Same underlying mechanism, different control flow. Both rely on tool calling (also called function calling), where the model outputs a structured request like “run this function with these arguments” instead of a plain text answer, and the surrounding system executes it and feeds the result back in.
In a workflow, that loop is capped and directed by you. Step one calls the model with prompt A, step two takes that output and calls the model with prompt B, and so on to a defined end. If you’ve built a chain like summarize, then classify, then route to one of three prompts depending on the classification, that’s a workflow even if it has branches. The branches are still ones you defined.
In an agent, the loop is open-ended and the model drives it. It gets a system prompt, a goal, and a toolbox, maybe a file reader, a web search, a code executor. Each turn it decides which tool to call, evaluates the response, and loops again, until it hits a stop condition: either it decides the task is done, or it hits a limit you set, like a max turn count, a token budget, or a timeout. The Model Context Protocol is one of the standards that’s made this practical at scale. It gives agents a common way to discover and call tools across different systems instead of every integration being bespoke glue code.
OpenAI’s practical guide to building agents makes roughly the same split, and adds a useful practical note: start with a single model and a small toolset, and only add a full agent loop, multi-agent handoffs, or a supervisor pattern once a single call plus retrieval genuinely can’t do the job.
why it matters
Cost and latency. An agent loop can call the model five, ten, twenty times chasing one task, each call burning tokens and adding seconds. My publish pipeline runs a fixed three or four model calls per article. Turning that into an agent that “figures out” how to write and publish a post would mean paying for a loop every single day for a job that doesn’t need judgment, it needs repetition.
Handling the long tail. Some tasks genuinely can’t be flowcharted because you don’t know the steps until you’re in it. Debugging a failing build, researching a topic where the right next search depends on what the last one turned up, triaging a support ticket that could go five different directions. That’s where an agent earns its cost, because a workflow would need a branch for every case you can think of, and there’s always a case you didn’t.
Predictability versus flexibility, and you don’t get both for free. A workflow does the same thing every time, which makes it auditable and cheap to debug when it breaks, you know exactly which step failed. An agent can solve a problem you didn’t anticipate, but a bad run can also spiral, call the wrong tool repeatedly, or burn through a budget without finishing. I’ve had a coding agent get stuck re-reading the same file over and over, convinced the fix was in there. It wasn’t. I had to stop it and point it elsewhere.
Where the two actually meet. Most real systems aren’t purely one or the other. A workflow step can hand off a bounded, well-defined subtask to an agent (research this one topic, then return a summary), and an agent can call a workflow as one of its tools. Treat it as a spectrum you’re choosing a point on for each piece of the system, not a single decision for the whole product.
common misconceptions
“Agent” is marketing for anything that uses AI. A lot of products labeled “AI agent” are workflows with good UX. If the tool follows the same script every time regardless of what it finds along the way, it’s a workflow. Nothing wrong with that, most production software should be one, but call it what it is.
More autonomy is more advanced, so it’s the better choice. Not for most jobs. Anthropic’s own guidance is blunt about this: start with the simplest setup that works, and only add complexity when it demonstrably improves the result. A workflow that runs the same four steps reliably beats an agent that’s right 90% of the time and expensive or unpredictable the other 10%, especially for anything customer-facing.
Agents don’t need supervision once they’re set up. They do, arguably more than workflows, because their failure mode is different: not a wrong output at a known step, but doing something you didn’t expect, possibly several tool calls deep before anyone notices. That’s why every serious writeup on this, Anthropic’s included, talks about sandboxing, permission scoping, and hard stopping conditions as part of the design, not an afterthought. If your agent has broad tool access and no budget cap, you’ve built a system whose downside you haven’t measured.
You have to pick a lane for the whole project. No, you’re picking per task, and the same product usually has both. My content pipeline is a workflow end to end for anything routine, and the parts of my day where I actually reach for an agent are the ones where I don’t know the steps until I’m in the middle of them.
where to go from here
A few places to go deeper once you’ve got the workflow versus agent split straight. If your next question is what an agent actually needs to safely reach outside its own sandbox, tools, files, other services, that’s worth pairing with something on data exposure and permission scoping, which theprivacywire.com covers from the security side more than I do here.
On this site: if you’re deciding whether your use case even needs an LLM making decisions versus just retrieving the right document, read our piece on retrieval-augmented generation versus fine-tuning. If you want the mechanism behind how agents and workflows both talk to tools in a standard way, see how to give an agent tools it can actually call. And if you’re trying to get more reliable output out of either architecture without redesigning it, our explainer on how much context a model can actually use is the practical next step. For everything else we’ve written in this vein, the blog index has the full list.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-09.