When an AI Worker Dies Halfway Through
I had a batch of parallel workers running, each writing a file, and about a third of the way through the provider started returning server errors. Not a few. A wave. When it settled I had files that existed, files that did not, and a handful that existed but were half written.
That last category is the one worth talking about, because nothing in my setup was designed to handle it.
Agents break the assumption retries are built on
A plain API call is atomic from your side. Send a request, get a response or an error, and either way nothing in your world changed except that you now hold some text. Retry costs you nothing but time.
An agent works differently. It acts as it goes: reading files, writing files, running commands, calling services. At any moment during its run it has already changed things, and if it dies at that moment there is no rollback and no transaction. Those changes sit there representing some fraction of a task that never finished.
So “the worker failed” is messier than it sounds. It rarely means nothing happened. It usually means something happened and you do not know how much.
Why they die
Provider-side errors are the most common cause in my experience, and their important property is that they arrive correlated. You do not lose one worker. You lose eight, inside the same thirty seconds, because whatever is happening sits upstream of all of them. Any design assuming failures are independent and rare will be wrong exactly when it matters.
Timeouts are second, yours or theirs. Worse than a clean error, because a worker you have written off may still be alive on the provider’s side, still doing things, while your orchestrator has moved on.
Local resource exhaustion is third. Forty workers on a machine that handles fifteen, and the OS starts killing things. Self-inflicted, obvious once seen.
Fourth and rarer, the worker gets stuck. Not erroring, not finishing, looping or waiting on something. From outside, a stuck worker and a slow worker look identical.
The retry is where it actually hurts
Your orchestrator sees a failure and retries, which is the sensible default. But the retry starts from scratch while the dead worker’s partial output is still on disk. Now a second worker is producing the same artifact, and depending on how your code writes, you get either a clean overwrite or duplication.
I have had both. Clean overwrite happens when the output path is deterministic, derived from task identity, so the retry lands on the same file and replaces it. The bad case is any output path containing something that varies, a timestamp, a counter, a generated name, because the retry creates a second file and you end up with two versions of the same work and no obvious way to tell which is good.
Make the output location a pure function of the task. Same task, same path, every time. That single decision turns retries from a correctness problem into a non-event.
Write atomically
Never write directly to the final path. Write to a temporary file and rename it into place on completion, since rename on the same filesystem is atomic everywhere you care about.
This is what handles the half-written category. A file that exists but is truncated is genuinely dangerous, because anything downstream checking whether the file exists concludes the work is done. An empty or partial file passes an existence check perfectly.
That lesson generalises well past AI workers. It has bitten me elsewhere with audio files, where a truncated output shipped downstream because the file was there and nobody checked its length. The existence of an artifact is not evidence that the artifact is complete.
Keep a manifest, written before the work starts
A small record, one entry per unit of work, with a state field. Before dispatching a worker you mark the task claimed. When it finishes and you have verified the output, you mark it done. Anything still claimed when the run ends started and did not finish, and that is exactly the list you want.
Without it you are reconstructing history from which output files exist, which is guesswork. With it you can restart the batch and skip everything already done, which is the property you actually want from a long-running job.
Verify the artifact, not the exit status
A worker exiting cleanly has not necessarily produced anything useful. An agent can decide it is finished, explain what it did, and exit zero without ever writing the file you wanted.
Check the output instead. Does it exist, is it plausibly the right size, does it parse if it is meant to be structured. Cheap, and it catches a class of silent failures no status code will.
Resume, do not re-run
People get this wrong constantly, usually because re-running the batch is easier than resuming it. With the manifest above, resuming is nearly free, and re-running everything means paying again for work that already succeeded. At any real scale that is serious money.
Back off, and check whether you caused it
If the provider is erroring at eight workers simultaneously, launching eight replacements immediately makes it worse. Use exponential backoff with randomness in the delay so retries do not land at the same instant, and when the error rate is high, drop concurrency rather than holding it. A sustained wave is a reason to stop the run and look, rather than grind through hundreds of failed attempts.
Worth saying plainly: sometimes you caused it. Fire off far more parallel work than your account allows and the errors coming back are the system telling you so. No retry logic fixes that. Check your own concurrency before blaming the provider.
Isolate workers from each other
If two agents can touch the same file, eventually they will, and the result is hard to attribute afterwards. Give each worker its own output directory, or its own copy of the working tree when they are modifying a codebase. That isolation costs a little setup time and removes an entire category of confusing bug.
Files are the easy case
Everything above assumes the partial work is a file. Files forgive you. You can overwrite one, delete it, check its size, rename over it, and a retry landing on the same path leaves a clean state.
External side effects do not forgive you. If your agent sent an email, posted something, made a payment, created a record in someone else’s system, or called a per-call paid API, a retry does not correct it. It doubles it. And you cannot tell from your side whether the dead worker got as far as the send, because the thing that would have told you is the thing that died.
My position on this, and it is a position rather than a consensus: keep agents away from irreversible external actions. Have the agent produce an artifact describing what should happen, then let separate, deterministic, extremely boring code apply it.
That split buys a lot. The agent step becomes retryable, since producing the same plan twice is harmless. The applying step is ordinary software you can make idempotent in the usual ways, with a request ID, a uniqueness constraint, a check before write. And when something goes wrong you can read the plan the agent produced and see what it intended, which is impossible when the intention only ever existed inside a process that has since died.
It costs a little immediacy. It saves you from discovering at scale that some of your workers sent things twice.
There is a smaller local version: child processes. An agent that runs commands may have spawned things, and killing the agent does not necessarily kill what it started. You end up with orphans holding files open, holding ports, consuming resources long after their worker is gone. If you kill workers on a timeout, kill the process group rather than the process.
Failed work still costs money
Every worker that died partway through burned tokens on the way there, and you are paying for all of them.
That spend appears in no success-oriented metric, because it is attached to runs you threw away. Track it deliberately. A fifth of your workers failing on expensive tasks is a real line on the bill, larger than you would guess, and invisible unless you count failed attempts alongside successful ones.
On the stuck worker
You need a timeout, and you need to accept it will occasionally be wrong.
There is no reliable way from outside to distinguish genuinely slow work from a wedged process. So pick a bound based on how long the work should take, kill anything past it, and accept that you will sometimes kill something that would have finished. That is the correct trade, because a wedged worker holds a slot forever, and one of those in a pool of ten costs you ten percent of throughput indefinitely.
If you can get progress signals out of the worker, time out on lack of progress rather than total elapsed time. Better signal, and it lets you be generous with genuinely long tasks.
What I am unsure about
Where the line sits between checkpointing inside a worker and keeping workers small enough that losing one does not hurt.
I have gone with small units, because a worker doing one bounded thing is easy to retry and easy to verify, and internal checkpointing adds complexity I have not needed. If your individual units are genuinely long and expensive that calculation changes, and I have not run that shape at enough scale to say where it flips.
The reframe is the part to keep. We are used to API calls, where failure means nothing happened. Agents break that, and every piece of advice above follows from taking it seriously.
For more breakdowns like this on how AI tools actually work under the hood, head back to the AI Tool Gazette homepage.