← all articles

Coding agents vs autocomplete: the split I actually use

ai coding developer-tools ai-agents

An agent wrote me a tidy little retry wrapper in June. Exponential backoff, jitter, sane defaults, about forty lines, worked first try.

We already had one. Same behaviour, four directories over, written by me in March.

It had pulled in eleven files before it started writing. Ours was in the twelfth. Nothing about that was a model failure. From where it was standing, our helper did not exist.

I run both most days on the same repos: the grey text that finishes my line as I type, and the thing I hand a whole task to and walk away from. They get sold as one product category and they fail in completely different ways. What separates them has little to do with which is smarter. Half the time it is the same model underneath.

The axis that matters is unread code

How much work piles up between you and the result without anyone looking at it.

Completion never lets that number climb. The suggestion appears, you read it in a fraction of a second, you take it or you keep typing. Your review backlog sits at zero because reviewing is the same motion as writing.

An agent inverts that. The unit of work is a task, so the unit of review is a task. You get handed a diff you did not watch being made, and you owe the codebase a careful read you will probably not perform at full attention, because reading code is slower and duller than writing it.

Almost everything else about these two falls out of that.

What completion is actually selling

Latency. Not intelligence.

The suggestion arrives inside the same second you were already thinking, so taking it costs nothing. You were going to type that line anyway. Something you had already decided to do took 200 milliseconds instead of eight seconds, twenty times an hour.

So it earns its keep on the code you resent. I wired fifteen brand configs into a content pipeline last month, each one a near copy of the last with different slugs, different number bands, different output paths. Completion did most of that. The fourth almost identical error branch. The test that is the previous test with two values changed.

Boilerplate you can check at a glance, where wrong is obvious instantly.

Completion agrees with you, including when you are wrong

Took me an embarrassing amount of time to spot this one.

Completion is very good at continuing a pattern and has no opinion about whether the pattern is any good. I had three cron scripts on a box that swallowed exceptions and exited zero, so a failing job looked identical to a successful one in the logs for six weeks. When I wrote the fourth, my editor offered me the same swallow, correctly indented, matching its neighbours.

Consistency with a mistake spreads the mistake.

So completion accelerates whatever direction you were already going. In a repo you have kept tidy, that is a gift. In one drifting somewhere unfortunate it is a tailwind pointed the wrong way, and the deeper you have entrenched the problem the more correct every suggestion feels.

There is a second cost, harder to notice. You stop composing lines and start grading proposals. Grading is much worse at catching that the whole approach is wrong, because your attention sits on a detail while the direction goes unexamined.

Where I hand over the whole task

Three situations, narrower than the marketing suggests.

Mechanical change across many files. Renaming a concept everywhere, moving every call site onto a new signature, updating the same pattern in forty places. Tedious, well defined, and the tests either pass or they do not.

Unfamiliar territory. When I do not know a library’s shape, an agent that reads the installed source and tries something beats me reading documentation, because it closes the loop by running the thing.

Anything with a fast automatic check bolted on. A failing test, a type error, a build that either completes or does not. When a machine can say yes or no in seconds, an agent iterates against it far faster than I can.

And where they take a day off you

Agents are strongest wherever verification is cheap and automatic. They are worst wherever it is expensive.

The failure I hit most is the confident wrong direction. It hands back something that runs, passes whatever checks exist, and solves a subtly different problem. The code itself is fine. What it does is answer a question I never asked, and because the answer is coherent it survives a casual read.

Obviously broken output announces itself. Plausible output aimed at the wrong target gets committed and found in November.

The other one is scope. Give a vague instruction, get a large diff, and somewhere in it sit four changes you never asked for that seemed reasonable at the time. I asked for a change to a renewal window once and got expiry handling helpfully tidied up in three other places along the way. Now the review is archaeology, and reading code you did not write is slower than writing it, which is precisely the cost the tool was supposed to remove.

The rule I ended up with

Task size has nothing to do with it. The question is who gets to judge the result.

Hand it to the agent when a machine can tell you whether it succeeded. Keep your hands on the keyboard when the only judge is your own taste.

Anything touching money, auth, or a schema migration stays manual here, and I will argue about that with anyone. A model can write a migration perfectly well. My problem is that a green test suite says nothing about whether a billing table is correct, and I find out I was wrong when a customer emails me.

Refactors, plumbing, test scaffolding, anything with a red or green light at the end: hand it over and go do something else.

Retrieval, not reasoning, is the ceiling

Back to that duplicated retry helper.

An agent can only reason about what it has pulled into context. On a codebase of any size that is a small fraction and it cannot read the rest. Everything outside the slice may as well not exist, which is how you get a perfectly reasonable utility function written from scratch a short walk from its twin.

This is why naming files beats writing an elegant brief. An instruction that opens with “here are the four files this touches” beats a beautiful paragraph about intent every time, because you have solved the retrieval problem the agent cannot solve for itself.

It also explains why agents look brilliant on a small project and turn erratic on a large one. The model did not change. The fraction of the relevant world it can see collapsed.

And some of that world sits in no file at all. We pin one client library to HTTP/1.1 because of a threading bug that cost me most of a Saturday. Nothing in the repo says why. An agent reading that line sees a weird constraint and, given a free hand, tidies it away.

Two review habits worth more than the tool choice

Demand small diffs. One task, one concern, one review. A broad instruction produces a broad change, and a broad change is where unread code hides. Three narrow instructions in sequence give you three diffs you can hold in your head.

Read the diff before you run the tests. Not after.

Once you have seen a green tick your standards drop, measurably and without asking permission. You start skimming. A passing suite proves the code does not fail in the ways somebody thought to check, which is a far weaker claim than it feels like at one in the morning.

Tests are where my own rule breaks

I used to give test writing straight to an agent. Tests run, so verification is automatic and cheap, which by my own rule makes them an ideal handover. I was wrong about that for months.

What came back, over and over, were tests that exercised the code as written instead of the behaviour I wanted. They asserted that the function does what the function does. All green, coverage climbing, and not one of them would have caught a real regression, because they came from the implementation rather than from what it is supposed to guarantee.

The check is automatic and the check measures the wrong property. So now I write the assertions myself, or I write the intended behaviour down in words first and have the agent work from that, never from the finished code.

Small distinction, and it is the difference between a suite that protects you and one that only makes your build slower.

The bill arrives late

The pricing argument gets the most airtime and deserves the least. A few dollars to a few tens of dollars a month, plus tokens on long agent sessions, set against an hour of your own time.

The cost that hurts sits in a different account. It is the code you accepted without understanding, found eight months later, and have to reverse engineer at the worst possible moment.

Economise on unread code. The tokens are cheap.

What I would tell someone starting

Use completion constantly. It is nearly free, the review cost is zero because you see every character, and when it goes wrong you find out immediately.

Use agents deliberately: on tasks with an automatic check, in small pieces, with the files named up front, and read the diff properly.

Then there is a third category nobody advertises. Some tasks are faster to do yourself, and the tell is that you have spent longer explaining the task than doing it would have taken. I still catch myself negotiating with a machine over fifteen minutes of typing, usually around 11pm, because handing it over feels like progress and typing feels like work.

Current pricing and the full side by side comparison table live here.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →