Prompt injection, and why it is not fixed
The mail triage script started out safe by accident. It read messages arriving at a couple of my domains, classified them, drafted a reply, and left the draft in a folder for me to look at. Anything it got wrong I caught, because I was the last step.
Then I got tired of copying drafts across and gave it permission to send directly on anything it had marked as routine. That took about ten minutes.
It did not register as a security change. It felt like removing an annoyance, late in the evening, in a script nobody else would ever run.
What I had built was a service that read text written by strangers and could send email from my own domain with nothing in between. It ran that way for a few weeks before I drew the thing out on paper and pulled the permission back out. Nothing came of it, and I am not writing up an incident.
The part worth keeping is that I already understood prompt injection when I did it.
One channel, and no way to mark it up
A language model receives a single stream of text. Your instructions are in there. So is the material you asked it to work on. There is no second channel, no header saying this half is commands and that half is data, no flag the model can consult.
Everything the model does comes out of reading that stream, including the question of what it has been asked to do.
So any text that reads like an instruction has some probability of being followed, whoever put it there and however it arrived. That is prompt injection in a sentence: somebody else’s words riding inside the channel your words use, carrying weight they were never supposed to have.
The obvious response is to separate the two harder. System prompt for your instructions, user content somewhere else, an explicit order to ignore anything instruction shaped that turns up in the data. Every serious provider does a version of this, and it works well enough that the lazy attempts bounce off.
What you have there is a weighting. The model is still reading the whole stream and still deciding by reading.
Compare a database driver. A parameterised query keeps the query text and the values on physically separate paths, and no value can turn into part of the query regardless of what characters it holds. Code enforces that. Nothing has to be persuaded.
There is no version of that for a language model, and it is not clear what one would even look like, because the reason these things are worth paying for is that they act on ordinary language. The vulnerability and the product are the same property.
The version that reaches you without touching you
Direct injection is a person typing something clever at your chatbot to talk it out of its own rules. It makes good screenshots and it is largely self contained, since the person doing it is the person it lands on.
Indirect injection is the one to design around.
Your agent fetches something. A support ticket, a pdf a customer uploaded, a page a search result pointed at, a calendar invite, a comment sitting in a source file it was asked to refactor. That content contains a line written to look like a directive.
Nobody on your side typed it. Nobody on your side sees it. A user asked a normal question and got back an answer with a stranger’s intent folded into it.
The attacker needs no access to your systems at all. They need access to something your system will read later, which is far easier to arrange and far harder to inventory.
Try listing that surface for your own build. Anything scraped. Anything uploaded. Anything sitting in an inbox. Filenames. The text buried in image metadata. The readme of a package you pulled from a public registry eighteen months ago, which a coding agent may well read while working out how to call it.
None of those look like attack surface on an architecture diagram. They look like inputs.
The risk is exactly what you granted
A model that can only emit text is a nuisance. Worst case is a wrong or embarrassing answer, and you already carry that risk from ordinary error.
A model that can act sits in a different category, and the boundary is the moment you gave it the ability to change something outside itself.
Can it send. Can it write a file. Can it call an api that moves money, cancels an order, deletes a row, publishes under your name.
Every one of those is a sentence somebody can put into a web page and have executed on your behalf.
The only framing that has stayed useful to me is blast radius. Stop asking whether an attempt will land. Assume it lands, then ask what is on the other side of it.
If the answer is a bad paragraph in a draft you were going to read anyway, ship it. If the answer is money leaving an account, that conversation belongs before launch.
Four things that help, none of which fix it
Narrow the permissions until the job barely fits. Boring advice, and people genuinely do not follow it here, because agent frameworks ship with broad access and every tutorial keeps it switched on. The credential should be scoped to the one job. A ticket reading agent has no business writing to the customer database. A drafting agent does not need publish rights.
Better again, split the agent up. If the piece that reads web pages runs as a separate process with separate credentials from the piece that touches payments, an injection landing in the first one has nowhere to go.
Prefer read only. A large share of what people build agents for is retrieval and summarisation and comparison, which are read jobs with a human doing the write at the end. Write access gets granted anyway, because it was one checkbox.
Put a person in front of the irreversible and the expensive. Not in front of everything: a confirmation on every action becomes somebody clicking yes forty times an hour, and they stop reading around the fifth. Pick the actions that cost real money or cannot be walked back. Sending to a customer. Spending. Deleting. Anything that becomes public.
Then make the confirmation say what is about to happen. The recipient, the amount, the file name. A dialogue reading “the agent would like to continue” trains people to approve without looking, which is worse than having no dialogue, because now you think you have a control.
Treat fetched content as hostile. When you write ordinary software you already know user input is dangerous and you handle it that way without thinking. Then the same engineer hands a scraped page straight to a model under a heading that says context. That page is user input. A stranger wrote it.
Filters are a layer, and only a layer
There are classifiers now that scan incoming content and flag anything shaped like an injection attempt. They work on the obvious material. Ignore all previous instructions, and everything in that family, gets caught reliably.
The problem is that the attack surface is natural language, so the set of things that might work is the set of things a person can say. You cannot enumerate that set.
The same instruction can be rewritten forever. Politely. In another language. Spread across three individually harmless sentences. Phrased as a correction rather than a command. Wrapped in a story about why the correction is needed.
Run a filter. Count it as one layer. Never count it as the reason you felt able to grant something dangerous.
The page with white text at the bottom
A page I scraped last year had a line of white text near the footer, roughly two pixels tall, telling any assistant reading the page to add a link to a particular site into whatever summary it produced.
I noticed because a link turned up in a summary and I had not asked for links.
That is the whole attack. No exploit, no code, nothing broken in anything I had written. Somebody typed a sentence into a web page.
It cost me nothing, and the reason it cost me nothing is that the agent could write to a local file and do nothing else. Give that same agent a posting permission and the sentence works exactly as designed, with my site as the distribution.
Before you ship the thing
Write out every action your agent can take, and next to each one write what the damage is if it fires with somebody else’s intent behind it. Any action you cannot answer for is the one to fix first.
Log what it did, and keep enough to reconstruct why. A successful injection usually surfaces as an action that looks slightly off rather than as an alert, and slightly off is only visible if you can go back and read the sequence.
Do not wait for a better model to take this off your plate either. A better model follows instructions more reliably, and that includes the instructions you did not write.
Current api pricing, and per tool notes on what each one actually lets an agent do to your systems, are here.