What the model context protocol actually does
I counted what my client was shipping to the model one afternoon last month. Nine servers attached, 63 tools between them, every description going out in full on every request. What I had actually asked for was a file rename.
That is where most setups land about a month after someone discovers this protocol.
So before the explanation, the number. 63 descriptions at roughly 50 words each is over 3,000 words of overhead riding on a nine word question, and I paid that on every call that week. In tokens, and in the pause before the first word came back.
The duplication it deletes
For about two years the same code got written over and over. A wrapper so the model could search the web. Another to read from Postgres. Another for whatever ticket system the company was on. Switch frameworks and you wrote all three again, because every framework had its own way of describing a tool to a model.
That description is the whole game, incidentally. Name, purpose, arguments, return shape. Get it wrong and the model calls the thing at the wrong moment with the wrong arguments, and you spend a week blaming the model.
So the same integration, against the same system, got built four times by four teams in four incompatible shapes. MCP exists to stop that.
Two roles and one question
A server exposes some capability. A filesystem, a database, an internal api, whatever it is. It publishes a list of the tools it has, each one carrying a name, a description and a schema for its arguments.
A client is whatever is driving the model. It connects to servers and asks each one what it can do.
Asks is the word that matters.
A client written in March can talk to a server published in July, because it never needed to know in advance what that server had. That is the difference between a protocol and an integration, and it is the genuinely new thing here. Build the server once and every compliant client can use it.
The loop, walked slowly
Worth going through a single call slowly.
Your client starts. It connects to a server and asks for the inventory. The server answers with a list: maybe three tools, maybe thirty, each with a name, a sentence or two of description and a schema saying what arguments it expects.
The client hands that list to the model along with whatever the user typed.
The model reads the question, reads the list, decides one applies, and emits a request to call it with arguments filled in. The client forwards that to the server. The server does the work and returns a result, which drops back into the conversation, and the model carries on.
That is all of it.
Notice what is being standardised. The shape of the list. The shape of the request and the response. The model’s part was already happening before anyone wrote a spec, and it works exactly as well or as badly as it did. What changed is that the plumbing around it looks the same everywhere.
The part that actually changed
Integrations became shareable. Before this, a good connector was trapped inside whichever framework you happened to be using, and anybody on a different stack got nothing from your work.
It also moved where extension happens. Adding a capability used to mean changing the application: commit access, a release, a review. Now it is a config entry pointing at another server. The person who knows the system being connected no longer has to know anything about the application consuming it, which is a bigger organisational shift than a technical one.
Where the server actually runs
Two transports, and the difference between them is entirely about trust.
A local server runs as a process on your own machine and talks over stdin and stdout. That is the normal case for anything touching local files or local tools. It also means the server runs with exactly the access your user account has, which on a developer machine is close to everything.
A remote server runs somewhere else and you reach it over the network. Better for a shared service, and the point at which authentication becomes a real design problem.
People underrate the local case constantly. Installing a local server takes about eight seconds and one command. It is the same act as installing any other program written by a stranger and it warrants the same pause. It never feels that way, because it arrives dressed as configuration.
The arithmetic nobody does
Back to the 63 tools.
Every description sits in the prompt. Every call. All of them, used or not, because the model cannot choose from a list it has not been shown.
Run the numbers on a setup slightly bigger than mine. Eighty tools at 50 words each is 4,000 words of overhead before the user has typed a character. You pay it on every request, in money and in the delay before the first token appears. Then ask how many of those 80 were relevant. One, usually. Sometimes none.
People accept this instantly as arithmetic and go on attaching servers anyway, because attaching is one line of json and the cost lands somewhere nobody is looking.
Pruning is the optimisation
The most effective change you can make to a system with too many tools is deleting tools.
Connect what this application needs and stop there. If you are building something that reads a codebase and files issues, it does not need a weather server, and every day it stays attached it makes the issue tool marginally harder to pick.
The advanced version is selecting which tools to expose based on what the incoming request looks like, so the model sees a short relevant list instead of everything you happen to own. That is real engineering and most teams will not do it. Pruning by hand takes an afternoon.
Names are your job
Here is the failure I see most, and it has nothing to do with the protocol.
A model picks a tool by reading its name and its description. That is the entire input to that decision.
So two tools called search and query, both described in one vague line, will get confused with each other for as long as they both exist. The model is not being stupid. You handed it two options that are indistinguishable in the only channel it can see.
Name them for what they do to what. Searching the customer database and searching the product documentation get different names, spelled out, no cleverness.
Then describe the situation the tool is for. “Fetches a record by id” tells the model what happens after the decision. “Use this when the user names a specific order number” tells it how to make the decision, which is the part it is actually stuck on.
This is unglamorous work and it pays back the same day. I have fixed more bad tool selection by rewriting a 40 word description than by any model swap or any amount of prompt surgery.
The security part I would not skip
Attaching a server means letting a model decide when to run code that does something real. If that server can write files, send messages or move money, the model’s judgement is sitting directly in front of consequences with nothing in between.
The risk that gets underweighted is what happens when the model reads untrusted input. Point it at a web page, let that page contain a paragraph written to look like an instruction, and it may act on it. It has no reliable way to separate content it was asked to read from instructions it was asked to follow. Both arrive as text in the same channel.
So treat any server that can act on the world with the care you would give an api key. Know what it can do before you attach it. Prefer read-only where read-only does the job. Put a confirmation step in front of anything destructive or expensive, and accept that it makes the product more annoying, because the alternative deletes things on the strength of a sentence it found on a web page.
What I got wrong
I spent the first month treating server count as a feature. More attached meant more capable, and the demos supported that, because a demo is one request against a clean context.
What it produced was a model that picked the wrong file tool roughly one time in six, and I blamed the model for weeks. Swapped it for a bigger one. Same failure rate. Then I detached four servers I had attached because they looked interesting, and the problem largely went away.
I also have not run this at serious scale. Everything above comes from a handful of internal tools and a setup I use daily, so the authentication story in particular is one I have read about more than fought with.
When it earns its place
Several different systems to reach: worth it, easily. You get existing servers for the common ones and write far less glue.
The same capabilities needed across different models or different clients: worth it, because that portability is the entire argument.
Building something other people will use: publishing it as a server is now the obvious way to distribute it. The alternative is asking every user to write your integration themselves.
One integration, to one system, that you already wrote and that works: leave it alone. Standardising a single connection is process for its own sake. The value here scales with how many connections you have, and at one connection it returns nothing.
The framing I would hold onto is that this is plumbing. It stops you rewriting the same connector every time your stack moves, and it does nothing about whether the model calls the right thing at the right moment. That part stayed yours, and it still lives in the descriptions. The comparison tables and current pricing I keep are here.