Why your tool calls fail and how to make them reliable
If you have shipped anything using LLM tool calling, you have watched a model confidently call a function with the wrong argument types, invent a parameter that does not exist, or call the right tool at the wrong time. This is not a model intelligence problem you can prompt your way out of. It is a systems problem, and it has systems solutions.
What tool calling actually is under the hood
When you give a model a list of tools, you are not handing it an API it can inspect. You are serializing your tool schemas (name, description, parameter types) into text that gets stuffed into the context window, usually as part of the system prompt or a dedicated tool-definition block depending on the provider’s format. The model has no runtime awareness of your functions. It has only ever seen a JSON-ish description of them, once, at the start of the conversation.
Then, instead of writing normal text, the model is trained to emit a structured object: a tool name and an arguments payload, usually JSON. The inference server intercepts that structured output, and your application code is responsible for actually executing it, catching the result, and feeding it back into the context as a new message before the model continues.
Every failure mode you will hit comes from a break somewhere in that chain: schema serialization, next-token prediction of the arguments, JSON well-formedness, or your own execution and result-injection code. Understanding which link broke changes what you fix.
Failure mode 1: schema drift between what you documented and what you pass
The most common bug I see in code review is not exotic. It is a tool schema that got updated in one place and not another. Maybe you renamed a parameter from user_id to userId in your backend but the tool description string still says user_id. The model has no way to know your code changed. It will keep emitting the old name because that is what it read in the schema at the top of the context.
This is not an LLM reliability problem, it is a source-of-truth problem. If your tool schemas are hand-written strings that live separately from your actual function signatures, they will drift, guaranteed, the same way any duplicated documentation drifts. The fix is boring: generate your tool schemas from your actual code, whether that is Pydantic models, TypeScript interfaces plus a JSON Schema generator, or whatever typed layer you already have. If the schema can only be correct because a human remembered to update two files, it will eventually be wrong.
Failure mode 2: argument hallucination on required fields
Even with a perfect schema, models will sometimes fill in a required argument with a plausible-looking value it invented rather than one from the actual conversation. This shows up a lot with things like dates, IDs, or filenames when the user has not actually provided one. The model has been trained to complete the pattern of “a filled-in tool call” more strongly than it has been trained to notice the value is missing from context.
Tightening the parameter description helps some, but the real fix is validation on your side, not persuasion on the model’s side. Treat every argument coming back from a tool call the way you would treat a value from an untrusted HTTP request: validate type, validate range, validate that referenced IDs actually exist in your system before you execute anything with side effects. If validation fails, do not silently retry with a slightly different prompt and hope. Return the validation error as the tool result, in plain language, and let the model see it and correct itself on the next turn. This is dramatically more reliable than trying to prevent the hallucination upstream, because you are now checking a concrete value against ground truth instead of trying to control probabilistic generation.
Failure mode 3: malformed JSON in the arguments payload
This one is smaller than it used to be. Providers that support native structured outputs or constrained decoding for tool arguments (where the sampler is restricted to only emit tokens that keep the output valid against your JSON Schema) have mostly closed this gap for their own SDKs. But it still bites you in three situations: when you are using an older model or endpoint that does not support constrained decoding, when you are running an open-weights model through an inference server that does not enforce the schema, or when your schema has deeply nested or unusual types (unions, unusual enum values, deeply nested arrays) that push against what the constrained decoder handles well.
If you are in any of those situations, do not assume JSON.parse or json.loads will succeed. Wrap the parse in a try block, and when it fails, feed the raw malformed string back to the model as a tool error with a message like “your last tool call was not valid JSON, here is the parser error, please retry.” Models are good at fixing their own malformed JSON when told specifically what broke. They are much worse at getting it right blind on a second attempt with no feedback, which is what happens if your code just crashes and you retry the same turn from scratch.
Failure mode 4: wrong tool selected among near-duplicates
If you register search_users and search_customers and search_accounts as three separate tools with overlapping descriptions, you are asking the model to make a judgment call your own team would get wrong in code review. The model picks based on the semantic similarity between the user’s request and each tool’s description text, and when three tools describe near-identical operations, small wording differences in the user’s message can flip which one gets called.
The fix is consolidation, not better prompting. Where two tools do genuinely overlapping things, merge them into one tool with a parameter that disambiguates, like a type or scope field. Fewer, clearer tools with non-overlapping descriptions beat many narrow tools every time I have measured it in production traffic. This also has a side benefit: fewer tool definitions means fewer tokens spent on schema text in every single request, which matters once you are paying per-token at scale.
Failure mode 5: silent context loss on long tool chains
Multi-step agentic tasks that call five or six tools in sequence fail more often near the end of the chain than at the start. Part of this is genuine compounding error, each step has some probability of a mistake and they stack. But a real chunk of it is context window pressure: as tool results accumulate in the conversation, older tool calls and their outputs get pushed further from the model’s attention, and some providers or your own context management will summarize or truncate older turns to save tokens.
If you are truncating or summarizing tool history to manage context length, you need to be deliberate about what you preserve. Losing the exact arguments of an earlier tool call, or the exact return value, is often worse than losing an equivalent amount of the system prompt, because the model may need to reference that earlier value in a later step (an ID it looked up three calls ago, a total it computed two calls ago). If you must trim, trim from the middle of the chain and keep the first tool call and the most recent few intact, rather than doing a naive sliding window that drops whatever is oldest regardless of relevance.
What actually moves the reliability number
None of this is about picking a smarter model. It is about tightening the boundary between the model and your system: schemas generated from real code instead of hand-maintained strings, validation on every argument before execution, structured error feedback instead of silent retries, consolidated tool sets instead of near-duplicate tools, and deliberate context management on long chains instead of naive truncation. Every one of these is a normal software engineering discipline you already apply to any other untrusted input source. Tool calling is not magic, it is a text-in, text-out interface with a JSON contract bolted on, and it fails in exactly the ways you would expect a text-in, text-out interface with a JSON contract to fail.
Fix the boundary and the model looks a lot more reliable, because the parts that were actually breaking were on your side of the wire the whole time.
For more breakdowns like this on how AI tools actually work under the hood, head back to the AI Tool Gazette homepage.