Agents fail more often because of bad tool ergonomics than bad reasoning. A model can plan a correct sequence of actions and still fail if the tool it calls returns an ambiguous error, does two jobs at once, or hands back a payload the model can't parse into its next step.
Quick Answer: Agent reliability is bounded by tool design, not model intelligence. Give each tool one clear job, self-describing errors, and idempotent, predictable return shapes — the same discipline you'd apply designing an API for a junior engineer's first week.
Most teams shipping AI agents in 2026 still treat tool definitions as an afterthought — a thin wrapper around whatever internal function already existed. That's backwards. The tool layer is the interface between a probabilistic reasoner and your deterministic systems, and interfaces are where reliability is won or lost.
Why Tool Design Outweighs Model Choice
Tool design outweighs model choice because the model only ever sees what the tool exposes — a bad interface caps performance regardless of reasoning quality, while a good one lets even a smaller model succeed. The bottleneck moves from "can it think" to "can it act correctly on what it's given."
Anthropic's own guidance on building effective agents (its 2024 engineering writeup on agent design) makes a point worth internalizing: most production agent failures trace back to tool definitions, not the underlying model's chain of reasoning. The tool is the model's UX. Just as a confusing button placement causes human users to click the wrong thing, a confusing tool signature causes an agent to call the wrong function, pass the wrong argument, or misinterpret a response.
Researchers studying tool-augmented language models (including the Berkeley Function-Calling Leaderboard team) have repeatedly found that models with strong general reasoning still post materially lower accuracy on tool-calling benchmarks when tool schemas are ambiguous or overloaded — the gap isn't closed by a bigger model, it's closed by a cleaner schema. This mirrors decades of API design wisdom: Joel Spolsky's "Law of Leaky Abstractions" and the API-design literature that grew up around REST and gRPC all converge on the same idea — a good interface hides complexity without hiding information the caller actually needs.
The Junior-Hire Mental Model
Treat every tool as an instruction you're handing to a smart but context-free junior hire on their first day. They don't know your internal conventions. They will follow the tool's name, description, and parameters literally. If the instructions are ambiguous, they'll guess — and guess wrong in the same ways an LLM does: plausibly, confidently, and without asking for clarification unless you've built that in.
This isn't just a metaphor. It's a design constraint you can act on directly, borrowed from the same discipline product teams already use to write clear requirements for customer jobs to be done — the job the tool exists to do has to be legible to the person (or model) doing it.
Tool Granularity: One Clear Job Per Tool
The right granularity for a tool is the smallest unit of work that still produces a meaningful, independently useful result — not a single API endpoint mirrored one-to-one, and not a mega-tool that bundles unrelated actions behind one name. Granularity errors are the single most common root cause of agent tool-use failures.
Too coarse looks like a manage_customer tool that accepts an action parameter (create, update, delete, search) and a grab-bag of optional fields depending on which action was chosen. The model has to correctly infer which fields are relevant for which action, and errors compound because the tool's contract shifts based on a hidden branch inside it.
Too fine looks like splitting get_user_email, get_user_phone, and get_user_address into three separate tools when they're always fetched together. This forces the model into three round-trips where one get_user_contact_info call would do, burning context and latency without adding clarity.
A Practical Granularity Test
Ask three questions before shipping a tool definition:
- Does this tool do one verb, on one noun, with one clear success condition? If you need "and" to describe it, split it.
- Would a human API consumer need to read the source to know what arguments are valid together? If yes, the schema is hiding conditional logic the model can't see either.
- Does the tool's name alone predict its behavior?
search_ordersshould search orders — not silently create one if none are found.
This is the same discipline behind well-scoped REST resources and the Unix philosophy of "do one thing well" — decades-old advice that turns out to matter even more when the caller is a language model instead of a human developer reading documentation.
Naming and Descriptions Are Load-Bearing, Not Cosmetic
Tool names and descriptions function as the model's only documentation at inference time — there's no onboarding, no tribal knowledge, no Slack thread to ask in. A vague name like process_data forces the model to infer intent from context, which is exactly where hallucinated arguments and wrong tool selection come from.
Name tools as verb_noun pairs that match the mental model of the task, not your internal service names. cancel_subscription beats update_billing_state, even if they hit the same backend code path. The model reasons in terms of user intent, not your database schema.
| Naming pattern | Example | Why it fails or works |
|---|---|---|
| Internal jargon | mutateEntityV2 | Meaningless to the model; invites wrong usage |
| Overloaded verb | handle_request | No signal about what actually happens |
| Vague noun | get_data | Ambiguous scope — data about what? |
| Clear verb_noun | refund_order | Predicts behavior and side effects precisely |
| Scoped and specific | search_orders_by_customer_email | Parameters are implied by the name itself |
Descriptions should state, in plain language: what the tool does, when to use it, when not to use it, and what it returns. Anthropic's tool-use documentation is explicit that description quality correlates directly with call accuracy — the description is effectively a few-shot prompt the model reads before every decision to invoke that tool.
Parameter Design
Keep parameter names self-explanatory and types narrow. An enum for status beats a free-text string the model might populate with a synonym your backend doesn't recognize. Required versus optional parameters should map to what's actually required by the operation — not what happens to have a default in your codebase, which the model has no way of knowing about.
Error Messages: Design for Recovery, Not Just Logging
A good tool error tells the agent what went wrong, why, and what to try instead — a bad one just says the call failed. Since the model has no stack trace to read and no colleague to ask, the error message is the entire debugging session, and it either lets the agent recover in the same turn or sends it into a retry loop or a hallucinated workaround.
Compare the two error philosophies directly:
| Poorly designed error | Well-designed error | |
|---|---|---|
| Message | Error 500: Internal Server Error | Cannot cancel order ord_8821: order already shipped. Use initiate_return instead. |
| What the model learns | Nothing actionable | The failure reason and a concrete next tool to call |
| Likely agent behavior | Retries blindly or gives up | Calls initiate_return in the next turn |
| Underlying cause exposed? | No | Yes, at the right level of abstraction |
This connects directly to the reliability work already established for agent action guardrails: a guardrail that blocks an unsafe action is only useful if its rejection message tells the agent what a safe action would look like. A guardrail that just says "blocked" produces the same blind-retry failure mode as a bad HTTP 500.
Principles for Self-Describing Errors
- State the specific cause, not a generic category. "Invalid date format" is better than "Bad request," but "Expected YYYY-MM-DD, got '07/10/2026'" is better still.
- Suggest the corrective action when one exists, ideally naming the exact tool or parameter to fix.
- Distinguish retryable from non-retryable failures. A rate limit should say "retry after 30s"; a permissions error should say "not retryable — requires admin role," so the model doesn't waste a turn retrying something that will never succeed.
- Never leak raw stack traces or internal identifiers the model can't act on — that's noise competing with signal in a context window that's already finite, a constraint covered in depth by any serious treatment of context engineering for agents.
Idempotency and Return Shape: Making Tools Safe to Retry
An idempotent tool produces the same end state no matter how many times it's called with the same input, which matters because agents retry — sometimes because a network call timed out, sometimes because the model itself decides to try again after an ambiguous result. Without idempotency, retries create duplicate side effects: double charges, duplicate tickets, repeated emails.
The practical fix is an idempotency key — a caller-supplied unique identifier (an order ID, a request UUID) that the backend checks before executing a mutating action a second time. This is standard practice in payment APIs like Stripe's, and it applies just as directly to agent tools: create_invoice(customer_id, idempotency_key) should return the existing invoice on a repeated call with the same key, not create a second one.
Return shape matters just as much as the call itself. A tool that returns a single flat success/failure boolean gives the model nothing to reason about next. A tool that returns a structured object — status, relevant IDs, a human-readable summary, and next-step hints — lets the model chain actions without a follow-up call just to find out what happened.
| Return design | Consequence for the agent |
|---|---|
true / false only | Model can't explain outcome to the user or chain further actions |
| Raw database row dump | Model burns tokens parsing irrelevant fields |
| Structured summary + IDs + status enum | Model can confirm, chain, or recover in one pass |
| Nothing on success (silent) | Model may assume failure and retry unnecessarily |
This is the same non-determinism concern that shows up in agent versus workflow design decisions: a workflow's deterministic step order tolerates ambiguous returns because a human wrote the next step in advance. An agent, choosing its own next step, needs the return value to carry that decision-making information — because there's no hardcoded next step to fall back on.
Before and After: A Poorly Designed Tool vs. a Well-Designed One
Take a common case: an agent tool for updating a support ticket's status. The before version mirrors an internal API 1:1; the after version is redesigned around what the model actually needs to reason correctly.
Before (poorly designed):
Tool: update_ticket
Params: { id: string, payload: object }
Description: "Updates a ticket."
Returns: { success: boolean }
Errors: "Update failed."
This tool fails the granularity test (payload is an untyped blob), the naming test (description doesn't say what fields payload accepts), and the error test (no cause given). A model calling this has to guess at payload's shape from training data patterns, not from anything the tool told it.
After (well designed):
Tool: resolve_support_ticket
Params: { ticket_id: string, resolution_note: string, idempotency_key: string }
Description: "Marks an open support ticket as resolved and logs the resolution
note visible to the customer. Use only for tickets in 'open' or 'in_progress'
status — call reopen_ticket first if the ticket is already closed."
Returns: { ticket_id, status: "resolved", resolved_at, summary }
Error (if not resolvable): "Cannot resolve ticket tkt_4471: status is 'closed'.
Call reopen_ticket first, then retry resolve_support_ticket."
The second version does one job, names it precisely, tells the model exactly when it applies, returns a structured result, and — critically — tells the agent what to do when the call fails instead of just reporting failure. That single error message alone often prevents an entire dead-end conversational turn.
Where Prodinja Fits Into This Work
Reasoning about tool design in the abstract is hard — it helps to have a concrete surface to work against. Prodinja's API Designing tool turns endpoints you sketch into real curl commands and spec output, which gives you an actual, inspectable interface to evaluate: is this the right granularity for an agent to call? Does the response shape give a model enough to act on? Is the naming legible without reading the implementation?
That's a genuinely useful exercise even before you wire anything up to an actual agent — seeing your API as concrete request/response pairs, rather than an abstract description, is often what surfaces a granularity or naming problem you'd otherwise only discover after an agent starts failing on it in production.
Key Takeaways
- Tool ergonomics, not model capability, is usually the ceiling on agent reliability — a bigger model rarely compensates for an ambiguous tool schema.
- Give each tool exactly one verb, one noun, and one clear success condition — split "and" tools, merge tools that are always called together.
- Name and describe tools for a context-free reader, matching user intent language rather than internal service or database naming.
- Design errors to be recoverable in the same turn: state the cause, suggest the next tool, and flag whether retrying will ever help.
- Make mutating tools idempotent with a caller-supplied key, and return structured, informative payloads — not bare booleans — so the model can chain its next action confidently.
- Treat the tool layer as a UX surface you design deliberately, the same way you'd scope a customer-facing feature around a real customer journey rather than your internal architecture.
Frequently Asked Questions
What is the biggest mistake in designing tools for AI agents?
The most common mistake is overloading one tool with multiple unrelated actions behind a single generic name and an action parameter. This forces the model to infer hidden branching logic it can't see, which is the single largest source of wrong-tool and wrong-argument errors in production agents.
How many tools should an agent have access to?
There's no fixed number — the right count depends on task scope, but most reliable agents work with a curated set of narrowly scoped tools rather than dozens of overlapping ones. Fewer, well-differentiated tools reduce the chance the model confuses two similar options, which matters more as you scale autonomy per the agent autonomy levels framework.
Does function calling design matter more than the underlying model?
For most production failures, yes — function calling design determines whether a capable model can act correctly at all, since the tool schema is the only channel between reasoning and execution. A stronger model with a poorly designed tool still inherits that tool's ambiguity.
How do I make an agent tool idempotent?
Add a caller-supplied idempotency key parameter, and have the backend check for an existing operation with that key before executing a mutating action again. This prevents duplicate side effects like double refunds or repeated notifications when an agent retries after a timeout or ambiguous result.
Should tool error messages be different for agents than for human developers?
Largely no in substance, but yes in framing — agent-facing errors need the corrective action spelled out explicitly, since there's no developer to search documentation or ask a colleague. State the cause, name the fix, and indicate whether retrying is worthwhile, every time.