A chatbot hallucination is a wrong sentence someone reads and (hopefully) catches. An agent hallucination is a wrong sentence that becomes a POST request, a refund, or a deleted record. Containing it requires treating every proposed action as an unverified claim: validate its schema, check its parameters against ground truth, and gate execution behind a commit point the fabrication can't cross.

Quick Answer: Agent hallucination differs from chatbot hallucination because it produces executable output — a tool call, API request, or state change — not just text. Contain it with schema validation, parameter grounding against real system state, and a hard verification gate before any action commits.

Why Action Hallucination Is a Different Category of Risk

Content hallucination produces a false statement a human can fact-check before acting on it. Action hallucination produces a false instruction the system may execute automatically, collapsing the fact-check step that used to sit between error and consequence.

The distinction matters because the failure modes look completely different in practice:

  • Content hallucination: the model states a wrong fact, cites a paper that doesn't exist, or misquotes a policy. Damage is bounded by what a reader believes and repeats.
  • Action hallucination: the model calls a tool that doesn't exist, invents a parameter value, targets the wrong record ID, or chains a real tool with fabricated arguments. Damage is bounded by what the downstream system will accept.

Anthropic's own guidance on building agents (in its published engineering writing on tool use) stresses that model outputs feeding directly into side-effecting code need a different trust posture than outputs feeding into a chat window — the model's confidence and its correctness are not the same signal. Simon Willison, who coined "prompt injection" for LLM systems, has separately made the related point that any system letting an LLM take actions on external data needs to assume the model will sometimes be fooled or simply wrong, and design containment around that assumption rather than around the hope it won't happen.

The Risk Curves With Capability, Not With Model Quality

A better model does not make action hallucination go away — it makes fabricated calls more fluent and harder to eyeball as wrong. This is the central shift technical PMs need to internalize when scoping any agent with write or execute permissions, distinct from the general unpredictability covered in agent vs. workflow non-determinism.

DimensionChatbot / content hallucinationAgent / action hallucination
Failure surfaceA sentence, citation, or summaryA tool call, API request, database write
Who catches it (today)A human reader, sometimesOften nobody, until it executes
Cost of one missEmbarrassment, correctionData loss, wrong refund, broken integration
DetectabilityRequires domain knowledge to spotCan be caught structurally (schema, type, existence checks) — before domain judgment is even needed
Right containment layerEditorial review, citations, disclaimersSchema validation + grounding + approval gate

The last row is the reassuring asymmetry: action hallucination is often easier to catch mechanically than content hallucination, because a tool call has a schema and a fabricated one frequently fails structurally before anyone has to judge whether it's semantically sensible.

What Fabricated Tool Calls Actually Look Like

Fabrication in agent output clusters into a small number of recognizable patterns, each catchable by a specific check. Recognizing the pattern tells you which layer of defense should have caught it.

  1. Nonexistent tool invocation — the model calls refund_customer_v2 when only refund_customer exists, having pattern-matched a plausible-sounding variant from training data or from a similarly-named tool seen earlier in context.
  2. Invented parameters — the model supplies a priority: "urgent" field a schema never defined, or fabricates a customer_tier value the actual record doesn't have.
  3. Plausible-but-wrong IDs — the model references order_10293 because it's structurally valid-looking, not because that ID exists in the current conversation's grounded context.
  4. Silent unit or type drift — a dollar amount the model treats as cents, or a string the schema requires as an enum, passed as free text.
  5. Chained fabrication — a real, valid first tool call whose output the model then misreads or invents a follow-up detail from, corrupting the second call in the chain even though the first was clean.

Example: an agent authorized to adjust subscription tiers receives "downgrade this account" and calls update_subscription(account_id=42, tier="free") — but the account in context was account_id=417. The tool call is schema-valid, type-correct, and completely wrong. This is exactly the shape validation alone will miss, and grounding must catch.

That last example is the crux: schema validation catches malformed calls; grounding catches well-formed lies. You need both, and they sit at different points in the pipeline.

Containment Layer One: Schema Validation

Schema validation rejects a tool call before it reaches any business logic, by checking structural conformance against a strict machine-readable contract rather than judging whether the call is a good idea.

Treat every tool definition as a strict contract, not a suggestion the model is free to interpret loosely:

  • Enforce a closed tool registry. Reject any call to a tool name not in the current allow-list — do not let the runtime fuzzy-match refund_customer_v2 to refund_customer, even generously. A near-miss name is a signal to fail loudly, not to guess helpfully.
  • Use strict JSON Schema (or equivalent) per tool, with additionalProperties: false so invented fields are rejected outright instead of silently ignored or silently accepted.
  • Type and enum enforcement at the boundary — a status field with three valid enum values should reject a fourth in code, not rely on the model having read the prompt carefully.
  • Required-field checks before dispatch — a call missing a mandatory reason_code should never reach the executor layer, regardless of how confident the model's surrounding text sounds.

Schema validation is necessary but explicitly insufficient — it validates shape, not truth. A perfectly-typed call to update_subscription(account_id=42, tier="free") passes every schema check while still being wrong, because 42 is a syntactically valid integer that happens not to be the account in question.

Containment Layer Two: Grounding Against Real System State

Grounding checks a proposed action's content against the actual state of the system it's about to touch, closing the gap schema validation structurally cannot close.

Grounding means resolving every parameter the model proposes against a live, authoritative source before dispatch — not trusting that a value present in the model's context window is a value that still exists or ever existed in the target system.

  1. Existence checks: does account_id=42 actually exist, and does it match the account the user's message referenced by name or session context?
  2. Referential consistency: if the conversation established "Acme Corp, account 417" three turns ago, does the proposed call's ID match that entity, not just an entity?
  3. Range and business-rule checks: is a proposed refund amount within the bounds the order actually supports, not merely a plausible-looking number?
  4. Freshness checks: was the data the model reasoned over fetched this turn, or is it stale context from earlier in a long session that may no longer reflect reality?

Grounding is where customer journey context and jobs-to-be-done framing matter for the PM writing the spec, not just the engineer implementing the check — the acceptable range for "is this a plausible action" depends on what job the user actually hired the agent to do, and an agent with a narrowly-scoped job has a narrower space of plausible-but-wrong actions to defend against.

Containment Layer Three: Verification Before Commit

The commit gate is the last containment layer, and the one that catches whatever schema validation and grounding miss — by requiring a distinct verification step between "agent proposes" and "system executes," rather than treating proposal as equivalent to approval.

This is where autonomy level directly determines blast radius, a distinction the agent autonomy levels framework lays out in more depth: an agent scoped to propose-only cannot cause action-hallucination damage no matter how confidently it fabricates, because nothing executes without a separate approval step.

Verification-Before-Commit Checklist

Use this before shipping any agent with write, financial, or irreversible-action capability:

  • Tool allow-list is closed — no fuzzy name matching, no silent aliasing of near-miss tool names
  • Every tool has a strict schema with additionalProperties: false and enum-constrained fields where applicable
  • Every parameter referencing an entity is re-resolved against live data, not trusted from model context alone
  • Irreversible or high-cost actions require explicit human approval before dispatch, regardless of model confidence language
  • A dry-run or diff preview exists showing exactly what will change, in plain terms, before commit
  • Rate and magnitude limits cap the blast radius of a single fabricated call (e.g., max refund amount, max records touched per action)
  • Every executed action is logged with its full grounding trail — what was checked, against what source, at what time — for post-hoc audit
  • A rollback or compensating action exists for anything that does commit, because no gate is perfect

The Guardrail Layer Ties It Together

None of these checks work in isolation; they're complementary layers of the same containment strategy the agent action guardrails approach formalizes — schema at the door, grounding at the boundary, human approval at the point of no return. Removing any one layer means the other two are carrying weight they weren't designed to carry alone.

How Prodinja's Approve-Before-Act Gate Applies This

Prodinja's product design includes an approve-before-act gate for any agent-proposed action: a fabricated tool call or invented parameter is surfaced for human review before it can execute, rather than being trusted to commit on the model's own confidence. It's a direct, practical instance of the verification-before-commit layer described above — not a claim that hallucination is eliminated, but that it's caught at a designed checkpoint rather than left to chance. For teams building their own action-taking agents, the same pattern — schema validation, grounding, and a human gate before commit — is worth designing in from day one rather than retrofitting after the first bad call executes.

Key Takeaways

  • Action hallucination is categorically different from content hallucination — it produces executable output, so a fabrication becomes a consequence instead of a correction.
  • Risk scales with capability, not with model quality — a more fluent model produces more convincing fabricated calls, not fewer of them.
  • Schema validation catches malformed calls (nonexistent tools, invented fields, wrong types) structurally, before any business judgment is needed.
  • Grounding catches well-formed lies — schema-valid calls referencing the wrong ID, stale data, or an entity that was never actually established in context.
  • A commit gate is the last and most important layer — irreversible or high-cost actions should require explicit approval regardless of how confident the model's language sounds.
  • Logging and rollback are containment, too — assume some fabricated action will eventually get through, and design for detection and reversal, not just prevention.

Frequently Asked Questions

What is the difference between agent hallucination and regular chatbot hallucination?

Regular chatbot hallucination produces a false statement a reader can fact-check before acting on it. Agent hallucination produces a false instruction — a tool call, parameter, or target ID — that a system may execute automatically, removing the human fact-check step that normally sits between an error and its consequence.

How do you catch a fabricated tool call before it executes?

Use two layers together: strict schema validation (closed tool registry, additionalProperties: false, enum enforcement) to reject malformed or nonexistent calls, and grounding checks that re-resolve every parameter against live system state to catch well-formed calls referencing the wrong entity.

Can schema validation alone prevent action hallucination?

No — schema validation only checks that a call is structurally well-formed, not that its content is true. A call like update_subscription(account_id=42, tier="free") can pass every schema check while still targeting the wrong account, which is why grounding and a human approval gate are both still necessary.

Does giving an agent more autonomy increase hallucination risk?

Yes, directly — an agent that can only propose actions has no action-hallucination blast radius, because a human reviews every action before it commits. Risk rises specifically as autonomy moves from propose-only toward auto-execute, independent of how good the underlying model is.

What should a verification-before-commit process include?

At minimum: a closed tool allow-list, strict per-tool schemas, live re-resolution of every entity parameter, human approval for irreversible or high-cost actions, a dry-run preview of the exact change, magnitude limits, full audit logging, and a rollback path for anything that does commit.