An AI agent isn't graded on whether its final message sounds right — it's graded on whether the task actually got done, using the right tools, in a reasonable number of steps, recovering cleanly if something broke along the way. Task completion, not response quality, is the unit of measurement.

Quick Answer: Agent evals score the whole trajectory — tool selection, argument correctness, error recovery, and final task state — not a single response. A "good answer" mid-trajectory means nothing if the task never actually finished.

Why Single-Turn Evals Break Down for Agents

A single-turn eval compares one output against one rubric: is this answer accurate, is it well-written, does it match a reference. An agent doesn't produce one output — it produces a sequence of decisions (call a tool, read the result, decide what to do next) before anything resembling a final answer appears, and grading only that last message ignores every point where the sequence could have derailed.

This distinction has a name in the research literature. The ReAct framework (Yao et al., Princeton and Google Research, 2022) formalized agent behavior as an interleaved loop of thought, action, and observation — the model reasons, picks a tool, reads what comes back, and reasons again. Evaluating an agent means evaluating that whole loop, not just its last utterance.

From Response Quality to Trajectory Quality

Consider what actually changes when you move from a chatbot to an agent:

  • Multi-step execution replaces a single forward pass — a task might take 3 steps or 30.
  • Tool-use correctness becomes gradable: right tool, right arguments, right order.
  • State matters — the world (a database, a calendar, a customer record) changes as the agent acts, and those changes can't be undone by a better final sentence.
  • Recovery from failure is now part of the skill being measured, not an edge case.
  • End-state success, not phrasing quality, is the pass/fail bar.

If you're building an evaluation program from scratch, the foundational framing in the complete guide to AI evals still applies — you still need a golden set, a rubric, and a scoring loop. What changes for agents is what you're scoring: a trajectory, not a turn.

Process Metrics vs. Outcome Metrics: Two Different Questions

Agent evals split into two families that answer different questions: process metrics ask "did it do the right things," and outcome metrics ask "did the task get done." You need both, because an agent can pass one while failing the other — perfect tool use with a wrong final state, or a completed task reached through a broken, unsafe, or lucky process.

Think of it as the agent-evaluation analog of unit tests versus an end-to-end test: process metrics check each step in isolation, outcome metrics check whether the whole system actually worked.

Metric typeWhat it measuresExampleFailure mode it catches
Tool-selection accuracy (process)Did the agent call the correct tool for the sub-taskCalling check_inventory before charge_cardRight instinct, wrong or premature action
Argument correctness (process)Were the parameters passed to the tool valid and well-formedPassing a real customer_id instead of a hallucinated oneMalformed calls that fail silently or corrupt state
Step-ordering / dependency respect (process)Did the agent respect real-world preconditionsVerifying refund eligibility before issuing a refundSkipped guardrails that produce a "successful" but wrong outcome
Task completion rate (outcome)Did the environment reach the defined success end-stateSubscription is actually canceled in the system of recordAgent stops early, or reports success without acting
Task correctness (outcome)Was the completed task the right oneCorrect flight booked, not merely a flight bookedTechnically-done, actually-wrong outcomes
Recovery success (process)Did the agent handle a failed step without derailingRetrying with corrected input after a validation errorSilent failure or infinite retry on the same bad call

A golden dataset for agents needs to encode both sides — expected tool calls at each step, not just an expected final answer. The same discipline covered in building your first golden dataset applies here, except each of your 50 examples is a full expected trajectory rather than a single question-answer pair.

The Over-Indexing Trap: Process-Perfect, Outcome-Wrong

Teams that come from a traditional QA background tend to over-invest in process metrics, because they're the ones that feel most like familiar unit tests — deterministic, checkable, easy to automate. That instinct misses the failure mode that actually costs money: an agent that calls every tool correctly, in the right order, with valid arguments, and still lands on the wrong final state because a business rule got misapplied along the way.

Treat process metrics as diagnostic, telling you where a trajectory went wrong, and outcome metrics as the gate, telling you whether you can ship. A trajectory that fails outcome but passes every process check is your highest-priority debugging target, not a passing eval.

Scoring a Trajectory: A Worked Example

Scoring a trajectory means comparing the agent's actual step sequence against a golden trajectory (or an accepted set of valid trajectories, since more than one correct path often exists) and scoring both individual steps and the final state. Below is a stripped-down example for a support agent asked to cancel a subscription and confirm refund eligibility.

StepExpected actionAgent's actual actionProcess scoreNote
1get_customer_by_emailget_customer_by_emailPassCorrect tool, correct argument
2get_subscription_statusget_subscription_statusPass—
3Evaluate refund eligibility against policyEvaluated correctlyPassReasoning step, graded by rubric
4cancel_subscriptioncancel_subscription called twicePartialRedundant call — inefficiency flag
5send_confirmation_emailSkippedFailTask technically "done" but incomplete per spec

Even though the subscription was canceled — a plausible outcome pass — the missing confirmation email and the redundant cancellation call mean this trajectory should fail a strict eval. A rubric that only checked "is the subscription canceled" would have missed both problems.

A workable trajectory-scoring rubric, in order:

  1. Define the golden trajectory (or an accepted set of valid trajectories) for each task in your eval set.
  2. Score each step for tool-selection and argument correctness against that reference.
  3. Score the end-state against ground truth: binary completion, plus correctness of the completed state.
  4. Score recovery behavior on any step where a tool call failed or returned unexpected data.
  5. Aggregate into a per-episode score, then compute a pass rate across the whole eval set.

Whatever aggregation method you pick, the episode needs a hard numeric bar for pass/fail, not a vague sense of "mostly fine" — the same discipline behind setting numeric thresholds for failable evals applies just as directly to a five-step trajectory as it does to a single graded response.

Reasoning steps (like step 3 above) are harder to grade mechanically than tool calls, since there's no single correct string to match. Many teams use an LLM to judge whether the reasoning was sound — but that judge needs its own calibration first, which is exactly the problem covered in when to trust an LLM-as-judge.

Error Recovery: The Skill Static Evals Never Had to Measure

A single-turn eval has no concept of recovery — a wrong answer is just wrong, full stop. An agent operates inside an environment where tools time out, APIs rate-limit, and results come back empty or malformed, so recovery from a mid-trajectory failure is itself a scored skill, not an edge case to shrug off.

Recovery behavior generally falls into a few recognizable patterns:

  • Retry with correction — the agent detects a validation error and retries with fixed arguments, rather than repeating the identical failed call.
  • Fallback path — the agent detects an empty or ambiguous result and switches to an alternate tool or asks a clarifying question.
  • Clean exit — the agent detects it's genuinely stuck and reports an accurate failure status instead of fabricating a success or hanging indefinitely.
  • Backtrack — the agent notices a precondition wasn't actually met and reverses or re-does an earlier step before continuing.

A model that quietly hallucinates a successful tool result instead of surfacing an error is more dangerous than one that fails visibly — it passes the eval and fails the customer.

AgentBench (Liu et al., Tsinghua University, 2023), one of the earlier standardized benchmarks for evaluating LLMs as agents across environments like operating systems, databases, and web shopping, found a pattern worth internalizing: models with strong general reasoning still stumbled specifically on tool-interaction fidelity — malformed calls, misread outputs, poor recovery. Language competence and agentic competence turned out to be measurably different skills.

A graceful-failure rate — the share of unrecoverable-fault tasks where the agent exits with an accurate status instead of hanging or fabricating success — is worth tracking as its own number, separate from raw completion rate.

Cost, Loops, and the Economics of Letting an Agent Run

Agents can burn unbounded tool calls and tokens if nothing is watching, so a task that "completes" only after 40 redundant retries can pass a naive outcome metric while being commercially unviable and masking a reasoning defect. Cost and loop detection are eval dimensions that single-turn evals simply never needed, because a single-turn call has no way to loop.

Three things worth instrumenting, specifically:

DimensionWhat it flagsTypical guardrail
Tool calls per completed episodeInefficiency or thrashingStep budget + alert on P90/P95 outliers
Repeated identical calls in a rolling windowAn agent stuck in a cycleCycle detector that flags 2+ identical consecutive calls
Cost per successful completion (not per call)Economic viability at scaleDollar ceiling per task, not just per token
pass^k across repeated trials of the same taskReliability vs. one lucky runRequire passing at k ≥ 3, not k = 1

That last metric comes from τ-bench (Sierra AI, 2024), a benchmark built specifically around tool-agent-user interaction, which popularized measuring whether an agent succeeds consistently across repeated trials of an identical task rather than crediting a single successful attempt. A trajectory that only completes once out of five identical attempts is a different reliability profile than one that completes five out of five — a single pass/fail run tells you nothing about which one you're looking at.

The stakes for getting this right are rising with model capability. METR's research on task time-horizons has found that the length of task AI agents can reliably complete has been roughly doubling every seven months in recent years — which means agents are increasingly being trusted to run longer, with less supervision, before anyone checks in. Cost ceilings and loop detection stop being a nice-to-have exactly when that autonomy window gets long enough for a stuck agent to burn real budget before a human notices.

Setting a Step Budget From Your Own Golden Trajectories

A step budget works best when it's derived from data you already have rather than picked arbitrarily. Look at the step count across your golden trajectories — the ones a competent human or a well-behaved agent actually needed to finish the task — and set the budget at a multiple of that typical length, generous enough to allow legitimate retries but tight enough to catch runaway loops early.

Anthropic's engineering guidance on agent design (Building Effective Agents, Anthropic, 2024) makes a related point about the broader pattern: the simplest workflow that reliably completes the task usually beats a more autonomous, open-ended agent loop, precisely because a bounded workflow is easier to evaluate, budget, and reason about failure in. Treat "does this really need full agentic autonomy" as an eval-design question, not just an architecture one — a shorter, more constrained trajectory is inherently cheaper to score and cheaper to run.

Where Prodinja Fits: Planning Success Before You Ship

Before any of the scoring above is possible, someone has to decide what "done" actually means for a specific multi-step feature — which tool calls are required, what end-state counts as success, and which recovery paths are acceptable to ship with. That definitional work is easy to skip and expensive to skip badly.

Prodinja's eval-planning flow, part of its agentic-critique tooling, walks you through that definition before you write a single test case: it's designed to help you specify success as a task-completion threshold across a trajectory — the tools that must get called, the end-state that counts as done, and the failure modes you're willing to tolerate — rather than a single pass/fail score on one response. It's a planning aid for framing the eval, not a system that runs the evaluation and hands you a verdict.

Framing your task the way you'd frame a customer's underlying job is a useful discipline here too. The Jobs-to-Be-Done lens is built around exactly this: the customer doesn't hire an agent to produce a plausible-sounding step, they hire it to get the whole job done end-to-end — and your eval should hold the agent to the same bar.

Mapping out the full trajectory the way you'd map a customer journey — stage by stage, with the moments where things typically go wrong — is often the fastest way to spot the steps a naive eval would otherwise skip.

Key Takeaways

  • Grade the trajectory, not the last message. An agent's final answer can look great while every step that produced it was wrong, redundant, or unsafe.
  • Process metrics and outcome metrics catch different failures. Tool-selection and argument correctness catch a broken process; task completion and correctness catch a wrong destination — you need both, not one standing in for the other.
  • Recovery from failure is a first-class metric now. A graceful, honest failure beats a silent hallucinated success — track a graceful-failure rate separately from raw completion rate.
  • Cost and loop detection are agent-specific eval dimensions. Watch tool-call counts, repeated-call cycles, and cost per successful completion, not just per-call accuracy.
  • Consistency matters as much as capability. A pass^k view across repeated identical trials, per τ-bench, reveals reliability that a single successful run hides.
  • Define "done" before you build the eval. A task-completion threshold — required tool calls, valid end-states, acceptable recovery paths — has to exist before any trajectory can be scored against it.

Frequently Asked Questions

What's the difference between an LLM eval and an agent eval?

An LLM eval scores a single input-output pair — was this one response accurate or well-formed. An agent eval scores an entire trajectory of tool calls, intermediate reasoning, and state changes across multiple steps, ending in whether the real-world task was actually completed.

How do you measure task completion for an AI agent?

Task completion is typically measured as a binary against a defined success end-state in the environment — did the database record change, did the email send, did the booking exist — combined with a correctness check that the completed state was the right one, not merely a completed one.

What is trajectory scoring in agent evaluation?

Trajectory scoring compares an agent's actual sequence of tool calls and reasoning steps against a golden trajectory (or set of accepted valid trajectories), grading tool selection, argument correctness, and recovery behavior at each step, then combining that with the final outcome into one episode-level score.

How many tool calls should an agent be allowed before it's flagged as stuck in a loop?

There's no universal number — it depends on the task's realistic step count — but a common pattern is to set a step budget from your golden trajectories' typical length, flag P90/P95 outliers, and separately flag any window where the same call repeats two or more times consecutively, since that's a stronger loop signal than raw call count alone.

Can you use LLM-as-judge to evaluate agent reasoning steps?

Yes, for the reasoning and tool-argument-quality steps that don't have a single matchable string, an LLM judge is often the practical choice — but it needs to be calibrated against human-graded examples first, since an uncalibrated judge can rubber-stamp confident-sounding but wrong intermediate reasoning.