Skip to content

AI & ML

Evaluating Agents in Production: Trajectories, Judges, and the Harness That Survives a Model Swap

Output-only evaluation misses the failures that matter. An architecture for scoring the path an agent takes, running judges you can afford, and turning production traces into the eval set.

7 min readIdeaxa Engineering

The demo passed. That was never the question.

Every agent project reaches the same week. The demo works, a stakeholder has seen it, and someone asks how you will know if it breaks. The honest answer, on most projects, is "a user will tell us" — and the second honest answer is that by then the agent will have been quietly wrong for eleven days.

Traditional software has a clean answer to this. The function returns the wrong value, the test goes red. Agents do not offer that. The same input produces different trajectories on different days. The output can be word-perfect while the path there called a tool it should not have, retried four times, and cost nine dollars. And the failure that ends up in a postmortem is rarely "the answer was wrong" — it is "the agent did something nobody imagined it could do."

This is an architecture problem, not a prompt problem. Here is the shape of a harness that works.

Why output-only scoring is not enough

Score only the final answer and you are blind to the interesting half of the system.

Consider an agent asked to resolve a billing dispute. Two runs, identical final message: "I've issued a $40 credit and emailed confirmation."

  • Run A looked up the account, checked the refund policy, applied the credit, sent the email. Four calls.
  • Run B looked up the account, called the refund tool, got a timeout, called it again, got a timeout, called it a third time and succeeded, then sent three emails.

Output-only evaluation scores these identically. One of them refunded the customer once and one of them may have refunded them three times. Trajectory evaluation scores the step-by-step path rather than just the destination, which is how you catch duplicate tool calls, irrelevant actions and unsafe intermediate steps.

The rule we work to: if a behaviour would show up in an incident review, it must be scorable. Cost, tool choice, retry count, and irreversible actions all qualify. None of them are visible in the final message.

Four layers, and what each is for

An eval harness that survives contact with production has four distinct layers. Teams get into trouble by collapsing them — usually by trying to make one LLM judge answer every question.

LayerQuestionCostRuns
AssertionsDid it break a rule?~0Every turn
Trajectory metricsWas the path sane?~0Every turn
LLM judgeWas the answer good?HighSampled
Human reviewIs the judge right?HighestWeekly batch

The economics force this shape. A full model call per judgment is too expensive to run on every production turn, so the cheap layers have to carry the volume and the expensive layers have to be aimed.

Layer one: assertions

Deterministic, boolean, and they run on every single turn because they cost nothing. These are not "evals" in the fashionable sense — they are invariants.

ASSERTIONS = [
    ("no_pii_in_output",      lambda t: not pii.scan(t.final_text)),
    ("no_admin_tools",        lambda t: not (t.tools_used & ADMIN_TIER)),
    ("budget_respected",      lambda t: t.cost_usd <= t.budget_usd),
    ("terminated",            lambda t: t.stop_reason != "max_steps"),
    ("cited_when_claiming",   lambda t: not t.made_factual_claim or t.citations),
]

Write these first. On every project we have done, the assertion layer catches more real problems in the first month than the judge does, because it encodes the things you already know are unacceptable. ADMIN_TIER here is the same tool classification described in the MCP gateway article. A max_steps termination is not a quality signal — it is a bug report.

Layer two: trajectory metrics

Also deterministic, also free, computed from the span tree you are already emitting if you instrumented against the OpenTelemetry GenAI conventions.

The five that earn their place:

  • Step count against a per-task-type baseline. The tail is your failure population.
  • Redundant calls — the same tool with the same arguments twice in one trajectory. Almost always a bug, occasionally a legitimate retry, never something you want to be unaware of.
  • Tool precision — of the tools called, how many contributed to the final answer. Requires a little judgement to label, but the trend matters more than the absolute.
  • Recovery rate — when a tool errored, did the trajectory recover or spiral.
  • Cost and latency at p50/p95, per task type rather than globally. A global average hides the 4% of turns that cost thirty times the median.

Layer three: the judge

Now, and only now, an LLM scoring quality. Three things make judges usable rather than decorative.

Score dimensions separately. A single 1–10 "quality" score is noise. We score faithfulness, answer relevance and instruction compliance as separate calls or separate structured fields, because a response can be perfectly grounded and completely fail to answer the question.

Use a rubric with anchors, not adjectives. "Rate helpfulness 1–5" produces a judge that returns 4 for everything. Anchored levels — "3: answers the question but omits a caveat stated in the source" — produce something you can act on.

Control for the known biases. LLM-as-judge carries length, position and self-preference biases and is non-deterministic. Practically: randomise the order in pairwise comparisons, keep the judge model different from the generation model, and re-score a fixed holdout set whenever you change the judge prompt so you can tell judge drift from system drift.

# Sample deliberately, not uniformly. The interesting turns are rare.
def should_judge(trace) -> bool:
    if trace.assertion_failures:      return True   # always
    if trace.step_count > p95(trace.task_type): return True
    if trace.user_signal == "negative": return True
    return random.random() < 0.02                   # baseline drift check

That last line is the one people forget. Without a uniform baseline sample you will only ever see the turns you already suspected, and you will never notice quality sagging in the boring 96%.

Layer four: humans

Not to score everything — to score the judge. A weekly batch of fifty traces, reviewed by someone who knows the domain, compared against what the judge said. If judge–human agreement drops below about 80%, the judge is broken and every number built on it is fiction.

This is also where the eval set comes from, which brings us to the part that actually compounds.

The loop: production traces become the eval set

A static eval set built at project start decays immediately, because it encodes the failures you could imagine in week one. The systems that hold up treat reviewed traces as feeding directly back into evaluation sets.

production ──▶ traces ──▶ sampler ──▶ judge ──▶ human review
                                                     │
                          eval set ◀── promote ──────┘
                             │
                             ▼
                   CI: run on every prompt,
                   model or tool change

The mechanics that matter:

  • Promotion is a pull request. A trace becomes a test case with an expected outcome, reviewed like code. Not an automatic pipeline — automatic promotion is how a wrong answer becomes an enshrined wrong answer.
  • Freeze the inputs, not the outputs. Store the user turn and the tool responses. Assert on properties (cited its source, stayed under budget, called the refund tool exactly once), not on exact strings. Exact-match assertions on generated text break on every model release and teach the team to ignore red builds.
  • Every incident produces a case. Non-negotiable. This is the ratchet — the thing that makes the harness stronger every week rather than a static endpoint.

Two hundred well-chosen cases beat two thousand generated ones. We hand-curate to roughly that number and let it grow only through the promotion path.

Running it in CI without waiting twenty minutes

The full set against a real model is slow and costs money, so tier it.

TriggerSetModelTarget
Every commitAssertions + cached-response replaynone< 60s
Every PR~60-case smoke setproduction model< 5 min
NightlyFull set + judgeproduction modelany
Model changeFull set, both models, pairedbothany

The replay tier is the one worth building early. Record tool responses from real traces and replay them against the current prompt and orchestration code. It runs in seconds, costs nothing, and catches the large class of bugs that live in your orchestration rather than in the model.

That last row is the reason the harness exists. When a provider ships a new model version, the question "is this safe to adopt" is otherwise unanswerable, and teams end up either pinning forever on an ageing model or upgrading on vibes. A paired run over two hundred cases turns it into a table you can read in ten minutes.

What we would skip on a first build

Being honest about scope, because eval infrastructure is easy to gold-plate:

  • Skip a bespoke eval platform. The layer that matters is your trace schema. Judges and dashboards are replaceable; a badly-shaped span tree is not.
  • Skip synthetic data generation until the curated set exists. Generated cases cluster around what the generator finds plausible, which is exactly the region you already handle.
  • Skip fine-tuning a judge until judge–human agreement has been measured for a month and found wanting. Usually the rubric is the problem.
  • Skip per-turn judging entirely at low volume. Under a few thousand turns a day, weekly human review of a sample beats an automated judge nobody has calibrated.

The bar

An agent is production-ready when you can answer four questions without hedging:

  1. What fraction of turns violated an invariant last week?
  2. Which task types have a step-count tail, and what is in it?
  3. If your provider deprecated the current model tomorrow, how long to qualify a replacement?
  4. When a user says "it did something weird on Tuesday" — can you find the trajectory, replay it, and add it to the set?

Most teams shipping agents today cannot answer the fourth. It is the cheapest of the four to fix and the one that compounds, because it is the mechanism by which the system gets less wrong over time instead of more.

Everything else is a demo with good uptime.

Services in this article