Talk to us
← All insights

Quality Engineering

Judging Agents When the Rubric Runs Out

Rubrics with right answers miss judgment work. Score coherence, trajectory, judge effects, and scaffolding to see what your eval suite leaves out.

You own an eval suite. It runs on every merge, the dashboard is green, and last week an agent changed a customer-facing number before anyone noticed. Nobody could point to the test that should have caught it. The suite had a column for correct, a column for incorrect, and nothing that described the shape of the answer.

That gap widens as agents take on work with no answer key. Which vendor should we renew. What is this policy worth to us. How much latency are we willing to trade for a tighter data retention stance. You can't score those against ground truth, because there isn't one to score against.

What a rubric certifies

A rubric with a right answer measures agreement with your key. That's cheap and it works, right up to the point where the question moves from retrieval into judgment. Validity Without Ground Truth, a preprint applying stated-preference economics to model evaluation, lays out the concepts an economist would reach for instead: content, construct, and criterion validity, reliability, incentive compatibility, consequentiality. The authors administered a published water-quality valuation survey (Vossler et al. 2023) to six models. Their tests looked like predictions from economic theory rather than answer keys. Demand should slope down. Willingness to pay ought to shift with the scope of the good and with income. Two older models failed the most basic test at a household income level of $75,000. The two newest passed every theoretical validity test the authors could score, then diverged on convergent validity.

The sentence worth pinning above your eval dashboard is theirs: passing validity tests shows a model's answers are coherent, not that they're correct. Coherence is a real, measurable property. Whether it maps onto whatever your customer bought is a separate question, and it's the one your dashboard quietly answers "yes" to.

One way to stop that: break the single pass bit into named dimensions and write, next to each one, the specific assumption it tests. A downward-sloping demand curve certifies internal consistency under a stated framing. It says nothing about the price. If your report collapses that into one number, the person who finds out will be whoever owns the P&L.

Score the trajectory

Post-hallucination reasoning gives you a second axis that has nothing to do with the answer key. PHRBench, a preprint, is a controlled benchmark across four domains and 18 large language models, covering 4,820 controlled instances. It labels each reasoning trajectory independently of final-answer correctness as Hallucination Compliance, Hallucination Avoidance, or Heuristic Correction. Successful recovery stays relatively rare, and it's associated with more frequent belief updates along the trajectory. The authors also report that properties of the hallucinated prompt carry predictive signal for recovery, with a lightweight predictor reaching an AUROC of 0.847.

The operational read is that a model which lands the right answer after swallowing a hallucinated premise is a different asset from one that refused the premise at the top of the run. Your pass column can't separate them. Add a transcript label with three values: the agent accepted the premise, avoided it, or corrected it mid-run. That label still applies when the run ends wrong, which is the case your current rubric drops on the floor.

Your judge is part of the measurement

Project Arena is a vendor-neutral Kubernetes incident benchmark you can deploy yourself: a disposable cluster fixture, injected faults, investigation records saved from any product, scoring by a configurable judge, and a comparison view. It ships a 21-scenario suite and a six-scenario smoke fixture. Its published table compares two platform-native investigation products against an external agent driven through each platform's command line, with detection at 18/21 and 12/21 for the natives and root cause, blast radius, supported final mitigation and implementation readiness scored across all four columns.

The most useful line in that project is the caveat under the second table. It says the same-12-incidents comparison controls which scenarios are included, then names what it doesn't control: "differences in launch prompts, timing or available evidence." That's the discipline most internal eval programs skip. Re-score an old run after changing the judge prompt, the judge model, or the evidence window, and the delta tells you about your instrument, not about the agents.

Know what your scaffolding contributes

Cassis published a benchmark comparing eight LLMs on analytics questions: 30 cases on a public Formula 1 database and 60 business cases on a large business schema. They measured answer quality, reasoning effort, and cost, and ran every model twice, once through their agent's context-navigation layer and once with plain text search. The top business score was 75% correct at $0.12 per question; the most expensive model tested scored 53% at $0.57. Their navigation layer added 10 points to the leader, moving it from 65% to 75%, and added 11 and 15 points to two other models. One model went the other direction, from 58% to 53%, and the authors say they can't explain it yet. Scores across the business set spread from 38% to 75%.

If you compare models through your own agent, you're scoring a pair, and swapping either half moves the number. A vendor's clean test rig tells you little about your stack. The ablation you can run this week is the same scenario set with your navigation layer removed, because that's what tells you whether a new model is worth the switch.

What to do Monday

Pick one eval dimension that currently reports pass or fail. Write down the assumption underneath it, in one sentence, and put that sentence in the report. Add a single trajectory label to transcript capture, with values for accepted, avoided, and corrected. Then put the judge prompt and the judge model name under version control next to the scenarios, re-run the last suite with no other change, and compare. If the scores move, you've found the thing the green dashboard was covering for.

FAQ

Frequently asked questions

What is the difference between a rubric and a validity test?

A rubric with a right answer checks agreement with your key. Validity tests check whether answers are coherent under properties from economic theory, such as downward-sloping demand or willingness to pay shifting with scope and income. Passing those tests shows coherence, and correctness remains a separate question.

What do hallucination labels add to agent evaluation?

PHRBench labels each reasoning trajectory independently of final-answer correctness. The labels are Hallucination Compliance, Hallucination Avoidance, or Heuristic Correction. Successful recovery stays uncommon, and it's linked to more frequent belief updates along the trajectory. A model that lands the right answer after accepting a hallucinated premise differs from one that refused the premise at the top of the run.

Why compare models through my own agent rather than a vendor's test rig?

If you compare models through your own agent, you're scoring a pair. Swapping either half moves the number. A benchmark of eight LLMs on analytics questions found its navigation layer added 10 points to the leader and 11 and 15 points to two other models, while one model dropped from 58% to 53%. A vendor's clean test rig tells you little about your stack.

What should I add to an eval report that only shows pass or fail?

Pick one eval dimension that currently reports pass or fail. Write down the assumption underneath it in one sentence, and put that sentence in the report. A downward-sloping demand curve certifies internal consistency under a stated framing, and it says nothing about the price. Add a transcript label with three values so accepted, avoided, or corrected premises stay visible even when the run ends wrong.