Principle
Put uncertainty around consequential eval differences
A two-point score gap is not a verdict until you know how noisy the estimate is.
When it fits
- Two agent versions have close evaluation scores and the result may change a launch, model-routing or safety decision.
When to avoid it
- Statistical uncertainty does not repair biased sampling, weak labels or an irrelevant metric.
Why it matters
Report sample size and an appropriate uncertainty estimate around important metrics. Pair the aggregate with error counts and failure categories. A small apparent improvement should not drive a consequential release if the evaluation cannot distinguish it from sampling variation or grader noise.
An example
Two support agents score 91% and 93% on a small eval. Before switching, inspect confidence around the difference and whether severe escalation failures changed.
Check your result
The release note reports uncertainty and meaningful error changes, not only a single point estimate.
Keep this limit in mind
- Statistical uncertainty does not repair biased sampling, weak labels or an irrelevant metric.
Evidence and sources
Consequential evaluation metrics should be reported with uncertainty so small observed differences are not treated as certain product improvements.
An uncertainty interval does not repair biased samples, weak labels or the wrong metric; it quantifies uncertainty conditional on the evaluation design.
Build Evals That Actually Matter · 21:09-26:46, treat the judge as a classifier and put uncertainty around consequential numbers