Protocol
Re-benchmark AI when the model or task changes
An AI capability map expires faster than most process documentation.
When it fits
- A workflow has an old judgment about AI capability but the model, tools, prompt, data or task shape has materially changed.
When to avoid it
- A small benchmark can miss rare failures; keep production monitoring and consequential review even after a strong re-test.
Why it matters
Keep a small representative benchmark for important AI-assisted tasks and rerun it after material model or workflow changes. Compare quality, failure modes and review burden with the previous version. Promote or reduce AI responsibility from observed results rather than release notes or reputation.
Steps
- Representative task cases are saved with pass criteria.
- The model/tool/configuration version is recorded.
- Material workflow changes trigger re-evaluation.
- New failure modes are compared with old ones.
- Responsibility changes only after the new evidence is reviewed.
An example
After moving from one coding agent to another, rerun repository-specific tasks instead of assuming a benchmark leaderboard predicts your workflow.
Check your result
A material AI-system change produces an updated local capability judgment backed by comparable cases.
Keep this limit in mind
- A small benchmark can miss rare failures; keep production monitoring and consequential review even after a strong re-test.
Connected ideas
Useful withKeep an eval set that can embarrass the agent
Evidence and sources
The jagged-frontier study notes that AI capabilities and their useful task boundary change rapidly, so knowledge workers cannot rely on a fixed map of what AI can do well.
The study itself does not measure every subsequent model generation; reassessment is a practical implication, not a measured effect size.
Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality · Discussion of expanding and changing frontier