Protocol
Keep a human baseline for important AI tasks
Without a baseline, 'better with AI' can mean 'faster than I remember.'
When it fits
- AI output looks impressive but you do not know whether it actually improves the work.
When to avoid it
- Small samples are directional and can be biased by case selection; do not convert them into universal productivity claims.
Why it matters
For a small representative sample, compare AI-assisted performance with the current human or non-AI method using the same acceptance criteria. Measure quality, time and review effort separately. Repeat only often enough to detect meaningful capability changes; this is a calibration tool, not permanent double work.
Steps
- The same task and acceptance criteria are used in both conditions.
- Quality is judged independently from speed.
- Review or correction time is included in the AI-assisted cost.
- Cases include at least one known edge condition.
- The comparison is saved with the model and workflow version.
An example
Compare five real specification summaries produced with and without AI, including correction time and factual misses.
Check your result
You can say what AI changed relative to the current method instead of comparing it with an imagined baseline.
Keep this limit in mind
- Small samples are directional and can be biased by case selection; do not convert them into universal productivity claims.
Evidence and sources
The 2026 Organization Science field experiment found strong AI-assisted gains on a set of tasks within the tested capability frontier but lower correctness on a selected complex task outside that frontier.
The experiment used a particular model generation and consulting task set; current frontier location must be re-estimated for today's model and workflow.
Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality · Abstract and results
Across the recent field and lab studies, AI's effect on speed and quality varies by task and workflow, so a productivity claim should measure these outcomes separately rather than assume faster means better.
This synthesis draws a practical measurement implication from heterogeneous studies; it is not a pooled meta-analysis.
Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality · Abstract and task-contingent results