Principle
Route cases using both human and AI confidence only after both are calibrated
Comparing two uncalibrated confidence numbers creates a precise-looking coin toss.
When it fits
- A team wants a rule such as 'AI decides when AI is more confident; human decides otherwise.'
When to avoid it
- The theoretical complementarity model assumes confidence signals more disciplined than most ad hoc workplace ratings.
Why it matters
Before using relative confidence to assign authority, test whether human confidence and AI confidence each discriminate correctness on the relevant task. If one signal is weak, use task class, evidence or independent review instead. Relative-confidence routing is valuable only when the inputs carry reliable metacognitive information.
An example
Do not let an LLM override an experienced operator merely because it outputs 0.94 and the operator says 80% until those scales are validated.
Check your result
Confidence-based routing outperforms or meaningfully complements simpler task-specific routing on held-out cases.
Keep this limit in mind
- The theoretical complementarity model assumes confidence signals more disciplined than most ad hoc workplace ratings.
Evidence and sources
Metacognitive sensitivity concerns how well confidence distinguishes correct from incorrect decisions, which is different from average confidence or simple calibration.
Formal metacognitive metrics require enough labeled decisions; a single confidence value cannot establish sensitivity.
Modeling the joint impact of human and AI metacognitive sensitivity on human-AI collaboration · Abstract and theoretical model
The 2026 mathematical model shows that human and AI metacognitive sensitivity jointly affect achievable combined accuracy when confidence is used to combine decisions.
The Bayes-optimal assumptions are stronger than ordinary workplace decision support.
Modeling the joint impact of human and AI metacognitive sensitivity on human-AI collaboration · Analytic results