Protocol
Measure appropriate reliance as two errors, not one trust score
Good reliance means knowing both when to listen and when not to.
When it fits
- A human-AI workflow is evaluated with one question such as 'Do users trust the AI?' or 'How often do they agree?'
When to avoid it
- Some outputs are not objectively binary correct/incorrect; define adjudication and uncertainty carefully.
Why it matters
Measure two behavioral mistakes separately: rejecting correct AI advice and accepting incorrect AI advice. Add task-specific costs because one error may be much worse than the other. Overall agreement can rise while appropriate reliance gets worse.
Steps
- The evaluation can distinguish under-reliance from over-reliance and connect each to real task cost.
An example
In data deletion support, accepting one wrong recommendation may matter more than rejecting several correct suggestions, so agreement rate is a poor primary metric.
Check your result
The evaluation can distinguish under-reliance from over-reliance and connect each to real task cost.
Keep this limit in mind
- Some outputs are not objectively binary correct/incorrect; define adjudication and uncertainty carefully.
Connected ideas
Useful withMeasure AI speed and quality as separate outcomes
Evidence and sources
The same 2026 study found decision performance depended on AI recommendation quality and that adopting poor advice could reduce performance.
Viewing advice without adopting it is behaviorally different from relying on it.
Who listens to ChatGPT and when should they? A two-study examination of AI-assisted decision making · Abstract highlights
Reliance quality should distinguish accepting correct advice from accepting incorrect advice rather than measuring trust or overall agreement alone.
The best metric depends on task costs and whether false acceptance and false rejection have different consequences.
More is not better: Visual uncertainty cues and the fragility of trust calibration in LLM-assisted decision making · Appropriate reliance and behavioral calibration results