Principle
Do not assume a rubric reduced noise—measure it
A rubric can look structured while reviewers still interpret every line differently.
When it fits
- A team introduces a checklist or criteria set and declares the judgment process standardized.
When to avoid it
- The IFCN null result is from a clinical expert task; it motivates measurement, not a claim that rubrics generally fail.
Why it matters
Compare reliability before and after the structured criteria on repeated or shared cases. If agreement does not improve, inspect criterion definitions, training examples and how criteria are integrated into the final judgment. Keep the null result: structure that does not change reliability may still aid explanation, but it has not earned a noise-reduction claim.
An example
After adding a six-item architecture-risk checklist, compare whether reviewers actually converge more on the same proposals.
Check your result
The team can distinguish 'we added structure' from 'the structure measurably improved consistency.'
Keep this limit in mind
- The IFCN null result is from a clinical expert task; it motivates measurement, not a claim that rubrics generally fail.
Evidence and sources
The 2025 IFCN study found that explicitly adding six expert criteria did not materially improve inter-rater reliability, performance or overall calibration in the studied expert judgments.
This null result applies to one clinical judgment task and does not imply that structured criteria are generally useless.
Utility of the IFCN criteria for identifying interictal epileptiform discharges by experts: A decision hygiene approach to improve inter-rater reliability · Abstract results