Protocol
Record confidence before feedback, then score it against the outcome
Calibration needs two columns: what you believed before the answer and what happened after.
When it fits
- You make repeated forecasts, estimates, diagnoses or answerable judgments and want to improve how much trust you place in your own confidence.
When to avoid it
- Calibration feedback has been studied in forecasting tasks and does not automatically transfer to every domain. A confidence percentage is not a scientific probability unless the task and scoring support that interpretation.
Why it matters
For repeated checkable judgments, capture the answer and confidence before seeing the outcome. When the result arrives, score both correctness and confidence. Review the pairs in batches: look for ranges where you are systematically too sure or too hesitant. Forecasting experiments suggest individualized outcome feedback can reduce overconfidence, but the effect is task-dependent, so calibrate on the class of judgments you actually make.
Steps
- Write the judgment before feedback or outcome information arrives.
- Add a confidence estimate using the same scale each time.
- Record the verified outcome separately.
- Review several pairs together and adjust future confidence where a pattern is visible.
An example
Before checking a production defect, write your leading cause and confidence. After the trace or fix confirms the cause, add the outcome. Ten such pairs teach more than remembering only the dramatic wins.
Check your result
You can show a set of pre-outcome confidence–result pairs and name at least one recurring calibration pattern without rewriting old predictions after the fact.
Keep this limit in mind
- Calibration feedback has been studied in forecasting tasks and does not automatically transfer to every domain. A confidence percentage is not a scientific probability unless the task and scoring support that interpretation.
Evidence and sources
In two forecasting studies, individualized calibration feedback reduced confidence among initially overconfident forecasters; overall calibration improved in the more controlled second experiment.
The result is task-dependent and comes from forecasting experiments, not a general test of metacognitive training across all work.