Principle
Use verbal uncertainty as a brake, not a truth signal
Hedging can change behavior without being a calibrated probability.
When it fits
- The model says 'I'm not sure' and users treat the phrase as proof the system is well calibrated.
When to avoid it
- Natural-language uncertainty can be useful even when not numerically calibrated; just do not overclaim what it represents.
Why it matters
Treat uncertainty wording as an interface intervention. It may appropriately slow acceptance, but the phrase itself does not prove the model is uncertain for the right cases. Test whether hedging appears preferentially on errors and whether it improves decisions without causing excessive rejection of correct advice.
An example
'I may be wrong' can prompt a source check, but it should not be counted as calibrated metacognition without validation.
Check your result
The system distinguishes behavioral effect of hedging from evidence that the uncertainty statement itself is accurate.
Keep this limit in mind
- Natural-language uncertainty can be useful even when not numerically calibrated; just do not overclaim what it represents.
Connected ideas
Useful withTest uncertainty cues on behavior, not only user ratings
Evidence and sources
Supports
A large preregistered experiment found natural-language uncertainty expressions reduced agreement with an LLM and increased user accuracy in the tested question-answering setting.
Uncertainty wording can also reduce useful reliance on correct answers; calibration matters.