Protocol
Run failure-mode analysis before the consequential change
Ask how the process can fail before production teaches you.
When it fits
- You are about to change a process, migration, automation or workflow where one missed failure mode could have a large blast radius.
When to avoid it
- FMEA prioritization numbers are not calibrated probabilities. Use them to structure attention, not simulate certainty.
Why it matters
Walk each important step and list plausible failure modes, their effects, existing controls and detectability. Use rough severity/likelihood/detectability ratings only to prioritize attention. Spend the review on weak controls and hard-to-detect high-impact modes, not on debating a precise risk score.
Steps
- At least one high-impact or hard-to-detect failure mode has a stronger prevention, detection or recovery control.
An example
Bulk update: wrong selection filter → thousands of unintended records → current control: manual review → detection: after activation → safeguard: pre-run count and sampled IDs.
Check your result
At least one high-impact or hard-to-detect failure mode has a stronger prevention, detection or recovery control.
Keep this limit in mind
- FMEA prioritization numbers are not calibrated probabilities. Use them to structure attention, not simulate certainty.
Connected ideas
Useful withPrefer a system-strengthening action over another reminder
Evidence and sources
FMEA is a prospective method for identifying possible failure modes and considering their effects before or during process improvement.
Priority scores are heuristics; they should not be mistaken for calibrated failure probabilities.
Failure Mode and Effects Analysis · Description and uses