Protocol

Run failure-mode analysis before the consequential change

Ask how the process can fail before production teaches you.

When it fits

  • You are about to change a process, migration, automation or workflow where one missed failure mode could have a large blast radius.

When to avoid it

  • FMEA prioritization numbers are not calibrated probabilities. Use them to structure attention, not simulate certainty.

Why it matters

Walk each important step and list plausible failure modes, their effects, existing controls and detectability. Use rough severity/likelihood/detectability ratings only to prioritize attention. Spend the review on weak controls and hard-to-detect high-impact modes, not on debating a precise risk score.

Steps

  1. At least one high-impact or hard-to-detect failure mode has a stronger prevention, detection or recovery control.

An example

Bulk update: wrong selection filter → thousands of unintended records → current control: manual review → detection: after activation → safeguard: pre-run count and sampled IDs.

Check your result

At least one high-impact or hard-to-detect failure mode has a stronger prevention, detection or recovery control.

Keep this limit in mind

  • FMEA prioritization numbers are not calibrated probabilities. Use them to structure attention, not simulate certainty.

Connected ideas

Useful with
Prefer a system-strengthening action over another reminder

Evidence and sources

Supports

FMEA is a prospective method for identifying possible failure modes and considering their effects before or during process improvement.

Priority scores are heuristics; they should not be mistaken for calibrated failure probabilities.

Failure Mode and Effects Analysis · Description and uses

All sources (1)