FMEA: Failure Mode and Effects Analysis Explained
FMEA is a structured method for finding how an asset can fail, ranking each failure by risk, and acting on the biggest risks before they cause downtime.
What is FMEA?
Failure mode and effects analysis (FMEA) is a structured, step-by-step method for identifying every way an asset, component, or process can fail, evaluating the effect of each failure, and prioritizing those failures by risk so the highest-risk ones get addressed first. Rather than waiting for a breakdown to reveal a weak point, a team works through the equipment on paper, asks “how could this fail and what happens when it does,” and scores each answer.
Historically, the method came out of aerospace and defense engineering in the 1950s and 60s and is now standard practice in automotive, manufacturing, medical device, and facility reliability programs. The international standard governing it is IEC 60812:2018, Failure modes and effects analysis (FMEA and FMECA), and the American Society for Quality publishes a widely used FMEA reference as well. In a maintenance context, however, FMEA turns a vague sense that “this line is unreliable” into a ranked list of specific failure modes you can build tasks around. It is proactive by design: the whole point is to act on a failure mode before it happens.
Why maintenance teams use FMEA
Maintenance teams use FMEA so that limited PM and inspection resources go where they actually reduce risk. After all, not every failure mode deserves the same attention — a bearing that fails silently and shuts down a production line is a very different problem from a gauge light that burns out. Consequently, FMEA makes that difference explicit and easy to justify.
Done well, an FMEA therefore delivers several things a maintenance manager can use immediately:
- A prioritized risk list — failure modes ranked by a single number, so you know what to fix first.
- Better preventive tasks — each high-risk failure mode gets a specific detection or prevention action instead of a generic calendar task.
- Institutional knowledge — the analysis captures the reasoning of your most experienced technicians before they retire or move on.
- An audit trail — a documented, revisited analysis that supports reliability and, where relevant, regulatory reviews.
Typically, teams run FMEA after an asset criticality analysis — criticality tells you which assets earn the effort, and FMEA digs into how those assets fail.
The FMEA process, step by step
In practice, an FMEA follows a repeatable sequence that a cross-functional team works through together — typically operators, technicians, and a reliability or maintenance engineer. The steps run as follows:
- Define the scope. Pick the asset, system, or process and the boundary of the analysis (one pump, or the whole pumping skid).
- Break it into functions. List what each component must do.
- Identify failure modes. For each function, describe every way it can fail to perform — seized bearing, blocked line, loss of seal.
- Describe the effects. Capture what each failure does to the asset, the operation, and safety.
- Identify causes. Note the mechanisms behind each failure mode; this is where root cause analysis feeds in.
- Score Severity, Occurrence, and Detection. Rate each failure mode 1–10 on all three scales.
- Calculate RPN and act. Multiply the three scores, rank the results, and assign corrective actions to the top risks.
Importantly, FMEA is not a one-time exercise. After the team completes the actions, it re-scores those failure modes to confirm the risk actually dropped — which is why it lives best inside a system that keeps the history.
How FMEA scoring works: RPN (Severity x Occurrence x Detection)
RPN — the risk priority number — is the score that ranks failure modes, and it is the product of three ratings each set on a 1-to-10 scale: Severity × Occurrence × Detection. As a result, the range runs from 1 to 1,000, and the higher the number, the more urgent the failure mode.
- Severity (S) — how serious the effect is if the failure happens. A 10 is catastrophic (safety hazard, total line stoppage); a 1 is trivial.
- Occurrence (O) — how likely the failure mode is to occur. A 10 is near-certain and frequent; a 1 is remote.
- Detection (D) — how likely you are to catch the failure before it causes the effect. This scale runs backwards: a 10 means you almost certainly will not detect it in time, while a 1 means your team catches it early every time.
Watch the Detection scale. A high Detection score is bad — it means the failure hides until it hurts. Improving detection (adding a sensor, an inspection, or condition monitoring) is often the fastest way to pull a dangerous RPN down without redesigning the asset.
A worked FMEA example
For example, below is a short worked FMEA: a short FMEA worksheet for a centrifugal pump on a critical process line. Notice how two failure modes with the same severity land in very different places once you factor in occurrence and detection.
| Failure Mode | Effect | S | O | D | RPN | Action |
|---|---|---|---|---|---|---|
| Bearing seizure | Pump stops, line down | 9 | 6 | 7 | 378 | Add vibration monitoring; set PM re-lube interval |
| Seal leak | Product loss, slow degradation | 6 | 7 | 4 | 168 | Quarterly seal inspection route |
| Impeller wear | Reduced flow / pressure | 5 | 5 | 5 | 125 | Add flow reading to operator round |
| Coupling misalignment | Vibration, premature wear | 4 | 4 | 3 | 48 | Verify alignment at each rebuild |
Notably, bearing seizure wins the queue not because it is the most likely failure, but because it is severe and hard to catch in time (D = 7). Instead, the action targets that detection gap directly with condition monitoring — a far cheaper fix than tearing the pump down on a schedule that may never match the actual failure.
How FMEA feeds a maintenance and reliability program
Of course, FMEA only helps once its output becomes work — scheduled tasks, inspection routes, and monitoring that actually run. That handoff is where a CMMS turns the analysis into a living program instead of a spreadsheet that ages on a shared drive.
Specifically, each high-RPN action in the worksheet becomes something concrete in eWorkOrders:
- Preventive tasks — the re-lube interval and alignment check become recurring PMs in the preventive maintenance engine, tied to the specific asset.
- Condition-based triggers — vibration monitoring can auto-generate a work order the moment a reading crosses a threshold; eWorkOrders integrates with predictive and condition-monitoring vendors such as AssetWatch, so the alert becomes a work order without a person in the loop.
Closing the FMEA loop with real failure data
- Standardized failure codes — when the failure does occur, technicians log it against consistent failure codes, which builds the occurrence history you need to re-score the FMEA with real data.
- A feedback loop — actual failure and downtime records tell you whether your severity, occurrence, and detection estimates were right, so each FMEA revision gets sharper.
Finally, FMEA sits naturally alongside reliability-centered maintenance: FMEA identifies and ranks the failure modes, and RCM logic decides the most cost-effective strategy — preventive, condition-based, run-to-failure, or redesign — for each one.
Turn your FMEA into scheduled, tracked work
Ultimately, a ranked list of failure modes only prevents downtime once it drives real tasks. eWorkOrders takes your FMEA actions and runs them as PMs, condition-based triggers, and tracked work orders — with the failure history to re-score your analysis over time. We have been at this for over 30 years, hold a 4.9 rating on Capterra and G2, and configure the system to fit your assets in a single 90-minute to two-hour session.
Frequently Asked Questions
What does FMEA stand for?
FMEA stands for failure mode and effects analysis. It is a structured method for identifying how an asset or process can fail, evaluating the effect of each failure, and ranking those failures by risk so your team tackles the highest-risk ones first.
How is RPN calculated in FMEA?
To calculate RPN, the risk priority number, multiply three 1-to-10 ratings together: Severity × Occurrence × Detection. The result ranges from 1 to 1,000, and higher numbers mean higher-priority failure modes. Note that the Detection scale runs backwards — a high score means the failure hides until it hurts.
What is the difference between FMEA and root cause analysis?
FMEA is proactive: it looks forward at how an asset could fail and ranks those risks before they happen. Root cause analysis is reactive: it looks backward at a failure that already occurred to find why. The two connect — RCA findings feed the cause column of an FMEA and help you re-score occurrence.
How often should an FMEA be updated?
Review an FMEA whenever you take corrective action on a failure mode, when a new failure turns up that was not on the list, or when the asset, process, or operating conditions change. Many teams also schedule a periodic review — annually is common — using real failure history from their CMMS to re-score severity, occurrence, and detection.
About the author: Janet Jaquis is Marketing Director at eWorkOrders, where she writes about reliability engineering, maintenance strategy, and CMMS software for maintenance teams.