Operator’s Guides — ISO 14971 · Module 6 of 12

Your FMEA Score Is Not Your ISO 14971 Risk

RPN and ISO 14971 risk look like the same number. They are not. Here's why the difference matters, and how to score every hazard on a device honestly.

All ISO 14971 modules
Module 06 31:47 video + article

Watch this module

Module 6 of the video course. The article below covers the same ground in written form, so you can watch, read, or both.

# Your FMEA Score Is Not Your ISO 14971 Risk

Someone finished running an FMEA on their device. Severity, occurrence, detection, all scored, every row filled in, a risk priority number calculated for each one and ready to drop straight into the risk management file. Then they went looking at ISO 14971's own definition of risk, and they left a comment on a training video asking a question nobody had answered: why aren't we considering the detection measures in risk evaluation?

That is not a small mix-up. It is the difference between two entirely different numbers ending up in your risk management file, one of them wrong for the standard you are trying to satisfy, sitting there looking exactly as official as the right one would have. And the person asking is not alone. Failure modes and effects analysis is the technique almost everyone learns first, so almost everyone arrives at this exact fork: an RPN in one hand, an empty risk column in the other, and a guess about which number goes where.

I'm Dave Saunders. I've spent more than 30 years commercializing technology, almost 20 of them in medical devices and regulatory affairs, including several surgical robots. This is Module 6 of The Operator's Guide to ISO 14971, a twelve-module series building a complete risk management file from scratch using one running device: WearPump, a wearable subcutaneous insulin pump for Type 1 diabetes. Module 5 identified eleven hazards. This module gives every one of them a number.

Two ingredients, not three

Start with the definition the standard itself uses. ISO 14971 defines risk as the combination of the probability of occurrence of harm and the severity of that harm. Two ingredients: probability and severity. That is the entire recipe.

You already know how the probability piece is built. From Module 2, and again in Module 5: P1 is the probability the hazardous situation occurs, P2 is the probability that situation actually turns into harm, and P1 times P2 gives you the probability of occurrence of harm. Now put that next to what FMEA gives you. A risk priority number is severity times occurrence times detection. Three factors, not two. Occurrence is really just another way of estimating probability, but detection, how likely you are to catch the failure before it reaches anyone, has no equivalent anywhere in ISO 14971's definition of risk.

So when you calculate an RPN and drop that same number straight into your hazard traceability matrix, you are smuggling a detection score into a place the standard never asked for one. The numbers can look similar. They are not measuring the same thing, and ranking your hazards by one instead of the other can change which ones look urgent enough to fix first.

Here is the actual answer to that comment. Detection tells you how good your ability to catch a problem is. It says nothing about how likely the harm is, or how bad it would be once it happens. Fold detection into the risk number and a hazard you are simply bad at detecting looks artificially dangerous, while one you are excellent at catching looks artificially safe, even when the actual harm, if it occurs, is identical. Picture two hazards with the same severity and the same probability of harm, one with an obvious warning sign and one that fails silently. By RPN, the silent one looks catastrophically worse. By ISO 14971's definition they are correctly identical, because the danger to the patient never changed. None of this makes detection worthless. It belongs in risk control, later, as something you improve to reduce residual risk. Keep it in FMEA. Keep it out of the number you write in the risk column.

Score it as if you built nothing

A foundational rule trips people up constantly: you estimate this risk assuming nothing has been done to reduce it yet. No packaging, no protective circuit, no warning label, no software interlock. You score the device as if none of your engineering existed. That does not mean you assume chaos. The device is still used in its intended environment, by its intended user, the way Module 4 defined them. Only the deliberate safety engineering disappears.

It feels backwards the first time. Why imagine your own safeguards away? Because if you score the risk after your controls are already baked in, you can no longer tell whether those controls are doing anything. You have hidden the exact thing you are trying to measure. This is why the standard wants two estimates for the same hazard: an initial risk, scored with nothing in place, and a residual risk, scored again once your controls exist. Compare the two and you can see what your engineering bought you.

A second mistake sits right next to the first, and I see it almost as often. When a team adds a control, a warning label or an alarm, they often lower the severity score to reflect it. That is usually the wrong axis to move. A warning label does not make the harm any less severe if it occurs. What it does, if it works, is make the harm less likely, because a warned user behaves differently. Severity describes the worst reasonable outcome of the harm itself; a control that changes behavior belongs in probability. Move the wrong number and your residual risk looks better than it really is, which is exactly what an auditor is trained to notice. There is one genuine exception. If a control physically changes what harm is even possible, not just how likely someone is to encounter it, severity can legitimately move. A hard mechanical cap on the maximum dose deliverable per hour does not just make an overdose less likely; it shrinks the worst-case overdose itself. That is a real severity change.

Where the numbers come from

All of this assumes you have a number to plug in for probability. The standard's guidance gives you a hierarchy, in order of how much to trust each source. At the top: published standards, articles, and data on similar devices already in use. Somebody has probably already measured how often the failure you are worried about actually happens, and an auditor will recognize the citation. Next: your own product's field history, if you have one, from complaints, post-market surveillance, and service records. Next: testing, simulation, and analysis you run yourself, including fault tree analysis from Module 5. Last: expert judgment, pulling your team into a room and asking honestly how often this happens.

If your estimate came from that last category, you are not doing anything wrong. For a genuinely new device with no track record, expert judgment is often all you have, and most first-generation devices are estimated this way. The limit is simple: the worse the potential harm, the less comfortable you should be resting on a guess alone. A catastrophic outcome deserves real data or real testing before you write the number down.

A rating should also mean something specific. A probability rating of 3 on a five-point scale is not a feeling; it maps to a frequency band you defined ahead of time. A workable scale might read: 1 is under 1 in 100,000 uses, 2 is 1 in 100,000 to 1 in 10,000, 3 is 1 in 10,000 to 1 in 1,000, 4 is 1 in 1,000 to 1 in 100, and 5 is over 1 in 100. Severity gets the same treatment, described as outcomes a patient's family would recognize: negligible, minor, serious, critical, catastrophic. Five points is a good default. Fewer than three and you cannot tell risks apart; more than about ten and the distinctions get too fine to defend. Whatever you choose, the whole team uses the same scale the same way every time, agrees in advance whether probability means per use or per year, and writes down the reasoning behind every score. A rating with no reason is worthless to the next engineer and to an auditor. A rating with a one-sentence reason and a data source is a defensible estimate. The difference costs you one sentence.

Scoring WearPump, live

WearPump's matrix already had four columns from Module 5: hazard, sequence of events, hazardous situation, harm. Now it gains three more: probability, severity, and the risk rating those two combine into. On the five-by-five grid, the coloring is severity-weighted on purpose. The top-right corner, high severity and high probability together, is deep red. The bottom-left is green. A severity of 5 stays in an elevated band even at a low probability, because a catastrophic outcome, even a rare one, is never allowed to hide in green.

Take the motor driver stuck on, delivering insulin continuously. Failing stuck-on is rare, a single hardware fault, so P1 is a 2. With no controls yet, the harm follows almost inevitably, so P2 is a 4; multiplied and rounded, the probability of harm is a 2. Severity is fast-onset, life-threatening hypoglycemia, a 5. Probability 2, severity 5 lands in the amber band. Now the hazard fault tree analysis found and FMEA structurally could not: an occlusion, a silenced alarm, and a sleeping patient, all at once. Multiply one wear cycle in twenty by one in ten by a third and you land near 1 in 600, the lowest probability rating, a 1. Same severity, still a 5. Probability 1, severity 5 sits one column to the left of the motor driver, in yellow instead of amber, despite being every bit as severe if it happens. Two hazards, identical worst case, very different probability, purely because one needs three ordinary things to line up and the other needs one part to fail. That is why you need both techniques feeding this matrix.

The rest fill in the range. The moisture combination, extended wear plus water degrading the seal, scores probability 2, severity 3, comfortably green, a reminder that not every combination fault tree analysis finds turns out to be dramatic. The connectivity path, an OR gate where either a DIY loop command or a mid-command Bluetooth drop is enough, scores probability 3, severity 4, and lands in red. The reservoir seal, a gradual wear-out, scores 3 and 3, dead center, where most of a device's risk really lives. And the Bluetooth radio, whose frozen reading looks exactly like a working one, also scores 3 and 3, for a completely different reason: a conditional harm that depends on what a distracted person does next. That last one is exactly the hazard RPN would have ranked far too high, purely for being hard to detect.

What the grid does, and does not, decide

Once all eleven are scored and plotted, the value of a matrix like this is plain: eleven rows of dense reasoning, compressed into something you can read across a conference table in about four seconds. Anyone on the team, or in an audit, can see where the pressure is.

But notice what the grid does not tell you. It does not say which of these are acceptable to ship with, or which need a control before you can move forward. A color is a description of where a risk sits. It is not yet a decision about what to do with it. That separation is deliberate, and it is one of the standard's better design choices: estimation stays objective, and evaluation, the judgment call of what your organization is willing to accept, is kept as its own step so the math never quietly gets contaminated by it.

Back in Module 5 the promise was that WearPump's hazard table would reach a dozen or more situations once both techniques finished running. Eleven are scored here, with more of the device still to analyze the same way. The promise holds.

If you take one thing from this, take this: your FMEA's risk priority number and your ISO 14971 risk are not rivals, and they are not the same number wearing two names. They are two different questions. Answer both, on the same hazard, and you know something neither one could tell you alone.

Module 7 applies WearPump's actual acceptability criteria to each scored hazard and decides which ones require a control before this device ever reaches a patient. More on product strategy, roadmaps, and fractional CPO work at baserealitygroup.com, and I write more broadly for founders too, in a newsletter called The Build, at davesaunders.net.

Get the next guide

ISO 14971 is the first Operator’s Guide. ISO 13485, IEC 60601, and the GHG Protocol are in the queue. Leave your email and each new guide lands in your inbox the day it ships. Nothing else, no drip sequence.

Done. You’ll get the next guide when it ships.

The standard is free. So is this course. Building the product is the work.

Work With Dave

Prefer a smaller first step? Book a $500 one-hour working session →

Dave Saunders

Dave Saunders is the founder of Base Reality Group and a Fractional CPO for hard-tech founders. He was a founder and operator at Galen Robotics, where the surgical-robotics platform earned FDA De Novo authorization in 2023, and he managed a 35-patent portfolio licensed from Johns Hopkins. He wrote Founders Who Finish and publishes The Build. More about Dave →