Operator’s Guides — ISO 14971 · Module 9 of 12

Residual Risk, Done Honestly: When \"We Reduced It\" Isn't the Same as \"It's Acceptable\

ISO 14971 rarely lets you drive a risk to zero. Here's how you re-score what's left, build a benefit-risk case an auditor will accept, and tell the limit of your engineering apart from a risk actually worth shipping.

All ISO 14971 modules
Module 09 28:31 video + article

Watch this module

Module 9 of the video course. The article below covers the same ground in written form, so you can watch, read, or both.

# Residual Risk, Done Honestly: When "We Reduced It" Isn't the Same as "It's Acceptable"

Most risk-management tutorials end with a tidy matrix full of green cells. Real ones don't. When you control a hazard honestly, you usually don't get to zero. You get to smaller. The risk is lower than it was, and it's still sitting there on the page. So you end up staring at a hazard that responded to everything you threw at it and still didn't clear your line, and the question that stops teams cold is a simple one: now what?

That is the whole job of this module. When a risk won't go to zero, how do you decide it's acceptable, and how do you prove that to someone who wasn't in the room? ISO 14971 has a specific process for exactly this moment, and it runs straight into one of the most consequential arguments you'll ever write in a risk file.

This is Module 9 of The Operator's Guide to ISO 14971. I'm Dave Saunders. I've spent more than 30 years commercializing technology, almost 20 of them in medical devices and regulatory work, including several surgical robots. The series builds a full risk management file on one running device: WearPump, a wearable subcutaneous insulin pump for Type 1 diabetes. Four clauses close that file, and we take them in order: 7.3, residual risk evaluation; 7.4, benefit-risk analysis; 7.5, the new hazards your controls create; and 7.6, the completeness check.

Where WearPump stands

Coming out of Module 8, the WearPump file is not a clean sheet of green. The motor driver fault, the one that can over-deliver insulin, came down out of the elevated band into something we can live with. The connectivity hazard moved off red. The yellows came down too. And along the way we created a brand-new hazard of our own, alarm fatigue, by reaching for a louder alarm. That whole picture is what we close out here.

Clause 7.3: re-scoring what's left

Residual risk is whatever is left after your controls are in place. Clause 7.3 says you evaluate it, one risk at a time, against the criteria you already wrote down. Not new criteria you invent now that you can see the answer. The exact acceptability line you set back in your risk management plan in Module 3, before you knew how any of this would land. You set the bar first, then you measure against it. That order is the entire point.

Here's what re-estimation actually does. After a control goes in, you score the risk again. Both numbers can move, but they don't move the same way. A control almost always changes probability, and it rarely touches severity, for a clean reason: a control interrupts the chain of events between the hazard and the harm. It makes the harm less likely. That's a probability effect, every time.

Severity is different. If the control fails and the patient is exposed anyway, the harm is the same size it always was. An overdose is an overdose whether or not there was an alarm that failed to stop it. So the default rule is blunt. You may lower probability when you can explain why. You don't get to lower severity because you feel better about the risk now. The harm is the harm.

When severity legitimately drops

There is one real exception, and WearPump used it twice, so it's worth getting exactly right. Severity can drop when the control changes the injury itself, not the odds of it.

The motor driver fault started as high as severity goes, because a full electronic runaway could deliver a fatal amount of insulin. Then we added two hard caps: a firmware dose ceiling, and a 3 mL reservoir that physically limits how much can ever come out. Walk the worst case again and it's no longer a lost life. It's a serious overdose the body can survive with treatment. The design changed the size of the harm, so severity genuinely comes down, from an S5 to an S3.

The connectivity hazard did the same thing a different way. It came in at severity 4, a critical harm, because a bad wireless command could drive a dangerous dose. The fix wasn't a warning about pairing. It was a design decision: the pump is completely safe on its own, with the phone switched off and thrown in a drawer. Once the phone can never command anything unsafe, the worst case of a connectivity failure is no longer a dangerous dose. Severity falls to an S2.

That's the line you have to hold in the documentation. If severity drops, you show the mechanism that made it drop: the dose ceiling, the reservoir geometry, the safe-standalone design. The specific thing that shrank the injury. What you can never write is that you lowered a number because the team was more comfortable. An auditor will challenge any severity drop that doesn't come with a mechanism attached, and they'll be right to.

Probability gets the same discipline in reverse. You don't just write a smaller number. You name the control, state how it cuts the odds, and point at the evidence. Not "the sensor helps." Something concrete: the occlusion sensor caught the blockage in 95 out of 100 bench tests, so probability drops a full band, and here is the report. The bad version is "we added a control and re-estimated lower." The good version has a number and a report attached.

The mistake that fails more files than any other

It hides in the gap between two sentences that sound identical and are not.

"We can't reduce this any further" is an engineering fact, a statement about the limits of what you can practicably build. "This risk is acceptable" is a judgment about whether what's left is worth it. One is about your hardware. The other is about the patient.

Teams weld those two together and it costs them. Reaching the end of your engineering does not automatically make the leftover risk okay. So for a stubborn residual you owe two things, not one. First, the demonstration that you hit the limit of practicable control: you tried the design changes, you tried the guards, and here's why you stopped where you stopped. Second, the argument that the risk that remains is outweighed by something. Write only the first and you've described an engineering dead end and called it a safety case. It isn't one. And the second argument has a clause of its own.

Clause 7.4: benefit-risk is not an escape hatch

When a residual risk won't clear your criteria and you genuinely can't reduce it further, Clause 7.4 gives you one more move: you may weigh whether the benefit of the device outweighs that specific risk. This clause has a bad reputation it half deserves. Done lazily, it's the escape hatch teams reach for when they don't want to engineer something out. Done properly, it's the most evidence-heavy paragraph in the entire file.

Two rules govern it. First, the benefit has to be clinical. A better outcome for the patient: longer life, fewer emergencies, a real improvement in how they live. What does not count, and the standard is explicit, is anything on your side of the table: not the schedule, not the money you already spent, not what a redesign would cost this quarter. Benefit means benefit to the person wearing the device.

Second, you never weigh a device against zero risk, because zero isn't on the menu for anyone. You weigh it against what the patient would actually be doing instead. For a Type 1 diabetic, the alternative isn't a risk-free life. It's multiple daily injections, 4 to 6 needles a day, for the rest of their life, and that therapy carries its own serious risks. That's your comparator, and you have to name it.

The residual that stays catastrophic

Now the hazard that survives everything. It's the silenced-alarm situation: insulin backs up behind an occlusion, the alarm that should catch it has been silenced, and the patient is asleep. When the blockage finally gives way, the backed-up dose goes in all at once. A severe low, at three in the morning, with no alarm to wake anyone. We drove the probability way down. But the severity hasn't moved, because if that chain completes, a person can die in their sleep. Rare, yes. Catastrophic, still.

And we're honestly at the limit on it. We already made the smart move in Module 8, an alarm that fires only when it truly matters so it doesn't get tuned out. We can't make an undetected occlusion during sleep impossible without taking away function the patient needs. So probability is as low as we can practicably get it, and severity is inherent. That's exactly the situation Clause 7.4 was written for. We stop trying to shrink it, and we start weighing it.

The evidence, and how much it can carry

Here's the clinical evidence we've held back since Module 7. Pickup and Sutton, a study in Diabetic Medicine in 2008, pooled the trials comparing pump therapy against injections in Type 1 diabetes. Severe hypoglycemia, the dangerous lows that put people on the floor, happened far less often on a pump: in the randomized trials, injections carried close to three times the rate; across all the studies together, more than four times. The device attacks the exact harm we're worried about, and it attacks it hard.

A benefit-risk case is only as strong as the evidence underneath it, so it helps to know the hierarchy. Randomized controlled trials and the meta-analyses that pool them sit at the top. Registry data and real-world evidence sit below that, useful but noisier. A single expert saying "it seems fine" sits at the very bottom. For an insulin pump you're lucky, because the literature comparing pumps to injections is deep and decades old, so you get to stand on the strongest rung.

Put a floor under how common the harm is, too. In the DCCT, the landmark trial published in the New England Journal of Medicine in 1993, severe hypoglycemia on injections ran around 60 events per 100 patient-years. That's not a freak occurrence. It's a routine danger of the therapy this pump replaces. There's a quality-of-life piece as well: injections mean 4 to 6 needles a day, somewhere between 1,460 and 2,190 a year, while the pump replaces that with one cannula change every three days, roughly 120 a year. For a person managing a lifelong disease, that difference shows up in whether they stick with the therapy at all.

Now assemble it for the silenced-alarm risk. Yes, there's a rare path where the device fails to catch a dangerous low and someone is harmed. But the same device, used as intended, prevents far more of those exact dangerous lows than the injections it replaces. The benefit and the risk are measured in the same currency, severe hypoglycemia, and the benefit is the larger number by a wide margin. The residual is real, and the benefit outweighs it.

You don't say that with confidence. You say it with citations. A real benefit-risk argument has four parts, every time:

  1. 1. Name the residual risk, its score, and the fact that you couldn't reduce it further.
  2. 2. Name the comparator therapy the patient would use instead.
  3. 3. Cite the clinical evidence, by author and year, with the actual numbers.
  4. 4. Draw an explicit conclusion. No "it appears that," no "we believe."

The conclusion reads like a verdict, not a wish: The residual risk of an undetected occlusion during sleep is accepted. It cannot be reduced further without removing clinical function the patient depends on, and the device delivers a documented reduction in severe hypoglycemia, relative to multiple daily injection therapy, that outweighs it. That's a sentence a reviewer can check. They can pull the study, confirm the comparator, and run the logic. And that is exactly the checklist a notified body reviewer runs: did you name the risk, name the comparator, cite real clinical data, and reach a conclusion. Four boxes. If one is missing, it reads as an opinion dressed up as an analysis, and it comes back. One more thing: accepting a residual doesn't mean hiding it. A significant risk you've decided to live with gets disclosed to the person living with it, in the instructions for use. Accepted is not the same as buried.

Clause 7.5: the loop nobody runs

Clause 7.5 is the one I see skipped more than any other in the whole standard: the risks your own controls create. Every control you add changes the device, and a changed device can fail in a brand-new way it couldn't before. A safety feature is never free.

We don't have to hunt for an example, because we walked straight into one. On the silenced-alarm hazard, the obvious instinct was to make the alarm louder, harder to silence, add a backup tone. All reasonable. All aimed at safety. And every one of those moves quietly makes the device more dangerous. It's called alarm fatigue, and it's not a theory. The Joint Commission, which accredits hospitals across the United States, reviewed alarm-related events over a three-and-a-half-year window: of 98, 80 ended in a patient death, and by their estimate somewhere between 85% and 99% of alarm signals need no clinical action at all. So people stop trusting the alarms and tune them out. A louder alarm doesn't add safety. It adds noise, and the noise gets someone killed.

Read that back against what we were about to do. We were going to protect the patient with a more aggressive alarm, and a more aggressive alarm is itself a documented cause of patient death. The control became the hazard. The real fix wasn't more alarm, it was a smarter alarm, one that fires only when it genuinely matters so it keeps its meaning. And here's what 7.5 actually demands: that new hazard, alarm fatigue, doesn't just get noticed and waved at. It goes back into the process as its own line, with its own estimate, its own control, its own residual. You added a control, you asked what it just broke, you found something, and you run it through the whole loop again. That's not the process failing. That's the process working as designed. Once you see it, it's everywhere: we also considered dropping the adhesive patch for a mechanical strap and clip, which would have solved the skin reaction cleanly, and introduced a new pressure-injury hazard and a device more likely to pop off mid-exercise. Every control is a trade, and 7.5 is where you're forced to count the cost, not just the benefit.

Clause 7.6: the completeness check

The last clause is the gate before you're allowed to freeze the design. Before this design goes anywhere, you confirm that every hazardous situation you ever identified has been dealt with. No gaps. No hazard you wrote down in Module 5 and quietly lost track of. You run it as a cross-check: take the hazard list from your analysis and walk it against your risk file, one row at a time. Every hazard needs four things, and you're hunting for three kinds of gap.

Every hazard needs:

You're hunting for:

Run it on WearPump's eleven hazards. Seven we scored and worked on camera across these modules: the motor driver, the connectivity path, the silenced alarm, the reservoir seal, the Bluetooth lockup, the moisture combination, and the adhesive site. The other four we named but didn't fully walk, the pressurized reservoir, the concentrated insulin, the electrical energy, and the original adhesive and cannula hazard, each still gets its row, documented at its level, with the reason it sits where it sits. Eleven identified. Eleven addressed.

And now the honest part the tidy tutorials skip. This file does not end all green. The silenced-alarm risk is accepted, not eliminated. The motor driver still carries a residual, smaller than before, but not zero. That's correct. A risk file where every single cell is green is usually a file that lied somewhere, or a device so trivial it never needed this process. Real risk management ends with residuals you've chosen, on purpose, with your eyes open. So be precise about what "closed" means: not risk-free, but every hazard addressed, every acceptance documented, every benefit-risk argument on paper with its citations, every control backed by verification. Complete for this version of the design, not forever. A change tomorrow reopens it. And closing is a formal action, not a vibe: you capture the date, name the person with authority to approve it, and they sign the completeness check.

The question the standard won't answer

You can do all of this perfectly, the re-estimation, the benefit-risk, the completeness check, every citation in place, and there's still a question underneath it the standard won't answer for you. If the person wearing this pump were someone you love, would the residual risk you just documented as acceptable still feel like enough? Not on paper. In your gut, at two in the morning.

That question doesn't have a clean answer, and it isn't supposed to. The standard gives you a disciplined, evidence-based way to make the decision. It doesn't make the decision for you. Because the decision still belongs to people: engineers, a clinical lead, a quality manager, a regulatory lead, who have to look at each other and say, out loud, that they believe this device is safer than the alternative and they're ready to defend that to anyone who asks. That's what "acceptable" actually means past the matrix cell. It's a claim real people put their names to, and the framework is how you make that claim honestly instead of hopefully.

The recap

When you control a risk honestly, you usually reach smaller, not zero. Clause 7.3 makes you re-score what's left against the line you set first: controls move probability, and only move severity when the design changed the injury itself, with the mechanism shown. Clause 7.4 handles the residuals that won't clear, by weighing clinical benefit against the risk, in the same currency, against the real comparator, cited by author and year, with an explicit conclusion. Clause 7.5 sends the new hazards your controls create back through the loop. Clause 7.6 confirms every hazard is addressed before you freeze the design. The file closes with residuals you chose, documented and signed, not with a wall of green.

Module 10 asks the question a row-by-row review can't: are all of these residual risks, together, on one patient, still acceptable? That's Clause 8, the overall residual risk, and it catches what the line-by-line view misses.

---

If this way of working through a device is useful to you, I write more about challenges like this in my newsletter, The Build, at davesaunders.net, and you can find the rest of my work at baserealitygroup.com. Subscribe so the next module in this series finds you when it lands.

Get the next guide

ISO 14971 is the first Operator’s Guide. ISO 13485, IEC 60601, and the GHG Protocol are in the queue. Leave your email and each new guide lands in your inbox the day it ships. Nothing else, no drip sequence.

Done. You’ll get the next guide when it ships.

The standard is free. So is this course. Building the product is the work.

Work With Dave

Prefer a smaller first step? Book a $500 one-hour working session →

Dave Saunders

Dave Saunders is the founder of Base Reality Group and a Fractional CPO for hard-tech founders. He was a founder and operator at Galen Robotics, where the surgical-robotics platform earned FDA De Novo authorization in 2023, and he managed a 35-patent portfolio licensed from Johns Hopkins. He wrote Founders Who Finish and publishes The Build. More about Dave →