Watch this module
Module 10 of the video course. The article below covers the same ground in written form, so you can watch, read, or both.
# Every Risk Was Acceptable. The Whole Device Might Not Be. (ISO 14971 Clause 8)
Picture your risk file with every row green. Every hazard found, every one controlled, every residual signed off as acceptable, every box ticked. You're done. Except ISO 14971 has one more question, and it's the one that catches the most careful teams, because it isn't about any single risk on the list. It's this: are all of them, together, on the same patient, every day, for years, still acceptable? Because ten green risks can add up to a device you shouldn't ship.
Here's how that actually goes wrong. Imagine a device with twenty residual risks, every one individually acceptable, every one signed off. Then someone finally lays them out on one table and notices that eight of the twenty land hardest on exactly the same kind of patient: the elderly one, managing the device in low light, with poor eyesight, at the end of a long day. Each risk was fine on its own. Stacked on that one person, they're a systematic problem nobody saw, because nobody looked at the whole at once.
Individual risk evaluation answers one question at a time: is this hazard, on its own, acceptable? It cannot answer the question about the whole device, and those are genuinely different questions with different answers. This is the step where a file that looks finished turns out not to be, and it gets skipped more than almost any other, because it feels like paperwork after the real work is done. It isn't paperwork. It's the real work, one more time, at a higher altitude.
This is Module 10 of The Operator's Guide to ISO 14971. I'm Dave Saunders. I've spent more than 30 years commercializing technology, almost 20 of them in medical devices and regulatory work, including several surgical robots. Across this series we've built the WearPump risk file from nothing: eleven hazardous situations, identified, estimated, evaluated, and controlled, each one closed individually in Module 9. Now we stop looking at rows and look at the shape of the whole file.
Clause 8, and the door that isn't there
Clause 8 is the overall residual risk evaluation. After every individual risk is controlled and evaluated, you gather all of that residual risk together and make one judgment: considering the whole picture, and considering the medical benefit the device gives the patient, is the overall residual risk acceptable? If yes, you document the conclusion with your reasoning and the file closes on it. If no, you have exactly two options. Add more control, or abandon the device concept.
Sit on that for a second, because it's the part with teeth. There's no third door. There's no option where the overall residual risk comes back unacceptable and you ship anyway because the launch date is set, or the investors are waiting, or the market window is closing. The standard gives you no commercial override on this decision. It's one of the few places where a standard about paperwork quietly has real power over a business.
And it has to be its own evaluation, because the mistake is subtle. Teams finish the individual evaluations, see a file full of acceptable verdicts, and assume the overall evaluation is automatically satisfied. Every part passed, so the whole must pass. That's not how complex systems fail. A device can be a collection of individually reasonable risks that, stacked on one patient over years, add up to something you wouldn't put your name to. Clause 8 exists precisely to catch what the row-by-row view is structurally blind to.
What "overall" actually means: three dimensions
The word overall is carrying three different kinds of weight.
The first is cumulative burden. A patient wearing the WearPump isn't exposed to one risk at a time. They aren't carrying the occlusion risk on Monday and the connectivity risk on Tuesday. They're carrying all of it, at once, every hour the device is on their body. Every failure mode in the file is live, simultaneously, on one person. The overall view has to reflect that total load, not the eleven separate slices you evaluated it in.
The second is clustering by harm, and this is the one that earns its keep. When you lay every residual risk on a single table instead of reading them one at a time, patterns jump out that were invisible at the row level. Risks that looked unrelated turn out to be different roads to the exact same place. On the WearPump, that's not hypothetical. Look at where our residuals actually lead. The motor driver fault over-delivers insulin. The silenced-alarm situation dumps a backed-up dose all at once when an occlusion finally lets go. The connectivity path could push an unintended command. The dosing algorithm could stack a double bolus. Four separate hazards, from four different corners of the device, hardware, firmware, wireless, and software, and every one of them ends in the same place: a severe low. Severe hypoglycemia, the dangerous drop that can put a person on the floor, or worse.
Each of those four was individually acceptable. We closed each one honestly, on its own terms. But row by row, you never once see that they all drain into the same catastrophic outcome, and that matters for a hard, quantitative reason. The real-world probability of a severe low from this device isn't any one of those four risks. It's all four added together. Four small streams feeding one river, and the river is the number that actually reaches the patient. Individual evaluation hides that sum. Clause 8 is what makes you compute it.
The third dimension is exposure over time. The WearPump is a chronic-use device. A patient wears it around the clock, changes it every three days, and does that for years, maybe decades. Every three days is roughly 120 wear cycles a year. Over a decade, that's more than a thousand exposures to every one of these risks. A per-cycle probability that looks reassuringly tiny gets multiplied by a thousand, and reassuringly tiny can become near-certain across a life.
Put those together and overall residual risk isn't a headcount of how many yellow cells you have left. It's a structured look at the whole weight of the device resting on one real person over the entire life of their therapy. You can see why a stack of individually green rows doesn't answer it.
Judge it against what?
Acceptable is a meaningless word in a vacuum. You need something to compare that total picture to, and if you pick the wrong comparison, the entire evaluation is worthless no matter how rigorous the rest looks.
Here's the rule that fixes it. You never judge a device against zero risk, because zero is not a choice anyone actually has. A Type 1 diabetic doesn't get to pick a risk-free life. The only comparison that means anything is against the state of the art: what would this patient be doing right now, today, if your device didn't exist? For the WearPump, the answer is multiple daily injections, and that therapy is not safe. It carries its own serious, well-documented risks, and those risks, not zero, are your benchmark.
So characterize the comparator in hard clinical terms. On multiple daily injections, severe hypoglycemia, the same dangerous low our residuals cluster on, is common. In the DCCT, the landmark trial of intensive insulin therapy published in the New England Journal of Medicine in 1993, it ran around 60 events per 100 patient-years. Read that slowly: 60 severe lows, per hundred patients, per year, on the standard of care. That's not a rare edge case. It's the routine, documented danger of the therapy this device is trying to replace, and it's the number every part of your evaluation gets measured against.
Now characterize the benefit in the very same currency, because a benefit measured in different units than the risk isn't a comparison, it's a dodge. Pump therapy versus injections, in the pooled trials, cut severe hypoglycemia by a wide margin. Pickup and Sutton, in Diabetic Medicine in 2008, found injections carried close to three times the rate in the randomized trials, and more than four times across all the studies together. Same harm, same units, one therapy with three to four times more of it than the other.
So run it. The WearPump's residual risks cluster on severe hypoglycemia, the honest weak point of the device. And the very same device drives severe hypoglycemia down by three to four times against the injections it replaces. In the one currency that matters most, the aggregate residual of our device sits far below the comparator. That's the overall residual risk argument, and it's strong precisely because it's measured in the same harm on both sides of the scale.
To see the shape of the calculation and not just the vibe of it, put an illustrative number on it. Suppose the four risks that cluster on severe lows, taken together and translated from their residual probabilities into an event rate, land the device somewhere around 15 to 20 severe lows per 100 patient-years. That's an estimate built from our own file, not a measurement, and it's labeled as such. Set it next to the comparator: 60 on injections, 15 to 20 on the pump. The device isn't risk-free, and it's also plainly, quantitatively safer on the exact harm we were most worried about.
Don't let a strong number make you sloppy, though, because there's one category where the device adds risk injections simply don't carry: the skin. The adhesive site, the cannula, the chance of a local infection or a reaction where the device meets the body for three days at a time. Injection patients wear nothing between doses; ours wear the device continuously. We own an entire class of harm the comparator doesn't have. The honest evaluation names that out loud, estimates it, and notes that these are lower-severity harms, mostly local and treatable, traded against the far more dangerous lows we're preventing.
One discipline on those numbers, because this is exactly where files overreach and get caught. You don't need one perfect, precise harm rate for your device. Nobody has that before launch, and pretending you do is a tell. What you need is a defensible direction, grounded in the probability estimates already in your risk file and in the published literature on pumps like yours. When you put a device number on the table, mark it as the estimate it is. The comparator figures are sourced and citable; your device figures are reasoned projections. Keep that line visible and no reviewer can accuse you of grading your own homework.
The statement, and who signs it
All of that resolves into a single required output, and it isn't optional: the overall residual risk statement. This is not a cover letter and not a marketing line. It's a formal document in your risk file that captures your team's evaluated judgment, backed by evidence, in a form a stranger can audit. It has four parts, the same shape as the benefit-risk argument from the last module, scaled up to the whole device.
- 1. The aggregate residual risk summary. What the residual picture actually is, the clustering, the estimated rates, honestly stated.
- 2. The clinical benefit evidence, cited, not asserted. Every benefit claim carries a reference.
- 3. The comparator risk profile. The injection data, documented as rigorously as your own device.
- 4. The conclusion. An explicit, affirmative statement that the overall residual risk is acceptable in light of the benefit. Or a statement that it isn't, and what you're doing about it. No hedging in either direction.
Written out, the conclusion reads like a verdict a team is standing behind, not a hope they're floating: The overall residual risk of the WearPump, after all controls, has been evaluated against the benefit of continuous insulin infusion in adults with Type 1 diabetes. The comparator is multiple daily injection therapy, which carries a severe hypoglycemia rate of roughly 60 events per 100 patient-years. The device's aggregate residual harm sits substantially below that. The overall residual risk is acceptable. Every clause in that paragraph traces to a document or a citation, and that's the entire point of writing it that way.
This judgment doesn't belong to one person. The technical report that supports the standard is explicit: it's not a call an engineer makes alone in a spreadsheet at eleven at night before a deadline. The overall residual risk statement gets signed by regulatory, by clinical, and by quality. Three disciplines, three kinds of expertise, looking at the same aggregate picture and agreeing to it. If your acceptance statement carries one signature from one function, a reviewer notices immediately, and they're right to, because the whole idea is that no single perspective can see the whole risk.
One more phrase lives inside Clause 8: state of the art. It doesn't mean the theoretical best device that could ever exist. It means the current, generally accepted standard of care as it's actually practiced today, by real patients who miss doses and forget things and live messy lives. You compare the WearPump against real-world injections, with their real-world 60-per-hundred rate, not against some idealized therapy that eliminates all risk. There's a useful gut check for it: would an informed, reasonable patient, one who genuinely understood both the benefits and the residual risks, choose this device over the alternative? It's not a legal test or a formula. It's a way of standing in the patient's shoes and asking whether the trade you're documenting is one a clear-eyed person would actually take. If the honest answer is yes, your evaluation has a spine. If you find yourself hoping they wouldn't look too closely, that's worth listening to.
Acceptable is not silent
Say you get to yes. The statement is signed, the file is closing. You're still not done, and this is the piece teams treat as a formality and regulators treat as a hard requirement. Clause 8, read together with the standard's rules on information for safety, requires that significant residual risks are communicated to the person using the device. Deciding the overall risk is acceptable means the device is allowed to exist. It does not mean the patient gets to be uninformed about the risks they take on when they put it on their body.
Which residuals count as significant? The standard doesn't hand you a bright line, but the working test is clear enough: a residual is significant if it's something the patient could act on, or something that would change their decision to use the device at all. If knowing about it lets them catch a problem sooner, prepare for it, or choose more carefully, it belongs in the instructions. If it's truly invisible to them and nothing they do changes it, it may not. When in doubt, disclose.
For the WearPump, the instructions carry a specific set: a warning about occlusion and dangerous lows, with guidance to check blood glucose more closely after a site change; a precaution about skin sensitivity, with a hard contraindication for anyone with a known reaction to the silicone adhesive; a firm three-day wear limit regardless of how much is left in the reservoir; what to do if the phone loses its connection mid-dose; and the specific fault codes that mean stop and call support. Every one traces back to a named hazard in the file, and that tracing runs both directions, which is exactly what an auditor checks. Every significant residual points to a section of the instructions, and every safety line in the instructions points back to a risk. If a significant residual has no home in the manual, that's a finding. If the manual warns about something nowhere in your risk file, that's a different and more embarrassing finding, because now your own two documents contradict each other in front of the reviewer.
Five ways it fails an audit
These are the exact patterns that generate findings and stall submissions.
- 1. The aggregate step is simply never done. Twenty acceptable rows, and the file just ends. No overall evaluation, no benchmark, no acceptance statement. The team assumed green rows summed to a green device, and the single most important judgment in the file was never made, only presumed. This is the most common one.
- 2. Benefit is claimed with nothing behind it. The statement says the device provides significant clinical benefit, and there's no trial, no registry, no citation near it. An assertion is not evidence; the reviewer will ask for the study, and if you can't produce it, the statement collapses.
- 3. The comparator is cherry-picked. The team quietly reaches for best-case injection numbers from a tightly controlled trial instead of the messier real-world rate. A flattering comparator makes your advantage look bigger than it is, and the distortion gets found.
- 4. The significant residuals never reach the instructions. Disclosure gets treated as the labeling team's problem, not the risk team's, and the traceability matrix ends up empty.
- 5. One function signs off alone. The engineer certifies the whole thing with no clinical and no quality judgment beside it, and the cross-functional requirement quietly didn't happen.
All five are avoidable, and all five come from treating Clause 8 as the administrative tail of a process you're ready to be finished with.
The recap
Individual evaluation answers one risk at a time; Clause 8 asks whether the whole device, on one patient, over years, is acceptable. "Overall" carries three dimensions: cumulative burden (all of it at once), clustering (many risks draining into one harm), and exposure (a thousand cycles across a life). You judge the aggregate not against zero but against the state of the art, in the same units, which for the WearPump means severe hypoglycemia at roughly 60 per 100 patient-years on injections versus a reasoned 15 to 20 on the pump. You write it up in four parts, cite the benefit, document the comparator, conclude without hedging, and get three functions to sign. Then you disclose the significant residuals in the instructions, traceable both ways.
Most risk files get read the way they were written, one row at a time. Clause 8 asks you to read yours all at once. So the question worth sitting with: if you laid every residual in your file on one table today, what harm would three or four of your risks quietly turn out to share, and who is carrying the heaviest version of that combined load?
Module 11 turns this evaluation into the one document that has to carry it: the risk management report, the thing a reviewer opens first, and the place where a brilliant analysis goes to die if you write it badly.
---
If this way of thinking through a whole device is useful to you, I write more of it in my newsletter, The Build, at davesaunders.net, and the rest of my work lives at baserealitygroup.com. Subscribe so the next module finds you when it lands.
Get the next guide
ISO 14971 is the first Operator’s Guide. ISO 13485, IEC 60601, and the GHG Protocol are in the queue. Leave your email and each new guide lands in your inbox the day it ships. Nothing else, no drip sequence.
Done. You’ll get the next guide when it ships.
The standard is free. So is this course. Building the product is the work.
Work With DavePrefer a smaller first step? Book a $500 one-hour working session →
Dave Saunders is the founder of Base Reality Group and a Fractional CPO for hard-tech founders. He was a founder and operator at Galen Robotics, where the surgical-robotics platform earned FDA De Novo authorization in 2023, and he managed a 35-patent portfolio licensed from Johns Hopkins. He wrote Founders Who Finish and publishes The Build. More about Dave →