Field Notes — August 30, 2026

The FDA Wants to Know How You'd Test a Device That Answers Differently Every Time

All Field Notes
August 30, 2026 Regulatory

Run the test twice, get the same answer twice. That assumption sits underneath every verification plan I have ever signed off on. Put a generative model inside the product and the assumption breaks. On August 18 the FDA published a discussion paper admitting it has no settled method to put in its place. The paper arrived with an open docket asking the industry for help.

The paper is called Considerations for the Regulation of Generative AI-Enabled Medical Devices, out of the Digital Health Center of Excellence inside CDRH. Nothing in it is draft guidance and none of it is binding policy. What the FDA put on the table is 26 numbered questions, a comment deadline of October 19, and a docket number to file against: FDA-2026-N-7874.

Read it even if you'll never file with the FDA. Nothing about the verification problem it describes is specific to medicine. If your product decides something, and somebody gets hurt when it decides wrong, these questions are coming for you too. The FDA is just the first regulator to write its confusion down in public.

Two axes, not one

The agency sorts these devices on two axes at once. The first is how independent the thing is, running from handing a clinician information that carries no recommendation all the way to acting on its own. The second is how bad the consequence is when the output is wrong.

Teams argue about the first axis constantly. The second one they write down almost never, and that's backwards, because the consequence is where the evidence burden comes from. A warehouse robot that suggests a pick path and a warehouse robot that executes it can run identical software. The reviewer is weighing the consequence profile, and those two robots don't have the same one.

"Assistive" is also a status you have to keep earning. Your roadmap is under constant pressure to break it. Each autonomy feature your sales team wants pushes the product further along the first axis, and somebody should have to approve that move in writing.

Competency instead of a verdict

For premarket work the paper sketches benchmarking a device across ten elements, covering safety behaviors, clinical proficiency, generalizability, and autonomous conduct. The benchmark then gets confirmed clinically, somewhere between a retrospective study and a prospective one. After launch the framework leans harder on ongoing monitoring than the agency does for a fixed-function device, on the grounds that the outputs vary.

Benchmark-plus-monitoring is a different shape of work from what most hardware teams budget for. A test plan that produces a verdict finishes. A test plan that produces a score has to be re-run, and somebody has to own that score for as long as the product ships. I've watched founders model regulatory work as a submission-shaped expense that ends the day the letter arrives. If the thing you shipped keeps changing its own behavior, the expense doesn't end, and the burn model is wrong by whatever it costs to staff the difference.

The part that decides is somebody else's code

The paper also floats a voluntary Foundation Model Device Master File. The company that built the underlying model files its confidential information once, and your submission points at it. Strip the paperwork away and the FDA is acknowledging something founders already know and rarely price: the component making the call inside your device was built by a supplier you don't control and can't fully inspect.

I've paid for that arrangement before. We sold internet software that ran on MacTCP, the Mac's TCP/IP stack, written originally by University of Michigan students and by then maintained by nobody at Apple. System 8 could page MacTCP out when virtual memory was on, and the stack threw up a debugger statement those students had left in the code. It wasn't even a crash. If you had MacsBug installed you told it to continue and you were fine. Our customers were defense contractors running a hundred thousand copies, and what they saw was a bomb icon, and it was all our fault.

None of that was our code. Every support call still said our product crashed. A foundation model is the same bargain with more zeros. You're shipping a dependency you didn't write, can't read, and can't patch, and in front of the customer and the reviewer the debt is yours. The master file gives that problem a paperwork channel. It doesn't solve it.

Dave's take

I'd file a comment. There are 26 questions open until October 19, and the operating reality inside the eventual framework will be whatever the companies who bothered to file put there. Everyone else gets to comply with what the people who showed up decided. Filing costs you an afternoon. You'll be living inside the answer for the next decade.

From Dave’s video library

Dave walks through how to tell a checkable AI claim from one nobody can check, using the model the government supposedly banned.

Dave Saunders

Dave Saunders is the founder of Base Reality Group and a Fractional CPO for product companies. He was a founder and operator at Galen Robotics, where the surgical-robotics platform earned FDA De Novo authorization in 2023, and he managed a 35-patent portfolio licensed from Johns Hopkins. He wrote Founders Who Finish and publishes The Build. More about Dave →