Medical device software has been verified the same way for decades: write down the inputs, run them, and compare each output to a known answer. On August 18 the FDA's device center said that method probably won't work for generative AI, because the range of possible inputs and outputs is too large to test. Its new discussion paper floats a replacement modeled on how clinicians get credentialed. The proposal shifts much of the proof from the day of submission to every day after it, and onto a foundation model the startup didn't build.
Credentials, then monitoring
The paper, Considerations for the Regulation of Generative AI-Enabled Medical Devices, is not guidance and sets no requirements. Comments close October 19. It sketches a risk chart that weighs how much the software does on its own against how bad the harm is when someone relies on a wrong answer. Then it describes a two-part evaluation. The finished device, in the configuration it will ship in, gets benchmarked on clinical knowledge, safety behavior, robustness and knowing when to defer to a human. A clinical confirmation step checks real-world performance, with the rigor scaled to risk, anywhere from a retrospective look at patient data to a randomized trial.
The agency also says it may accept more uncertainty before launch in exchange for more monitoring after. Its options include re-running the benchmark on a schedule and after defined changes, independent clinicians reviewing samples of real inputs and outputs, and watching for drift. The paper asks whether AI agents could do some of the supervising.
The component nobody can freeze
Most generative-AI devices will sit on a foundation model built by someone else. The paper's proposed help is a voluntary Foundation Model Master File. Model developers could give the agency confidential information on architecture, training data, known failure modes and update commitments, and a device maker could reference that file with the developer's permission. The paper is explicit that a master file authorizes nothing. The device maker stays responsible for a component it didn't build and can't fully inspect, from a supplier with no obligation to tell the FDA anything.
Cooley, the law firm, reviewed the public comments filed so far and reports worry that some models are updated hundreds of times a day. The paper names a change to the underlying model as a trigger for re-running the benchmark. A fast-moving vendor could hand a device maker a stream of regulatory events it never scheduled.
Who this lands on
The proposal reaches fewer founders than the word AI implies. The AI already shipping in cleared devices, and in the surgical robotics research I work around, is image and video processing with no language model in the loop. Those teams can read this as someone else's problem. The paper is aimed at founders putting a chat interface, a summarizer or an agent in front of a clinician or a patient, and it raises the idea of ranking a patient-facing function as higher risk than the same function aimed at a clinician.
For a founder putting a chatbot or an agent in front of patients or clinicians, the decisive work comes before any submission: four questions for the model vendor, with the answers written into the contract. Can we name the exact model version we validated? Can we keep running that version after the vendor moves on? Will we get notice before it changes? Can we re-run our own benchmark within days when it does?
If any answer depends on the vendor's goodwill, the model belongs in the risk file as a single-source part. Hardware teams already know how to handle one. A sole-source chip with no end-of-life notice gets a mitigation, a second source or a redesign. I expect the same treatment to reach industrial robots and defense systems that run vision-language models, because their buyers can't list every input either. They'll ask for competency evidence and a monitoring plan, and the seller will need the same four answers.
Dave's take
The competency model is a sensible idea and worth supporting in the docket. The master-file section is the weak spot. A voluntary program works only if the large model developers volunteer, and nothing in the paper requires them to, so the only master file a startup can count on today is the one it writes into its own vendor contract.
From Dave’s video library
Dave traces the study behind the claim that AI beats lawyers at contract review, and finds the AI helped build the answer key it was graded against.
I’m here to help you scale.
Work With DavePrefer a smaller first step? Book a $500 one-hour working session →
Dave Saunders is the founder of Base Reality Group and a Fractional CPO for product companies. He was a founder and operator at Galen Robotics, where the surgical-robotics platform earned FDA De Novo authorization in 2023, and he managed a 35-patent portfolio licensed from Johns Hopkins. He wrote Founders Who Finish and publishes The Build. More about Dave →