The FDA is preparing to give the healthcare industry something it has been waiting for: clearer rules for generative AI in medical devices.
According to reporting by STAT, the agency is developing a regulatory framework for devices using generative AI, planning both broad guidance on the technology overall and more narrowly tailored “specialty” guidance for particular applications that are especially complex or high-interest.
Why existing rules do not fit
Medical device regulation rests on an assumption that generative AI violates: that a device does the same thing each time it is used.
A blood pressure monitor, an imaging system, even a conventional machine-learning classifier produces a deterministic output for a given input. Validation therefore means testing the device across representative conditions and demonstrating it performs to specification — and if it performs correctly during testing, it will perform correctly afterwards.
Generative models do not work that way. The same input can produce different outputs, the range of possible outputs is effectively unbounded, and the model may generate something no test anticipated. Traditional validation asks whether the device gives the right answer; here there may be no single right answer, and no way to enumerate what the device might say.
The second problem: models change
Regulated devices are approved in a fixed configuration, and a manufacturer changing the product materially must return to the agency.
That sits badly with software built on foundation models that vendors update continuously. A clinical tool built on a general-purpose model can behave differently after an update the healthcare vendor did not make and may not control.
The FDA has developed mechanisms for anticipated software change — allowing manufacturers to pre-specify how a model may be modified and how modifications will be validated. Whether that framework extends to a model whose underlying capabilities shift beneath the application is precisely the sort of question guidance would need to address.
What officials said
“The ecosystem is expecting clarity. We seek to provide that clarity,” said Rick Abramson, director of the FDA’s Digital Health Center of Excellence.
He added that industry can expect “not only broad guidance on the overall topic of generative AI, but also some more narrowly constructed specialty guidance on particular…topics of special interest or special complexity.” No specific timeline was attached.
Why a two-tier structure makes sense
The split between broad and specialty guidance reflects how differently these tools are used.
A model drafting a clinical note from a recorded consultation, a model summarising a patient record, and a model suggesting a differential diagnosis raise entirely different risks. The first mainly risks transcription errors a clinician should catch on review. The last participates in clinical reasoning, and its failure modes are subtler and more consequential.
General principles about validation, monitoring and transparency apply across all of them. The specifics of what evidence is required cannot sensibly be uniform.
The regulatory gap that already exists
Worth noting what is happening while guidance is developed: a great deal of generative AI is already deployed in healthcare, and much of it sits outside device regulation entirely.
Ambient documentation tools, administrative assistants and drafting aids are widely used on the reasoning that they support rather than make clinical decisions, and that a clinician reviews the output. Whether that reasoning holds under pressure — when a clinician reviews the fiftieth AI-drafted note of a shift — is a question the current framework does not really ask.
Why clarity helps developers
Uncertainty is expensive in both directions. Developers unsure whether a product will be regulated as a device face a choice between building for the stricter standard, which is costly and slow, or building for the looser one and risking a later determination that the product needed clearance.
That uncertainty tends to favour incumbents, who can absorb regulatory risk, over smaller developers who cannot. Clear rules generally widen the field rather than narrowing it, even when the rules themselves are demanding.
What this is, and is not
This is a statement of regulatory intent, not a finalised rule, and no timeline was given.
Guidance documents are also not binding law — they describe the agency’s current thinking and what it expects to see, which in practice shapes behaviour strongly while leaving room for case-by-case judgement.
The question underneath all of this
Every specific issue — non-determinism, model updates, where to draw the device boundary — reduces to one problem regulators have not faced at this scale before: how do you validate a system whose behaviour you cannot fully characterise in advance?
Traditional validation is exhaustive in principle. You define the operating conditions, test across them, and demonstrate performance within specification. The approach works because the space of possible behaviours is bounded and enumerable.
Generative systems break that. Their output space is effectively infinite, and rare failures may be the important ones precisely because they are unusual and unanticipated. Approaches borrowed from other fields — adversarial testing, statistical performance bounds, continuous post-deployment monitoring with defined intervention thresholds — each address part of it and none is a complete substitute.
Which is why guidance is genuinely hard to write rather than merely slow to produce, and why the specialty-guidance approach is sensible: the answer probably differs by application, and pretending otherwise would produce rules that fit nothing well.
The substantive question guidance will have to answer is what evidence demonstrates that a non-deterministic system is safe and effective. Everything else follows from that, and it does not have an obvious answer. Regulatory news.