The US Food and Drug Administration (FDA) is officially changing the rules of engagement for generative AI in healthcare. The agency is ditching formal, static benchmarks in favor of a multi-tier assessment of clinical competence: models will now have to pass comprehensive qualification exams similar to those taken by practicing physicians. Traditional testing on fixed datasets has proven ineffective: generative systems process open-ended input, produce variable outputs, and systematically suffer from confabulations—plausible-sounding clinical hallucinations that can cost patients their lives.

The regulatory framework is anchored by a two-axis risk matrix, tying requirements directly to the system's level of autonomy and the potential severity of a medical error. The pre-market filter now combines standardized testing with real-world clinical simulations. Developers building wrappers on third-party foundation models will feel the sharpest pain: they remain directly legally and regulatorily liable for any drift or failure in the underlying LLM. Because neural network behavior degrades over time, the FDA is making continuous post-market monitoring a mandatory prerequisite for staying on the market.

For MedTech startups and enterprise systems integrators, this effectively shuts the door on rushing raw API wrappers to market. Certification cycles will stretch significantly, compliance overhead will become a major budget line item, and hospital market access will shift from a one-off formality to ongoing recertification.

AI in HealthcareGenerative AIAI RegulationAI SafetyLarge Language Models