FDA opens comment period on new framework for regulating GenAI-powered medical devices
On August 25, 2026, FDA's Center for Devices and Radiological Health (CDRH) released a discussion paper seeking public feedback on a regulatory framework for generative AI-enabled medical devices, products built on LLMs and foundation models that produce open-ended, variable outputs and can evolve after deployment. Unlike traditional software, these devices can't be exhaustively tested against a fixed set of inputs and outputs, so FDA is proposing a competency-based approach modeled loosely on how human clinicians are trained and credentialed: a two-stage process combining large-scale "device benchmarking" (testing safety, clinical proficiency, and generalizability against prespecified criteria) with "clinical confirmation" (real-world or clinically representative evidence, ranging from retrospective case review to full prospective trials).
Risk would be assessed via a proposed two-axis framework, plotting a device's activity level (from passive information-giving to fully autonomous action) against the severity of consequences if its output is wrong. This is to determine how much premarket evidence is needed. The paper also floats voluntary "Foundation Model Master Files" so third-party model developers can share safety-relevant data with FDA confidentially, and raises open questions on postmarket monitoring, agentic AI systems, and how to handle changes made by foundation-model providers rather than device makers themselves.
This matters because it signals FDA's clearest attempt yet to adapt medical device regulation to AI systems that behave less like static software and more like adaptive decision-makers. It is a shift with major implications for how digital health and AI-driven diagnostic/clinical-support companies will need to design evidence generation and monitoring programs going forward. Comments are due October 19, 2026, via docket FDA-2026-N-7874.
Why this matters for EDC, CTMS, and clinical data infrastructure
Several of FDA's Appendix B questions (particularly #11–15 and #19–24) map directly onto capabilities that clinical trial systems already provide, which is worth flagging in any comment submission. The "clinical confirmation" spectrum FDA describes (retrospective evaluation on real patient inputs → shadow deployment → clinician adjudication → prospective study) is essentially a maturity ladder that EDC and CTMS platforms are built to support: audit-trailed data capture, source-data verification workflows, and query management could all serve as the backbone for generating this evidence, provided vendors build export/ingestion pathways that let sponsors route GenAI device inputs/outputs through existing 21 CFR Part 11–compliant systems rather than bespoke logging. Similarly, FDA's postmarket questions on re-benchmarking cadence and "triggering events" (Q19, Q22) are a natural fit for CTMS-style change-control and version-tracking logic — sponsors already track protocol amendments and system changes with timestamped approvals; the same discipline could extend to model-version and prompt/guardrail changes. A useful piece of feedback here is to recommend FDA explicitly reference interoperability with existing GxP data systems (EDC, CTMS, safety databases) as an acceptable substrate for postmarket monitoring evidence, rather than requiring parallel, device-specific monitoring infrastructure — this would reduce burden and align with sponsors' existing quality systems.
On Foundation Model Device Master Files (MAF)
The MAF proposal (Q25) is one of the more consequential asks in the paper, and it's where clinical-ops perspective is especially relevant: it essentially proposes treating foundation models the way sponsors already treat drug substance from a third-party manufacturer, via a confidential master file that downstream device sponsors can reference by authorization. The practical gap FDA flags but doesn't resolve is incentive: unlike drug substance suppliers, foundation model developers have no existing regulatory relationship with FDA and, as the paper itself concedes, "limited incentive to disclose safety-relevant information." Worthwhile feedback would push FDA toward pairing the voluntary MAF with something that creates pull, e.g., recognizing MAF submission as a factor that streamlines review timelines for downstream device sponsors (similar to how referencing a Type II Drug Master File can expedite ANDA review), and requiring update-notification commitments to be contractually flowed down from model developer to device sponsor to CTMS/EDC-level change logs, so that a foundation-model version change is traceable through the same system that already tracks protocol and software version history in a trial.
A proposal worth submitting: a "Cumulative Competency Record" for devices used across multiple trials
One gap in the discussion paper is worth raising directly in comments on Q11, Q12, and Q19. FDA asks how sponsors should combine benchmarking and clinical-confirmation evidence into a single performance estimate (Q12), and how re-benchmarking cadence and triggering events should be determined (Q19), but says little about what happens when a single GenAI-enabled device is deployed across multiple trials and RWE programs simultaneously, each generating its own slice of performance and subgroup data. Today, a sponsor evaluating a device for a new trial has no structured way to know how many other trials have already used it, under what conditions, or with what results, information a sponsor would otherwise have to reconstruct manually, if at all, via ClinicalTrials.gov.
The paper already contains the two building blocks for a fix, just not connected to each other. Section VII.A proposes a Foundation Model MAF, a confidential, FDA-held file capturing information about the underlying model, referenceable across submissions. Section V.D.2 separately floats independent third parties maintaining sequestered benchmarking datasets and issuing performance certifications. Neither addresses the accumulation of real deployment history for a specific device across the trials that actually use it.
A Cumulative Competency Record (CCR), effectively a device-level counterpart to the Foundation Model MAF — would close that gap: a living, FDA-referenceable record, keyed to a device identifier (analogous to an NCT number), that aggregates a device's benchmarking results, clinical confirmation evidence, and subgroup performance (R.2) across every trial and RWE program in which it has been used. A sponsor considering the device for trial #6 could reference the accumulated evidence from trials #1–5 rather than starting evaluation from zero, and FDA reviewers would gain visibility into a device's real-world track record that no single submission currently captures. Critically, this doesn't require new FDA authority to be useful: the CCR could feed directly into a device's Predetermined Change Control Plan (PCCP) (Section VI.C, Q23), accumulated cross-trial performance becoming the evidentiary basis for whether a device qualifies for streamlined re-benchmarking under its existing change-control plan, rather than requiring a new premarket submission each time. Framed this way, the proposal builds on two mechanisms FDA has already floated (MAF and PCCP) rather than asking for new infrastructure — which is likely to make it a more actionable comment than a standalone registry proposal.
A note on Real-Time Clinical Trials (RTCT)
Readers following FDA's parallel Real-Time Clinical Trials (RTCT) initiative launched in April 2026 as part of the agency's broader "Operation TrialBlazer" effort, with proof-of-concept trials from AstraZeneca and Amgen streaming safety signals and endpoints to FDA in near real time may wonder if the two initiatives connect. As of this discussion paper, they don't formally reference each other; RTCT is a separate FDA workstream focused on continuous data reporting during active trials, governed under existing GCP and IND frameworks (21 CFR Part 312), not a device-regulation pathway. That said, the thematic overlap is real and worth watching: both initiatives push toward continuous, streaming evidence generation rather than periodic snapshots, and both lean on the same underlying data infrastructure (EDC, CTMS, automated signal-detection platforms) to make that continuous evidence trustworthy and traceable. A sponsor building postmarket monitoring and periodic re-benchmarking pipelines for a GenAI-enabled device today would likely reuse much of the same real-time data plumbing that RTCT participants are building for continuous trial oversight, so it may be worth flagging in FDA comments that the two efforts should be designed for technical convergence, even if they remain separate regulatory pathways for now.
Comment deadline: October 19, 2026 — Docket FDA-2026-N-7874 Sources: FDA Discussion Paper | RAPS coverage
(https://www.fda.gov/media/194242/download , https://www.fda.gov/medical-devices/digital-health-center-excellence/considerations-regulation-generative-ai-enabled-medical-devices-discussion-paper-and-request)
Comments
Post a Comment