Resources
Can AI Pass the Doctor Test? FDA Considers Physician-Inspired Evaluation of GenAI Devices and Other Regulatory Measures
AI is entering healthcare with indomitable will. It promises to increase efficiency, reduce administrative burden (which healthcare providers so desperately desire), reduce error, and use and analyze data to drive better, more equitable, and more accessible health outcomes.
Federal agencies are embracing this technologic revolution and measuring the benefits AI can offer in the healthcare space. Their goal is to establish clear pathways for AI’s integration while creating regulations which can protect consumers.
The FDA has issued a discussion paper acknowledging that “GenAI-enabled devices have unique characteristics and behaviors that are distinct from traditional software and AI-enabled devices that FDA regulates.” These devices may accept open-ended inputs, perform multiple subtasks, produce variable outputs to similar inputs, and evolve over time. Many are built on AI models developed by third parties—such as major and well-known technology companies—whose models offer “varying levels of transparency into their training data, architecture, and evaluation methods, making it difficult to attribute specific behaviors and errors to the device or its underlying model.”
I. What GenAI Risks have the FDA identified?
The FDA is now seeking public comments on its proposed regulatory approach for generative-AI (“GenAI”) enabled medical devices. In announcing their request, they have identified several risks associated with the use of GenAI:
- Hallucinated and clinically incorrect output – information may be generated that sounds plausible but is incorrect, incomplete, or unsupported, leading to delayed diagnoses, inappropriate treatment, or patient harm.
- Overreliance by clinician or patient – undue trust may be placed in AI-generated outputs, especially when they are presented in a highly confident or persuasive conversational format making it difficult to identify erroneous recommendations.
- Variable performance – an AI model may perform well in testing but may fail when confronted with different patient populations, specialties, care settings, or real-world workflows. Output may also vary depending on prior prompts and context making performance less predictable and more difficult to validate.
- Increasing autonomy – emerging AI systems may independently plan tasks, make decisions, use tools, or initiate actions which increase the risk for potential error.
- Excessive conservatism– AI may generate overly cautious recommendations, leading to unnecessary testing, referrals, interventions, costs, and clinician burden.
- Performance degradation – AI-enabled devices may change over time through updates, retraining, or changes to underlying foundation models. A system that was safe and effective at launch may not remain that way indefinitely.
- Limited transparency – the method in which an AI reaches a conclusion may neither be transparent nor may the information it relied upon.
- Inadequate premarket testing – current testing approaches may not be sufficient to evaluate safety, reliability, and clinical competency before market authorization.
The FDA has proposed a two-axis framework for thinking about risk – type of X (horizontal axis) versus level of Y (vertical axis). The X-axis considers the “activity” performed by the GenAI-enabled device—ranging from providing non-directive information to directing action to taking autonomous action. The Y-axis considers the “consequences” or severity of harm from relying on an incorrect GenAI-enabled device output. Risk increases as a device’s output moves from lower-left (non-directive, low-consequence) to upper-right (autonomous action, severe consequences).

II. Where is the FDA seeking input?
The FDA has not proposed specific requirements on how to address the above risk areas. Rather, it is seeking stakeholder input on the regulatory framework that should govern GenAI-enabled medical devices and how that framework can appropriately mitigate the risks described above. In particular, the FDA is requesting feedback on:
- Risk classification. Whether existing risk-based approaches adequately capture the unique risks posed by GenAI and whether factors such as device autonomy, intended users, and the potential consequences of an incorrect output should affect regulatory scrutiny.
- Premarket review and evidence standards. What evidence manufacturers should be required to generate before marketing a GenAI-enabled device, and whether traditional validation approaches are sufficient to demonstrate safety and effectiveness.
- A “Competency-Based” Evaluation Inspired by Medical Training. Perhaps most notably, the FDA is considering a “competency-based approach” to evaluating GenAI-enabled devices inspired by how physicians are trained and credentialed. Just as medical students undergo board examinations and supervised clinical rotations before practicing independently, GenAI devices may be evaluated through a two-phase process:
- Device Benchmarking (analogous to board exams) – Evaluating whether the device demonstrates clinical knowledge, safety behavior, communication quality, and generalizability across a range of test conditions to support reasonable assurance that the device is safe and effective for its intended use. The FDA proposes benchmarking across multiple elements including, without limitation: safety-critical recognition and escalation; scope maintenance and boundary adherence; calibration and uncertainty communication; clinical knowledge and task fidelity; information gathering and analysis; communication quality; robustness and reliability; and subgroup performance.
- Clinical Confirmation (analogous to supervised practice) – The FDA recognizes that even with rigorous device benchmarking, a competency-based approach may not fully predict how a GenAI-enabled device will perform in real-world clinical settings because its interactions with users, workflows, and patient populations can introduce variables not captured through structured benchmarking. Therefore, the FDA proposes to also use a clinical confirmation approach to evaluate and verify that devices actually perform appropriately and as intended in real or clinically representative conditions. The FDA suggests this may not require a prospective clinical study in every case; alternative approaches could include retrospective evaluation on real patient inputs, “shadow deployment” where the device operates but does not affect care, standardized patient interactions, clinician adjudication of real cases, or prospective clinical studies depending on the device’s risk profile.
- Post-market monitoring. How manufacturers and regulators should monitor AI-enabled devices after deployment to identify performance degradation, unexpected behavior, safety issues, and other risks that may emerge over time.
- Model modifications and lifecycle management. How changes to a model after authorization should be regulated, particularly where systems are updated, retrained, or rely on third-party foundation models that may themselves evolve.
- Autonomous and agentic AI systems. Whether AI systems capable of independently planning tasks, using tools, or taking actions with limited human involvement warrant additional regulatory safeguards beyond those applied to more traditional software.
- Transparency and foundation model oversight. FDA is also seeking input on whether foundation model developers should voluntarily submit “Foundation Model Master Files” containing information about their models’ architecture, training data, known limitations, and safety constraints. These files would be held confidentially by the FDA and could be referenced by device manufacturers in their premarket submissions, enabling more consistent regulatory review of devices built on common foundation models.
III. Deadline for Feedback
The FDA is accepting stakeholder feedback that may inform the development of future policy governing GenAI-enabled medical devices through October 19, 2026. Public comments may be submitted through Regulations.gov under docket FDA-2026-N-7874.
While the FDA has not yet established new regulatory requirements, it has provided an important indication of the issues the agency believes will shape the next generation of medical device regulation. Manufacturers, healthcare providers, developers, and other stakeholders should closely monitor these developments, as the feedback received may influence how the FDA evaluates, authorizes, and oversees GenAI-enabled technologies in the years ahead.
Verrill’s Health Care & Life Sciences attorneys continue to monitor developments in the use and regulation of artificial intelligence in healthcare.
To learn more about these developments or to discuss how AI may affect your business, please contact Andrew A. Ferrer, your Verrill relationship partner, or another member of Verrill’s Health Care & Life Sciences or Artificial Intelligence & Emerging Technologies teams.
Sources:
- U.S. Food & Drug Administration, FDA Seeks Public Feedback to Inform Regulatory Approach for Generative AI-Enabled Medical Devices (Sept. 23, 2025), https://www.fda.gov/news-events/press-announcements/fda-seeks-public-feedback-inform-regulatory-approach-generative-ai-enabled-medical-devices.
- U.S. Food & Drug Administration, Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback (2025), https://www.fda.gov/media/194242/download.