Healthcare Policy and AI Governance
FDA maintains a periodically updated AI-enabled medical devices list. Most listed entries use the 510(k) pathway, but FDA cautions that the list is not comprehensive and depends on identifiable AI-related terms in public authorization records (FDA AI-Enabled Medical Devices).
Epic’s widely deployed sepsis prediction model had a hospitalization-level AUC of 0.63 in external validation, compared with an AUC of 0.76–0.83 in the developer’s internal documentation (Wong et al., 2021). At the evaluated alert threshold, sensitivity was 33% and positive predictive value was 12% (Wong et al., 2021). Regulatory frameworks built for static medical devices must also govern version-controlled software, planned modifications, and performance changes caused by data or workflow drift. The EU AI Act adds obligations for qualifying high-risk systems, while U.S. requirements depend on intended use, device status, authorization pathway, and other applicable law. Regulatory authorization, clinical validation, local utility, and liability are distinct questions.
This chapter prepares clinicians and health-system leaders to:
- Understand the evolving FDA regulatory framework for AI/ML-based medical devices
- Evaluate international regulatory approaches (EU AI Act, WHO guidelines)
- Recognize reimbursement challenges and evolving payment models for AI
- Assess institutional governance frameworks for safe AI deployment
- Navigate liability, accountability, and legal frameworks for medical AI
- Implement hospital-level AI governance policies
Introduction
Medicine operates within complex regulatory and policy frameworks: FDA device approvals, CMS reimbursement decisions, state medical board oversight, institutional protocols, and malpractice liability standards. These structures emerged over decades to protect patients from unsafe drugs, devices, and practices. They assume products are static: a drug approved in 2020 is chemically identical in 2025.
AI challenges this assumption fundamentally. Machine learning systems evolve through retraining on new data, algorithm updates, and performance drift as patient populations change. How should regulators approve systems that change continuously? Who is liable when AI errs: developers who built it, hospitals that deployed it, or physicians who followed its recommendations?
The stakes are high:
- Patient safety: Poorly regulated AI can harm thousands before problems are detected
- Innovation: Over-regulation may stifle beneficial AI development
- Equity: Biased regulatory frameworks may entrench disparities
- Legal liability: Unclear accountability creates defensive medicine
Part 1: External Validation and the Epic Sepsis Model
The Case Study
What was known before external validation: Epic’s proprietary sepsis prediction model was embedded in its EHR. Internal documentation reported a hospitalization-level AUC of 0.76–0.83, but independent validation had not been published (Wong et al., 2021).
What happened:
In 2021, Wong et al. published an external validation study in JAMA Internal Medicine testing the Epic sepsis model on 27,697 patients at Michigan Medicine (Wong et al., 2021):
- Sensitivity: 33%
- 67% of sepsis cases never triggered an alert at any point
- Positive predictive value: 12% (88% false positive rate among alerts)
- Area under the curve: 0.63 (poor discrimination)
Why institutional oversight failed to prevent this:
- Limited public evidence: The proprietary model was deployed widely before independent validation was published
- Poor transportability: Performance at Michigan Medicine was substantially worse than the developer’s internal results
- Alert burden: At the evaluated threshold, 18% of hospitalizations would have generated an alert, while positive predictive value was 12%
- Regulatory boundary: The Epic Sepsis Model was deployed as EHR decision support without FDA clearance
Lessons:
- Deployment ≠ clinical validation
- Retrospective studies can overstate utility when temporal leakage, selection bias, or confounding is not controlled
- External validation at independent institutions is essential
- Post-market surveillance is inadequate
Part 2: FDA Regulation of AI-Enabled Medical Devices
Regulatory Pathways
510(k) Clearance (Substantial Equivalence):
- Device is “substantially equivalent” to a predicate device already on market
- Premarket notification pathway whose review requirements depend on the device and submission; a 2024 cohort analysis reported a median review time of 151 days for the 510(k) devices it studied (Almarie et al., 2025)
Premarket Approval (PMA):
- Premarket review for Class III devices requiring valid scientific evidence that provides reasonable assurance of safety and effectiveness
- Reserved for devices that support or sustain human life, are of substantial importance in preventing impairment of human health, or present a potential unreasonable risk of illness or injury
De Novo Classification:
- New device type with no predicate
- Establishes new regulatory pathway for similar future devices
- Example: IDx-DR diabetic retinopathy screening received De Novo authorization in April 2018 (DEN180001, Class II) (FDA De Novo Decision Summary)
Current State
By the numbers:
- FDA maintains a periodically updated AI-enabled medical device list, but cautions that it is not comprehensive (FDA AI-Enabled Medical Devices)
- 2024 cohort: One published analysis identified 168 machine-learning-enabled Class II devices authorized in 2024, 94.6% through 510(k), with 74.4% categorized in radiology. These are study-defined counts, not a substitute for FDA’s current list (Almarie et al., 2025)
- Transparency gap: Only 23.7% of FDA-authorized AI/ML devices report dataset demographics in public summaries, and only 1.5% report Predetermined Change Control Plans (Khunte et al., npj Digital Medicine, 2025)
- Lifecycle gap: A 2026 analysis of 956 FDA-authorized radiology AI medical devices found that 429 (44.87%) had undergone software version updates, while most companies lacked a closed-loop risk management system connecting adverse events, recalls, and software updates. Software defects accounted for 124 recall reports (68.13%) (Li et al., 2026).
Predetermined Change Control Plans (PCCP)
Many authorized AI-enabled devices use fixed, version-controlled models. When a manufacturer anticipates later modifications, FDA’s PCCP framework can define specified changes and the methods used to develop, validate, and implement them (FDA PCCP Guidance, August 2025):
What PCCP allows:
- Manufacturer specifies anticipated changes (retraining, performance improvements)
- FDA reviews and may authorize the PCCP as part of a marketing submission
- A PCCP is controlled preauthorization for specified modifications, not general permission for uncontrolled self-updating
Components required:
- Description of modifications: Itemization of proposed changes with justifications
- Modification protocol: Methods for developing, validating, and implementing changes
- Impact assessment: Benefits, risks, and mitigations
Final guidance issued August 2025 provides FDA recommendations for Predetermined Change Control Plans (PCCPs) for AI-enabled device software functions.
### General Wellness Products: What Escapes FDA Oversight {#sec-general-wellness}
Not all health-related software and devices undergo FDA premarket review. The FDA’s General Wellness guidance (updated January 2026) explains when healthy-lifestyle software is not a device and when FDA does not intend to examine whether a low-risk general wellness product is a device or to enforce applicable device requirements (FDA General Wellness Guidance, 2026).
The two-factor test:
A product falls within the general wellness policy if it meets both criteria:
- Intended for general wellness use only: Claims relate to maintaining or encouraging a healthy lifestyle (weight management, physical fitness, relaxation, sleep management, mental acuity) OR relate healthy lifestyle choices to reducing risk of chronic diseases where this association is well-established
- Low risk: Not invasive, not implanted, does not involve technology posing safety risks without regulatory controls (lasers, radiation)
January 2026 update on physiologic sensing:
The updated guidance explicitly addresses non-invasive optical sensing for physiologic parameters, directly relevant to consumer wearables:
Products using optical sensing (photoplethysmography) to estimate blood pressure, oxygen saturation, blood glucose, or heart rate variability may qualify as general wellness products when outputs are intended solely for wellness uses, provided they:
- Are non-invasive and not implanted
- Are not intended for diagnosis, treatment, or management of disease
- Do not claim clinical equivalence to FDA-cleared devices
- Do not prompt specific clinical actions or medical management
- Do not include clinical thresholds or diagnostic alerts
- Have validated values if displaying physiologic measurements
What makes a product NOT a wellness device:
| Characteristic | Example | Regulatory Status |
|---|---|---|
| Disease diagnosis claims | “Detects atrial fibrillation” | May meet the device definition; verify classification and applicable premarket pathway |
| Treatment guidance | “Adjust insulin based on glucose reading” | May be regulated as a device function, depending on intended use and risk |
| Clinical equivalence claims | “Medical-grade blood pressure” | Outside the low-risk general-wellness framing; verify device status |
| Diagnostic thresholds/alerts | “Heart rate dangerous, seek care now” | Clinical action claims may place the function within device oversight |
| Invasive measurement | Microneedle glucose sensor | Outside the low-risk general-wellness policy; separate device requirements may apply |
Illustrative examples from FDA guidance:
| Product | General-wellness policy may apply | Device oversight may apply |
|---|---|---|
| Wrist-worn activity tracker with heart rate, sleep, blood pressure for “recovery assessment” | Yes (if validated values, no disease claims) | |
| Same device claiming to “monitor hypertension” | Yes, because the disease-management claim changes the intended use | |
| Pulse oximeter for “monitoring during hiking” | Yes | |
| Pulse oximeter for “detecting hypoxemia” | Yes, because this is a diagnostic claim | |
| App playing music for “relaxation and stress management” | Yes | |
| App claiming to “treat anxiety disorder” | Yes, because this is a treatment claim |
What this means for physicians:
Consumer wearables within the general wellness policy do not undergo FDA premarket review for those wellness functions. When patients present data from Oura, Whoop, Apple Watch (non-FDA features), or similar devices:
- Treat outputs as informational, not diagnostic. These devices may display blood pressure, SpO2, or glucose estimates without demonstrating clinical accuracy
- Validation status varies by feature. Authorization of one function does not extend to every feature on the same device
- Manufacturers remain responsible for displayed values. The guidance states that physiologic values should be validated, although FDA does not conduct premarket review of those values under this policy
- Marketing language matters. The same hardware can be a wellness product or medical device depending on how it’s marketed and what claims are made
January 2026: Clinical Decision Support Software final guidance
FDA issued its final Clinical Decision Support (CDS) software guidance in January 2026 (issued January 29, 2026). The guidance clarifies which CDS software functions are excluded from the definition of a device under section 520(o)(1)(E) (Non-Device CDS), provides examples of Non-Device CDS and device software functions, and notes that FDA digital health policies continue to apply to software functions that meet the definition of a device, including those intended for use by patients or caregivers (FDA CDS Guidance, January 2026; PDF; FDA CDS FAQs).
What this means for physicians:
- Do not treat “physician in the loop” as a proxy for Non-Device CDS status. Regulatory status depends on intended use, claims, and whether the software function meets the section 520(o)(1)(E) criteria.
- Software intended to support time-critical decisions does not satisfy the independent-review criterion for Non-Device CDS. Use in an emergency department alone does not determine status; intended use and the four statutory criteria control.
- Patient-facing functions remain a key boundary. The guidance notes that FDA digital health policies continue to apply to software functions that meet the definition of a device, including those intended for patients or caregivers.
Autonomous clinical AI and licensure-like oversight
A 2026 JAMA Perspective proposes a licensure framework for autonomous clinical AI, distinct from ordinary clinical decision support (Bergman et al., 2026).
The operational boundary for health systems is governance, not model architecture. If a tool is intended to act without per-case review, treat it like credentialing rather than procurement:
- Define a scope of practice and the conditions it does and does not cover
- Require supervised deployment evidence in the intended setting
- Set renewal criteria tied to patient outcomes and safety events
- Assign authority to restrict, suspend, or revoke operational use after errors
Challenges and Needed Reforms
| Problem | Evidence | Needed Reform |
|---|---|---|
| Independent validation can follow deployment | Epic sepsis model was deployed widely before independent external validation documented poor performance | Require evidence proportionate to risk, including external and prospective evaluation when the intended claim requires it |
| Inadequate post-market surveillance | FDA cautions that MAUDE reports cannot establish event rates or causation and may contain incomplete, inaccurate, or unverified information (FDA MAUDE limitations) | Require active post-market surveillance with AI-specific adverse event analysis |
| Generalizability not assessed | Performance can change across populations and settings | Require relevant subgroup analysis and local acceptance testing |
| Transparency vs. trade secrets | Users may lack information needed to judge fit and limitations | Require disclosure proportionate to clinical risk, including intended use, validation population, performance, and known limitations |
| Autonomous clinical AI does not map cleanly onto device clearance | JAMA Perspective proposes a licensure framework for autonomous clinical AI (Bergman et al., 2026) | Develop a licensure-like pathway for autonomous clinical agents, distinct from ordinary CDS and narrow SaMD clearance |
2025 Federal AI Policy Shift
The federal approach to AI regulation changed significantly in 2025. On January 20, 2025, Executive Order 14148 formally rescinded Executive Order 14110 (“Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence”) (Federal Register, 2025). Three days later, Executive Order 14179 (“Removing Barriers to American Leadership in Artificial Intelligence”) directed agencies to identify and rescind or revise actions inconsistent with sustaining and enhancing U.S. AI leadership (Federal Register, 2025).
The July 2025 America’s AI Action Plan recommends regulatory sandboxes where agencies including FDA could support rapid testing of AI tools with data and results sharing. On December 11, 2025, Executive Order 14365 (“Ensuring a National Policy Framework for Artificial Intelligence”) directed the Attorney General to establish a task force to challenge state AI laws deemed inconsistent with federal policy and called for federal preemption legislation. The order did not itself broadly preempt state AI laws (White House, December 2025).
What this means for physicians: Federal policy favors faster AI deployment and challenges some state-level requirements, while the FDA’s medical device framework remains in place. Institutional governance and independent clinical validation remain necessary where federal and state requirements are unsettled.
Early 2026: FDA Regulatory Framework in Flux
Four developments in 2026 affect how AI medical devices reach clinicians:
QMSR amended Part 820 (effective February 2, 2026). The FDA’s Quality Management System Regulation amended 21 CFR Part 820 by incorporating ISO 13485:2016 by reference and adding FDA-specific requirements. As of February 2, 2026, FDA no longer uses the Quality System Inspection Technique (QSIT) for device inspections and instead follows the inspection process in Compliance Program 7382.850 (FDA, 2026; FDA, 2026). Manufacturers subject to Part 820, including manufacturers of AI-enabled devices, must comply with the amended framework (FDA QMSR).
FDA proposed additional Class II 510(k) exemptions (February 6, 2026). A Federal Register notice identified Class II device types that FDA proposed to exempt from premarket notification, subject to limitations, and accepted comments through April 7, 2026. The notice stated that it was not FDA’s final determination (Docket FDA-2026-N-0232; Federal Register, 2026).
FDA denied an industry petition to exempt AI radiology devices from premarket review. In October 2025, Harrison.ai filed a citizen petition (Docket FDA-2025-P-5560) requesting partial exemption from 510(k) requirements for certain radiology computer-aided detection, diagnosis, and triage devices under specified conditions (Citizen Petition, October 2025). FDA denied the petition on April 1, 2026, and noted that predetermined change control plans can support controlled post-clearance modifications without a new marketing submission for each authorized change (FDA Final Response, April 2026).
FDA launches the TEMPO digital health pilot (announced December 5, 2025). FDA’s Center for Devices and Radiological Health (CDRH) announced the Technology-Enabled Meaningful Patient Outcomes (TEMPO) pilot in connection with CMS’s Advancing Chronic Care with Effective, Scalable Solutions (ACCESS) model. The ACCESS model began July 5, 2026 (CMS ACCESS Model). Participating manufacturers may request enforcement discretion for premarket authorization, investigational device exemption (IDE), and informed consent/IRB requirements while collecting real-world performance data. The pilot covers four clinical use areas: early cardio-kidney-metabolic conditions (hypertension, dyslipidemia, obesity, prediabetes), established cardio-kidney-metabolic disease (diabetes, CKD, atherosclerotic cardiovascular disease), chronic musculoskeletal pain, and behavioral health (depression or anxiety), with up to approximately 10 manufacturers per use area (FDA TEMPO FAQ; Federal Register, December 2025). FDA selected the Dexcom Glucose Health Program as the first participant on July 22, 2026, while noting that the proposed intended uses had not yet been evaluated by FDA (FDA, July 2026).
A 2026 npj Digital Medicine interview with Stephen Gilbert, Tinglong Dai, and Shantanu Nundy framed TEMPO and ACCESS as a combined oversight, evidence-generation, and reimbursement pathway, with early reimbursement paired to controlled health technology assessment and public reporting of outcomes (Gilbert & Dai, 2026).
What this means for physicians: The QMSR modernizes quality requirements but does not add AI-specific provisions. The TEMPO pilot points toward earlier real-world deployment with periods of enforcement discretion while manufacturers collect performance data. FDA’s denial of the Harrison.ai petition preserved premarket review for the radiology AI categories at issue and identified PCCPs as a mechanism for controlled post-clearance change (FDA Final Response, April 2026). Institutional validation, independent performance monitoring, and clinical governance remain necessary.
State Regulatory Sandboxes
While federal policy shifts toward deregulation, several states have created “regulatory sandboxes” for healthcare AI, enabling controlled testing of innovations that would otherwise face regulatory barriers. These programs provide temporary relief from specific state requirements while maintaining safety monitoring.
Utah AI Prescription Renewal Pilot
Utah’s Office of Artificial Intelligence Policy authorized a phased pilot with Doctronic for AI-supported renewal of existing prescriptions. The pilot remains in Phase 1, in which every renewal requires authorization by a licensed medical practitioner (Utah Office of Artificial Intelligence Policy).
How the pilot works:
| Component | Details |
|---|---|
| Scope | Renewal of existing prescriptions within an approved formulary; no new prescriptions or dose or frequency changes |
| Exclusions | Controlled substances and medications outside the approved formulary |
| Phase 1 | Every renewal requires licensed-practitioner authorization |
| Phase 2 threshold | Each medication group requires 250 fills and Office of Artificial Intelligence Policy approval before direct submission to a pharmacist |
| Oversight | Licensed professionals oversee implementation; pharmacists may escalate renewals to a licensed physician |
| Duration | October 2025–October 2026, with an option for renewal |
Rationale: Medication non-adherence costs an estimated $100-289 billion annually in avoidable U.S. healthcare spending (Cutler et al., BMJ Open, 2018). Approximately 78% of prescription activity involves routine refills rather than new prescriptions (Optum, 2017). Administrative delays in renewals contribute to gaps in medication adherence, particularly for chronic conditions requiring continuous therapy.
Professional response: The American Medical Association has expressed caution. AMA CEO Dr. John Whyte stated that “without physician input [AI] also poses serious risks to patients and physicians alike” (Becker’s Hospital Review, January 2026).
States with AI Regulatory Sandbox Programs:
| State | Status | Key Features |
|---|---|---|
| Utah | Operational (2024) | Office of AI Policy with regulatory mitigation authority; healthcare AI pilots including mental health (ElizaChat), dental (Dentacor), prescription renewals (Doctronic) |
| Texas | Operational (2026) | TRAIGA took effect January 1, 2026. Includes a 36-month regulatory sandbox and requires plain-language disclosure when AI is used in relation to healthcare services or treatment, with emergency exceptions (Texas HB 149 / TRAIGA) |
| Delaware | Framework development directed (2025) | H.J.R. 7 directed development of an agentic-AI sandbox framework; it did not establish an operational sandbox (Delaware H.J.R. 7) |
| Arizona | Operational (2019) | General regulatory sandbox expanded beyond financial products; not AI-specific (Arizona H.B. 2177) |
| Wyoming | Proposed or under study | General regulatory sandbox proposals; no enacted AI-specific program verified (Wyoming Legislature, 2025) |
Federal sandbox proposals: The SANDBOX Act (S.2750), introduced in September 2025, would direct the Office of Science and Technology Policy to create a federal regulatory sandbox program. A waiver could run for an initial two years with up to four two-year renewals, for a maximum of 10 years. The bill remains introduced (Congress.gov, S.2750).
New state-law development: The Texas Responsible Artificial Intelligence Governance Act (TRAIGA), effective January 1, 2026, requires that when an AI system is used in relation to healthcare service or treatment, the provider must disclose that use to the patient or personal representative no later than the date the service or treatment is first provided, or as soon as reasonably possible in an emergency. The statute requires the disclosure to be clear, conspicuous, written in plain language, and not designed with dark patterns (Texas HB 149 / TRAIGA, Sec. 552.051).
What this means for physicians: State sandboxes create variation in what AI systems can legally do across jurisdictions. Utah’s current pilot still requires licensed-practitioner authorization for every refill in Phase 1. In Texas, AI-supported workflows may create a disclosure obligation. Physicians practicing in sandbox states should understand the specific regulatory relief granted, the safety monitoring requirements, and their own liability exposure.
Part 3: International Regulatory Approaches
EU AI Act (2024)
The EU AI Act entered into force August 1, 2024 (Regulation (EU) 2024/1689).
Risk-based categorization:
| Risk Level | Requirements | Medical AI Examples |
|---|---|---|
| Unacceptable (banned) | Prohibited practices under Article 5 | Prohibited practices can occur in healthcare contexts |
| High risk | Strict obligations | AI that is a safety component of, or is itself, a regulated product requiring third-party conformity assessment |
| Limited risk | Transparency requirements | Some medical chatbots or symptom checkers, depending on intended use |
| Minimal risk | No specific obligations under the AI Act | Health-related AI that does not meet the prohibited, high-risk, or transparency-risk criteria |
High-risk medical AI requirements (Regulation (EU) 2024/1689):
- Risk management: Identify, analyze, and mitigate foreseeable risks
- Data governance: Apply quality criteria to training, validation, and test data
- Documentation and logging: Maintain technical documentation and records for traceability
- Human oversight: Design appropriate measures for people overseeing system use
- Accuracy, robustness, and cybersecurity: Meet performance requirements throughout the lifecycle
Compliance timeline: The high-risk requirements for AI embedded in regulated products, including qualifying medical devices, apply from August 2, 2028 (Regulation (EU) 2026/1744).
WHO Guidelines (2021, 2024)
WHO published Ethics and Governance of Artificial Intelligence for Health in June 2021 with six principles (WHO, 2021):
- Protect human autonomy: Patients and providers maintain decision-making authority
- Promote human well-being and safety: AI must benefit patients, minimize harm
- Ensure transparency and explainability: Stakeholders understand AI logic and limitations
- Foster responsibility and accountability: Clear assignment of responsibility when AI errs
- Ensure inclusiveness and equity: AI accessible to diverse populations, mitigate bias
- Promote responsive and sustainable AI: Long-term monitoring, adaptation to changing contexts
In 2024, WHO published additional guidance on large multi-modal models (WHO, 2024), addressing risks specific to generative AI in healthcare:
- Hallucinations: LMMs generate confident but false medical information
- Outdated training data: Models trained on historical data produce obsolete recommendations
- Bias amplification: Training data from high-income countries encodes perspectives that may not generalize globally
- Liability gaps: The AI value chain (developer → provider → deployer) creates uncertainty about accountability when harm occurs
WHO’s 2024 guidance proposes liability frameworks including presumption of causality (shifting burden of proof to deployers), strict liability considerations, and no-fault compensation funds (see Liability chapter).
Separate from AI-specific guidance, the WHO Global Strategy on Digital Health 2020-2025 was extended through 2027 by the Seventy-eighth World Health Assembly in May 2025 (WHO, May 2025). The strategy covers national digital health infrastructure, interoperability standards, and health data governance. Since its adoption, 129 countries have established national digital health strategies. A follow-up framework for 2028-2033 is under development.
Limitation: WHO guidelines are aspirational, not enforceable. Countries adopt them voluntarily.
Other Regions
| Region | Current official route | Practical point |
|---|---|---|
| Canada (Health Canada) | Premarket guidance for machine-learning-enabled medical devices | Addresses risk management, data, clinical validation, transparency, monitoring, and PCCPs |
| United Kingdom (MHRA) | Software and AI as a Medical Device Change Programme | Applies medical-device rules according to intended purpose while the future regime develops |
| Japan (PMDA) | Software as a Medical Device review resources | Provides SaMD review and consultation routes, including AI-based device work |
| China (NMPA) | Medical-device standards system | National standards include work on AI medical-device performance testing; product-specific status still requires the controlling record |
International Governance and Multilateral Coordination
National regulations alone cannot govern AI systems that operate across borders. A foundation model developed in the U.S., fine-tuned in the EU, and deployed in hospitals across 50 countries presents governance challenges no single jurisdiction can address.
The WHO 2024 LMM guidance emphasizes the need for international coordination (WHO, 2024):
Networked multilateralism: Effective AI governance requires coordination across UN agencies, international financial institutions, regional organizations, civil society, and the private sector. No single body has authority over the global AI ecosystem.
Inclusive rule-making: AI governance must be shaped by all countries, not only high-income nations and the technology companies headquartered there. Rules developed without low- and middle-income country input risk encoding biases that harm those populations.
Cross-border accountability: Companies developing foundation models must be accountable regardless of where they are incorporated. Current frameworks allow regulatory arbitrage: locating operations in permissive jurisdictions while selling globally.
Current challenges:
| Gap | Description |
|---|---|
| No international AI treaty | Unlike nuclear, chemical, or biological domains, no binding international agreement governs AI development or deployment |
| Voluntary commitments lack enforcement | Corporate pledges on AI safety (e.g., Frontier AI Forum) have no accountability mechanisms |
| Regulatory arbitrage | Companies can base operations in jurisdictions with minimal oversight |
| Fragmented standards | No harmonized requirements for safety testing, transparency, or post-market surveillance |
Emerging coordination mechanisms:
- UN High-Level Advisory Body on AI: Recommendations published September 2024, but non-binding
- G7 Hiroshima AI Process: Voluntary code of conduct for foundation model developers
- OECD AI Principles: Adopted by 46 countries, but no enforcement mechanism
- Bilateral agreements: U.S.-EU Trade and Technology Council addresses AI but lacks specificity on medical applications
What this means for physicians: AI systems you use may be developed, trained, and updated by entities outside any jurisdiction’s effective control. Institutional governance and vendor due diligence become critical when regulatory frameworks are fragmented.
Part 4: Reimbursement and Clinical Adoption
The Historical Pattern: Government Incentives Drive Adoption
Regulatory authorization is necessary for regulated devices but does not establish a payment pathway. Coverage, coding, payment, workflow fit, capital costs, and evidence all shape clinical adoption. Some health systems fund AI directly, while others depend on service-specific reimbursement. Two U.S. technology-adoption cycles illustrate how payment policy can influence uptake.
Meaningful Use and EHR Adoption
Before 2009, electronic health record adoption was minimal. HITECH made available an estimated $27 billion in Medicare and Medicaid EHR incentives. Eligible professionals could receive up to $44,000 through Medicare or $63,750 through Medicaid, and Medicare payment reductions for noncompliance began in 2015 (CMS EHR overview).
The results were unambiguous. Annual EHR adoption rates among eligible hospitals increased from 3.2% pre-HITECH (2008-2010) to 14.2% post-HITECH (2011-2015), a difference-in-differences of 7.9 percentage points compared to ineligible hospitals (Adler-Milstein & Jha, Health Affairs, 2017). By 2017, 86% of office-based physicians and 96% of non-federal acute care hospitals had adopted EHRs (AHA News, 2017).
Telehealth and CMS Reimbursement
Telehealth followed the same pattern. Before COVID-19, Medicare telehealth coverage was restricted to rural areas for specific services. During the public health emergency, Congress and CMS waived geographic restrictions and expanded covered telehealth. Current policy combines permanent changes with temporary extensions through December 31, 2027: permanent provisions include home-based behavioral health telehealth and removal of frequency limits for subsequent inpatient and nursing facility visits and critical care consultations, while several non-behavioral flexibilities remain time-limited (HHS Telehealth Policy, 2026; CMS CY 2026 Physician Fee Schedule).
Why Public Payment Policy Matters
Medicare fees can influence commercial prices. Research found that a $1.00 increase in Medicare fees was associated with a $1.16 increase in corresponding private prices (Clemens & Gottlieb, J Polit Econ, 2017). This price relationship does not establish that commercial coverage decisions automatically follow Medicare.
Coverage decisions are distributed across Medicare statutes and regulations, national and local coverage processes, commercial payers, and health-system budgets. Medicare, which covers about 70 million people, can provide an important market signal, but a Medicare fee association does not prove that commercial payers will adopt the same coverage policy (CMS enrollment data).
The Capital and Talent Argument
Without clear reimbursement pathways, capital and talent flow elsewhere. Healthcare AI attracted $18 billion in venture investment in 2025, accounting for 46% of all healthcare investment (Silicon Valley Bank, 2026). Yet AI-enabled startups captured 62% of VC dollars because investors prioritize companies targeting operational efficiency over clinical decision support (Rock Health, 2025). The reason: operational AI has clearer revenue models. Clinical AI faces reimbursement uncertainty.
These investment figures describe financing patterns, not proof that one business model improves clinical outcomes. They are best used to frame reimbursement incentives rather than to forecast where individual engineers or companies will work.
Current Landscape: Adoption Remains Nascent
A 2024 analysis of 11 billion CPT claims found that only two AI applications exceeded 10,000 claims during 2018–2023: coronary artery disease assessment (67,306 claims via the then-current CPT 0501T–0504T codes) and diabetic retinopathy screening (15,097 claims via CPT 92229) (Wu et al., NEJM AI, 2024). Despite the FDA’s large and growing list of AI-enabled medical devices, clinical utilization remained concentrated in affluent, metropolitan, academic medical centers.
| AI Category | Reimbursement Status | CPT Code(s) |
|---|---|---|
| Diabetic retinopathy autonomous analysis | CPT 92229; Medicare payment is contractor-priced and coverage depends on applicable requirements (CMS, 2020) | No fixed national rate |
| Coronary FFR-CT | Medicare contractor-priced (CMS 2026 PFS dataset) | CPT 75580 (AMA, 2024) |
| Radiology AI | Often bundled into an existing service rather than separately paid | Service-specific |
| Clinical decision support | Payment varies by service and setting; many tools are health-system funded | Service-specific or none |
| Ambient documentation | Generally funded by clinicians or health systems | No dedicated national Medicare payment |
Deployment and Reimbursement Case: IDx-DR Diabetic Retinopathy Screening
IDx-DR (now LumineticsCore) illustrates how autonomous analysis can obtain a dedicated service code (FDA De Novo, 2018; CMS, 2020):
| Milestone | Details |
|---|---|
| FDA authorization | April 2018, De Novo pathway (first autonomous AI diagnostic) |
| CPT code | 92229 established 2021 (imaging with AI interpretation without physician review) |
| Medicare PFS status | Contractor-priced; coverage and payment vary by Medicare contractor and locality |
What the pathway illustrates:
- A defined intended use and autonomous workflow supported a dedicated service descriptor
- Primary-care sites can perform the authorized analysis without a specialist interpreting each image
- A prospective multicenter study supported De Novo authorization (Abràmoff et al., 2018)
- Coverage and payment remain payer- and locality-specific; authorization and a CPT code do not by themselves prove cost-effectiveness or universal access gains
The Policy Gap: What Clinical AI Needs
Medicare has no explicit benefit category for prescription digital therapeutics. The Access to Prescription Digital Therapeutics Act of 2025 (H.R.3288/S.1702) remains introduced and would direct Medicare and Medicaid coverage of qualifying prescription digital therapeutics (Congress.gov, H.R.3288). Some emerging AI-enabled services use Category III CPT codes; others are bundled into existing services, contractor-priced, or health-system funded.
The AMA’s 2026 CPT code set added codes for several augmented-intelligence services, effective January 1, 2026 (AMA, 2025). Inclusion in the CPT code set does not itself establish Medicare coverage or payment.
Emerging Payment Models
Fee-for-service does not incentivize AI adoption. Value-based models align incentives:
Value-based care contracts: Providers share risk with payers. AI that reduces hospitalizations and complications directly benefits providers financially.
Bundled payments: Single payment for entire episode of care. AI costs included in bundle; providers incentivized to use cost-effective AI that improves outcomes without separate line-item billing.
Outcomes-based contracts with vendors: Hospital pays vendor based on AI performance, not upfront license. Aligns incentives to reduce false positives and demonstrate clinical value.
The Path Forward
Clinical AI adoption will depend partly on whether payment rewards measured value rather than software acquisition alone. Policy options include:
- Dedicated CPT codes with adequate valuation for validated clinical AI applications
- Medicare coverage decisions that establish precedent for private payers
- MACRA/APM integration when an AI-supported workflow produces validated quality improvement
- Evidence-linked incentives that avoid penalizing nonuse before patient benefit, workflow safety, and equitable access are established
The technology and validation methods exist, but payment is only one missing component. A payment signal should follow evidence of patient or operational value, not substitute for it. Implementation also requires local validation, technical integration, clinician training, monitoring, and a safe fallback.
Part 5: Operational AI and Healthcare Economics
The Untapped Opportunity
While regulatory and reimbursement discussions focus on clinical decision support, operational and administrative activities consume a larger share of healthcare spending. Workforce staffing, care coordination, billing, claims processing, scheduling, and customer service contributed an estimated $950 billion in U.S. healthcare costs in 2019 (Sahni et al., McKinsey, 2021). This represents approximately 25% of total healthcare spending.
The National Bureau of Economic Research estimates that wider AI adoption could generate savings of 5–10% of U.S. healthcare spending, approximately $200–360 billion annually in 2019 dollars (Sahni et al., NBER Working Paper 30857, 2023). These estimates focus on AI-enabled use cases using current technology, attainable within five years, that do not sacrifice quality or access.
Why operational AI attracts capital: Rock Health reported that AI-enabled companies captured 62% of U.S. digital-health venture funding in 2025 and described administrative and workflow use cases as prominent investment targets (Rock Health, 2025). This financing pattern is not evidence that a particular operational tool produces savings. Each use case still requires a prespecified economic evaluation that counts implementation, monitoring, error correction, and displaced work.
The Productivity Paradox
Paradoxically, new clinical technologies have historically increased overall healthcare spending (Brynjolfsson et al., AEJ: Macroeconomics, 2021). This “productivity J-curve” describes a pattern where general purpose technologies initially reduce productivity before generating gains. The explanation: technology alone does not reduce costs. Organizations must redesign workflows, structures, and culture around the technology.
The implication for health systems: Rather than retrofitting AI into existing clinical workflows, effective implementation requires redesigning processes around AI capabilities. For example, follow-up appointment frequency is typically left to individual physician preference with high variability and little evidence guiding these decisions. AI-based risk stratification could reallocate appointment frequency based on patient need, substantially increasing capacity and reducing wait times. One study found that reducing follow-up frequency by a single visit per year could save $1.9 billion nationally (Ganguli et al., JAMA, 2015).
Such redesign requires institutional willingness to change practice patterns, not merely add AI to existing processes.
Cross-Industry Lessons
Other industries have adopted operational AI with documented returns on investment (Wong et al., npj Health Syst, 2026):
| Sector | Application | Healthcare Parallel |
|---|---|---|
| Retail | Inventory forecasting, demand prediction | Hospital capacity forecasting, supply chain optimization |
| Aviation | Predictive maintenance, weather disruption modeling | Biomedical equipment maintenance, OR scheduling |
| Logistics | Route optimization | Emergency response, patient transport |
| Financial services | Customer advisory chatbots, query resolution | Patient communication, prior authorization |
Key insight: The cross-industry examples argue for pairing technology investment with workflow redesign. Wong and colleagues cite UPS’s reported $300–400 million annual savings from route optimization as an illustrative company estimate, not a healthcare-effectiveness result (Wong et al., 2026).
Barriers to Operational AI in Healthcare
Several factors explain why healthcare has lagged other sectors in operational AI adoption:
- Fragmented data: Healthcare data is siloed across EHRs, claims systems, and departmental applications. AI requires integrated data access.
- Risk tolerance: Aviation and finance are safety-critical, yet healthcare is more risk-averse about algorithmic decision-making.
- Regulatory boundaries: Operational software may fall outside FDA device oversight when it lacks a medical-device intended use, but privacy, cybersecurity, employment, consumer-protection, and other requirements may still apply.
- Workforce concerns: Staff may view operational AI as threatening rather than enabling.
- Misaligned incentives: Fee-for-service rewards volume, not efficiency. Value-based contracts better align incentives with operational improvement.
The Learning Health System Framework
Effective AI integration requires continuous evaluation, not one-time deployment. The learning health system model provides a framework for pairing operational goals with evidence generation (IOM, 2012):
- Identify operational gap (e.g., OR utilization, scheduling efficiency, documentation burden)
- Deploy AI intervention with prospective measurement plan
- Evaluate outcomes against pre-specified metrics
- Iterate or terminate based on evidence
This approach addresses a critical gap: a 2024 study found that only 61% of U.S. hospitals performed any local performance evaluation of AI models prior to deployment (Nong et al., Health Affairs, 2025). Many health systems lack the expertise or infrastructure to validate AI performance or assess investment value.
National collaboratives can help: The Coalition for Health AI (CHAI), the Health AI Partnership, and the AMA’s Center for Digital Health and AI enable resource-limited health systems to leverage peer expertise, troubleshoot common problems, and disseminate findings on operational AI tools.
Strategic Implications for Health Systems
| Action | Rationale |
|---|---|
| Tie AI initiatives to measurable value | Avoid “productivity paradox” by requiring ROI demonstration before scaling |
| Redesign workflows around AI | Retrofitting AI into existing processes yields minimal benefit |
| Integrate AI operations with research | Learning health system model generates evidence while improving operations |
| Build distributed AI literacy | Frontline staff must understand AI capabilities to identify opportunities |
| Consider bounded operational use cases | Regulatory complexity may be lower than for device clinical decision support, but ROI and safety must still be measured |
Part 6: Institutional Governance
Why Hospital-Level Governance Matters
FDA authorization does not establish that AI will work at a particular hospital:
- Different patient population, workflows, EHR
- Local cost-effectiveness depends on implementation costs, workflow, volume, alternatives, and measured benefit
- Physicians will use AI appropriately
- Patients will not be harmed
The “Tip of the Iceberg” Problem: Hospital governance often focuses only on enterprise-procured, EHR-integrated AI tools. However, as Ötleş and colleagues argue in NEJM AI (2026), “Health Systems Govern Only the Tip of the AI Iceberg” (Ötleş et al., 2026). Banning generative AI or heavily restricting enterprise tools does not eliminate risk; it merely makes risk invisible. With large numbers of clinicians already using “shadow AI” on personal devices for drafting notes or reviewing cases, effective governance must acknowledge and manage this unseen majority of AI use rather than merely policing the visible tip of officially procured systems.
Institutional governance fills gaps left by regulation and addresses the reality of clinical practice.
Essential Governance Components
1. Clinical AI Governance Committee
Minimum composition:
- Chair: CMIO or CMO
- Physicians from specialties using AI
- Chief Nursing Officer representative
- CIO or IT director
- Legal counsel with medical malpractice and AI expertise
- Chief Quality/Patient Safety Officer
- Health equity lead
- Bioethicist
- Patient advocate
Responsibilities: Pre-procurement review, pilot approval, deployment oversight, adverse event investigation, policy development, bias auditing.
2. Validation Before Deployment
Do not assume vendor validation generalizes to your hospital.
| Phase | Duration | Purpose |
|---|---|---|
| Silent mode | Duration justified by event volume and uncertainty | AI generates outputs that do not affect care; evaluate pipeline behavior and in-setting performance |
| Human-factors or shadow evaluation | Duration justified by the questions being tested | Evaluate interface, workflow, and user behavior without presenting unvalidated output as clinical truth |
| Active pilot | Prespecified sample, duration, and stopping rules | Limited deployment with a comparator or credible baseline and clinically justified success criteria |
A scoping review of 75 silent evaluations found no standardized methodology and wide variance in what teams actually test during this phase; most measure only AUROC while neglecting subgroup equity, workflow integration, and human factors (Tikhomirov et al., 2026). Silent mode should evaluate all success criteria below, not just technical stability.
Success criteria should include:
- Technical: Prespecified discrimination, calibration, error, latency, and data-quality measures appropriate to intended use
- Clinical: Prespecified workflow or patient endpoints measured against a relevant comparator when benefit is claimed
- User: Physician satisfaction, response rate
- Safety: Defined adverse-event surveillance, escalation, rollback, and learning procedures
- Equity: Subgroup measures selected for the intended population, with uncertainty and clinically justified action thresholds
3. Bias Monitoring
Audit cadence should reflect clinical risk, use volume, update frequency, and the speed at which harm could occur. Relevant analyses may include:
- Race/ethnicity
- Age
- Sex
- Insurance status
- Language
No universal percentage defines unacceptable disparity. Governance teams should prespecify clinically meaningful thresholds, examine uncertainty and sample size, investigate causes, and restrict or deactivate the system when a material unresolved risk remains.
4. Vendor Contracts
Contract review should address, with qualified legal, privacy, security, procurement, and clinical input:
- Hospital retains ownership of patient data
- Performance representations, audit access, change notification, service levels, and termination rights tied to justified thresholds
- Disclosure of training data demographics, validation studies, limitations
- Allocation of responsibility, insurance, indemnification, and remedies appropriate to the jurisdiction and bargaining context
- A HIPAA Business Associate Agreement when the vendor is acting as a business associate, plus any other applicable data-processing terms
5. Value Alignment Assessment
Beyond technical performance and bias, institutions should assess what values AI systems embed. The RAISE consortium’s “Values In the Model” (VIM) framework proposes that AI systems disclose how they navigate value-laden trade-offs: intervention vs. conservative management, patient autonomy vs. paternalism, individual benefit vs. resource constraints (Goldberg et al., NEJM AI, 2026).
When evaluating AI systems, governance committees should ask:
- What optimization target was this system trained on?
- How does it handle scenarios where reasonable experts disagree?
- Does behavior differ between fee-for-service and capitated contexts?
See Ethics chapter: Value Alignment Frameworks for detailed guidance on assessing embedded values.
Enterprise AI Lifecycle Frameworks
Beyond committee composition and validation protocols, institutions need structured frameworks for how AI solutions progress from concept to deployment to monitoring. Stanford Medicine’s experience provides a model.
RAIL (Responsible AI Lifecycle):
Stanford Health Care established the Responsible AI Lifecycle (RAIL) framework in 2023 to codify institutional workflows for AI solution development (Shah et al., 2026):
| Stage | Key Activities |
|---|---|
| Proposal | Use case definition, risk tiering, stakeholder alignment |
| Development | Model building, integration with clinical data, prompt engineering |
| Validation | FURM assessment (see below), truth set curation, benchmark testing |
| Pilot | Controlled deployment with defined success criteria |
| Monitoring | System integrity, performance, and impact tracking |
FURM (Fair Useful Reliable Models):
The FURM framework specifies required assessments before AI deployment:
- Fair: Performance tested across demographic subgroups; disparities documented and mitigated
- Useful: Clear clinical or operational benefit demonstrated; workflows redesigned for integration
- Reliable: Consistent performance across settings; failure modes identified and documented
Why structured frameworks matter:
Most health systems lack the expertise or infrastructure to validate AI performance or assess investment value. A 2024 study found that only 61% of U.S. hospitals performed any local performance evaluation of AI models prior to deployment (Nong et al., Health Affairs, 2025). Structured frameworks convert ad hoc adoption decisions into systematic processes.
Implementation lesson from Stanford ChatEHR:
Stanford’s approach required embedding data science teams within IT organizations, providing direct access to personnel maintaining network security, EHR integrations, and cloud resources. This integration enabled what began as a sandbox (April 2023) to scale to health-system-wide deployment (September 2025) in approximately 2.5 years. The “build-from-within” strategy provides institutional agency: model-agnostic infrastructure that matches clinical tasks to appropriate LLMs, institutional data governance, and custom monitoring aligned to organizational priorities.
For institutions without Stanford’s resources:
National collaboratives provide pathways for resource-limited health systems:
- Coalition for Health AI (CHAI): Responsible AI Guide (RAIG) with developer/implementer accountability frameworks (see CHAI section)
- Health AI Partnership: Peer expertise sharing and troubleshooting
- AMA Center for Digital Health and AI: Policy guidance and educational resources
ISO/IEC 42001: A Management-System Overlay
ISO/IEC 42001:2023 specifies requirements for organizations to establish, implement, maintain, and continually improve an artificial intelligence management system (AIMS) (ISO, 2023). For a health system, an AIMS can provide the organization-level structure for policies, objectives, documented risk treatment, accountability, review, and continual improvement across AI use cases.
The management system should coordinate, rather than replace, use-specific clinical controls. Each clinical workflow still needs a defined intended use, accountable clinical owner, local evidence, change-control path, supplier and data-dependency review, incident escalation route, and post-deployment monitoring plan. An AIMS does not establish clinical validity, local workflow fit, patient benefit, or FDA authorization.
NIST’s AI Risk Management Framework can complement this layer through its GOVERN, MAP, MEASURE, and MANAGE functions. It is voluntary and use-case agnostic, so health systems must translate it into clinical and operational controls rather than treat it as a compliance checklist (Tabassi, 2023).
Part 7: Liability and Legal Frameworks
Hypothetical Liability Scenarios
The following scenarios are decision exercises, not reports of actual cases or predictions of a court outcome. Liability is fact- and jurisdiction-specific. It may depend on professional negligence law, product design and warnings, contracts, institutional governance, documentation, and the conduct of each participant. FDA authorization does not decide civil liability (Mello & Guha, 2024).
Scenario 1: Physician follows AI recommendation, patient harmed
Hypothetical: Radiology AI characterizes a lung nodule as low risk. The radiologist concurs. Six months later, the nodule is diagnosed as cancer.
- Questions for review: Was the tool used within its intended use? What information was available to the radiologist? Did the institution perform appropriate acceptance testing and monitoring? Were material limitations communicated? Did the clinician’s interpretation meet the applicable standard of care?
- Practice lesson: AI output does not replace the clinician’s own assessment when clinician review is part of the workflow. Documentation should reflect the clinically material reasoning, not merely “AI said low risk.”
Scenario 2: Physician overrides AI, patient harmed
Hypothetical: A sepsis model produces a high-risk alert. The physician evaluates the patient, records the relevant findings and rationale, and discharges the patient. The patient later returns in septic shock.
- Questions for review: Was discharge reasonable given all available evidence? Was the model reliable in this population? Did alert presentation create automation bias or alert fatigue? Were follow-up and return precautions appropriate?
- Practice lesson: A documented override is evidence of reasoning, not automatic immunity. The legal question remains whether the overall care met the applicable standard.
Scenario 3: Systematic AI error harms multiple patients
Hypothetical: An ECG algorithm systematically underestimates the QT interval. Multiple patients receive QT-prolonging medications and develop arrhythmias.
- Questions for review: Did the defect originate in design, labeling, integration, data transformation, an update, or local workflow? Were safety signals detectable? Did the manufacturer and institution respond appropriately? Which contracts and jurisdictional doctrines apply?
- Practice lesson: Responsibility may be distributed across developers, manufacturers, deployers, institutions, and clinicians. A chapter cannot predict allocation without the controlling facts and law.
Unsettled Legal Questions
- Black-box algorithms: How to prove negligence when AI logic is inscrutable?
- Continuously learning AI: Who is liable for harms from updated algorithms?
- Off-label AI use: Physician uses AI outside approved indications
- Training data bias: AI systematically harms certain demographic groups
Standard of Care and Nonuse of AI
There is no general rule that use, override, or nonuse of AI is automatically negligent. Professional standards of care evolve through evidence, customary practice, guidelines, resources, patient circumstances, and jurisdiction-specific law. Practice parameters also commonly state that they are not intended to establish a legal standard of care.
The strongest evidence for a tool may make its availability relevant to a future case, but speed or diagnostic-performance evidence alone does not establish that every facility must deploy it. Outcome evidence, patient selection, implementation capacity, alternatives, and professional guidance remain relevant.
Examples to monitor without predicting negligence:
| Application | Evidence question | Policy question |
|---|---|---|
| Large-vessel-occlusion workflow software | Does the evidence establish faster workflow, improved treatment, or patient benefit for the specific product and setting? | Do specialty guidance, local resources, and alternatives support adoption? |
| Intracranial-hemorrhage triage | Is evidence product-specific, externally validated, and applicable to the local image mix and workflow? | Are there safe fallback and monitoring processes? |
| Autonomous diabetic-retinopathy analysis | Does the authorized population match the patients being screened, and is the full referral pathway available? | Do coverage, staffing, and follow-up support equitable implementation? |
| Sepsis prediction | Does prospective comparative evidence show benefit without unacceptable alert burden or delayed care? | Has the institution validated the actual model version and threshold? |
A plausible legal argument is not a decided legal rule. Institutions should follow evidence and professional guidance without presenting speculative litigation narratives as settled law.
Governance recommendations:
- Document evidence review and rationale for adoption, restriction, or nonadoption of consequential tools
- Monitor current specialty guidance: Link directly to the controlling society document and record its effective status
- Use accountable governance: Multidisciplinary review improves consistency but does not automatically distribute or eliminate liability
- Reassess material changes: New evidence, model versions, safety signals, or workflow changes can alter the risk-benefit assessment
Risk Reduction for Physicians
- Document clinically material information: Record relevant AI recommendations and the reasoning for following or overriding them when they affect care; indiscriminate documentation can obscure rather than clarify decisions
- Understand AI limitations: Know validation populations, sensitivity/specificity, failure modes
- Maintain clinical independence: AI is decision support, not decision maker
- Use disclosure or consent when required or appropriate: Applicable law, institutional policy, the role of AI, and the decision’s stakes determine what patients should be told
- Report AI errors: If AI makes systematic errors, report to Quality/Safety and AI governance
Part 8: Guidance from Professional and Standards Bodies
AMA Principles for Augmented Intelligence (2018)
The American Medical Association uses the term “augmented intelligence” to emphasize the supporting role of AI in clinical practice (AMA AI Principles, 2018).
Six principles:
- AI should augment, not replace, the physician-patient relationship
- AI must be developed and deployed with transparency
- AI must meet rigorous standards of effectiveness
- AI must mitigate bias and promote health equity
- AI must protect patient privacy and data security
- Physicians must be educated on AI
WHO Framework (2021)
Six principles: protect human autonomy, promote well-being and safety, ensure transparency, foster accountability, ensure equity, promote sustainability (WHO, 2021).
Key Professional Society Positions
| Society | Current source and scope |
|---|---|
| American College of Radiology and Society for Imaging Informatics in Medicine | The ACR-SIIM Practice Parameter for Imaging AI, approved in May 2026 and scheduled to take effect October 1, 2026, addresses governance, inventory, acceptance testing, monitoring, privacy, and quality improvement. ARCH-AI is a recognition program; AI-LAB is an education and validation framework, not accreditation. |
| American Heart Association | AHA guidance addresses value assessment for cardiovascular imaging and risk-proportionate evaluation across predeployment, implementation, and postdeployment phases (Hanneman et al., 2024; Jain et al., 2025). |
| College of American Pathologists | CAP’s 2026 interpretive-diagnostic-error guideline states that fixed standards for AI tools in anatomic pathology have not yet been established; the guideline does not create a universal rule that every AI-flagged case must be reviewed (CAP, 2026). |
Coalition for Health AI (CHAI) and Joint Commission Partnership
The Coalition for Health AI (CHAI) and Joint Commission announced their partnership in June 2025 and released initial guidance in September 2025 (Joint Commission, 2025). CHAI released governance playbooks in May 2026 (CHAI, 2026), followed by Joint Commission’s organizational certification in June 2026.
Why this matters: In June 2026, Joint Commission launched its voluntary Responsible Use of AI in Healthcare certification for U.S. healthcare organizations and systems. The certification evaluates organizational governance and safeguards; it does not certify individual AI products or tools (Joint Commission, 2026).
Use-Case-Specific Work Groups:
CHAI’s approach differs from generic AI principles by providing guidance tailored to specific clinical applications (CHAI Use Cases):
| Use Case | Focus |
|---|---|
| Clinical Decision Support (LLM + RAG) | Scope definition, escalation rules for when AI defers to humans, continuous monitoring, evidence traceability |
| EHR Information Retrieval | Grounding retrieved information, verification in real-world contexts, handling fragmented patient data |
| Prior Authorization Criteria Matching | Explainability of match/non-match decisions, human review triggers, preventing “denial drift” |
| Direct-to-Consumer Health Chatbots | Accessibility (5th-6th grade reading level, multilingual), error handling, authoritative source grounding, citations |
Developer vs. Implementer Accountability:
The CHAI Responsible AI Guide (RAIG) distinguishes between “Developer Teams” (data scientists, engineers who build AI solutions) and “Implementer Teams” (providers, IT staff, leadership who deploy them). Each stage of the AI lifecycle specifies which team bears primary responsibility (CHAI RAIG):
| Stage | Developer Responsibility | Implementer Responsibility |
|---|---|---|
| Define Problem & Plan | Collaborate on technical feasibility | Define business requirements, clinical context |
| Design | Model architecture, training approach | Workflow integration design |
| Engineer | Build, train, validate solution | Provide real-world data, clinical input |
| Assess | Performance metrics, bias testing | Local validation, population fit assessment |
| Pilot | Technical support, iteration | Controlled deployment, clinician feedback |
| Deploy & Monitor | Ongoing maintenance, updates | Adverse event tracking, governance reporting |
Governance Structure Recommendations:
CHAI guidance emphasizes:
- Written AI policies: Establish explicit governance with technically experienced leadership
- Transparency to patients: Disclosures and educational tools about AI use
- Data protection: Minimum necessary data principles, audit rights in vendor agreements
- Quality monitoring: Regular validation, performance dashboards
- Bias assessment: Audit whether AI was developed with datasets representative of served populations
- Blinded reporting: Cross-institutional learning from AI-related events
Limitations to Acknowledge:
CHAI guidance is process-oriented and generally does not impose universal numerical thresholds. This allows each organization to calibrate controls to risk, but it also means implementers must prespecify and justify questions such as:
- What performance floor triggers intervention?
- What retrieval accuracy makes EHR summarization safe?
- How many false denials cross from efficiency to patient harm?
The absence of one universal cutoff is appropriate because event prevalence, clinical consequences, and uncertainty vary by use case. The operational gap appears when an institution invokes continuous monitoring without defining its own evidence-based triggers, owners, and stop rules.
Resources:
Questions Clinicians Ask
What changed in FDA’s January 2026 CDS guidance?
FDA did not broadly exempt diagnostic AI. The January 2026 guidance clarified the four statutory criteria for non-device CDS, emphasized that clinicians must be able to independently review the basis of recommendations, and pointed developers to FDA’s other digital health policies when a function falls outside that exemption.
What is the EU AI Act’s impact on medical AI?
The EU AI Act treats an AI system as high-risk when it is a safety component of, or is itself, a regulated product that requires third-party conformity assessment. The high-risk requirements for AI embedded in regulated products, including qualifying medical devices, apply from August 2, 2028.
Why has not clinical AI been widely adopted despite FDA clearance?
Payment policy materially affects adoption. EHR use increased after Medicare and Medicaid incentive programs, and Medicare telehealth policy expanded access. A 2018–2023 claims analysis found that only two AI applications exceeded 10,000 CPT claims.
Is AI reimbursed by Medicare?
Most clinical AI lacks a dedicated payment pathway. CPT 92229 describes retinal imaging with autonomous point-of-care analysis, but a CPT code does not itself establish universal Medicare coverage or a fixed national payment rate.
What happened with Epic’s sepsis prediction model?
External validation of the Epic Sepsis Model found a hospitalization-level AUC of 0.63 and, at the evaluated alert threshold, 33% sensitivity and 12% positive predictive value. The model was not FDA-cleared.
What is the Utah AI prescribing pilot?
Utah authorized a phased pilot for AI-supported renewal of existing prescriptions. The pilot remains in Phase 1, in which every renewal requires authorization by a licensed medical practitioner; progression by medication group requires 250 fills and Office of Artificial Intelligence Policy approval.
What is CHAI and why does it matter for healthcare AI?
The Coalition for Health AI (CHAI) and Joint Commission released initial governance guidance in 2025. In June 2026, Joint Commission launched its voluntary Responsible Use of AI in Healthcare certification for healthcare organizations; it does not certify individual AI products.
How much could AI save in U.S. healthcare costs?
NBER estimates $200–360 billion annually. Administrative activities consume an estimated $950 billion yearly. Achieving savings requires redesigning workflows around AI, not retrofitting AI into existing processes.
Conclusion
AI regulation and policy are evolving rapidly. Many authorized systems are version-controlled rather than continuously learning in production, yet they still create lifecycle questions about planned modifications, drift, integration, and monitoring. Challenges include variable evidence standards, incomplete post-market evidence, reimbursement barriers, unsettled liability, and fragmented international rules.
Key principles for physician-centered AI policy:
- Patient safety first: Prospective validation, external testing
- Evidence-based regulation: Demand prospective trials for high-risk AI
- Transparent accountability: Clear liability when AI errs
- Equity mandatory: Performance tested across demographics; biased AI not deployed
- Physician autonomy preserved: AI supports, never replaces judgment
- Reimbursement aligned with value: Pay for AI that improves outcomes
What physicians must do:
Individually: Demand evidence, validate AI locally, document AI use meticulously, report errors, maintain clinical independence.
Institutionally: Establish AI governance committees, implement bias audits, create accountability frameworks, provide training.
Professionally: Engage specialty societies, lobby for evidence-based regulation and reimbursement, publish validation studies.
The future of AI in medicine will be shaped by the choices made today: the regulations demanded, the reimbursement models advocated for, the governance structures built, and the standards held.