Executive Summary
Clinical AI performance depends on intended use, study design, population, threshold, workflow, and endpoint. Retrospective accuracy, regulatory authorization, adoption, and patient benefit are different claims. No single metric or authorization establishes that a system will improve care in a local workflow.
This summary enables clinicians and health-system leaders to:
- match a claim to the study design and endpoint that can support it;
- preserve product, version, population, and intended-use boundaries;
- recognize common clinical AI failure modes;
- identify implementation and monitoring controls proportionate to risk; and
- route to detailed specialty and implementation guidance.
Purpose
Clinical AI can fail when evidence, population, version, workflow, or endpoint is separated from the claim. Documented examples include unsafe cancer-treatment recommendations, weak external performance of one sepsis-model version at a studied threshold, and diagnostic systems that did not transport to new populations or operating conditions.
This handbook provides evidence-based frameworks for evaluating AI tools critically, understanding limitations, and implementing safely. It references peer-reviewed literature from JAMA, NEJM, The Lancet, Nature Medicine, and specialty journals.
The guidance addresses what practicing physicians need: critical evaluation of AI claims, understanding failure modes, safe workflow integration, and liability navigation.
Key Findings
AI Performance Often Falls Short of Marketing Claims
Clinical performance can differ from development or validation results when the population, prevalence, acquisition process, threshold, user, workflow, or software version changes. Controlled research environments may use curated datasets, standardized imaging protocols, and selected patient populations. Real clinical practice introduces motion artifacts, poor-quality images, atypical presentations, missing data, and populations underrepresented in training data.
External validation assesses transportability to conditions distinct from development. It does not by itself establish workflow benefit or patient outcomes. Performance claims from vendor studies require independent appraisal, and local evaluation should be proportionate to the tool’s risk, novelty, evidence gaps, and intended use.
Retrospective model studies can estimate performance in existing data but cannot establish how displaying an output changes clinician behavior, workflow, or patient outcomes. Prospective and comparative studies remain essential when those are the claims being made.
Evidence Must Match the Claim
| Evidence | Defensible claim |
|---|---|
| Retrospective model evaluation | Performance in the analyzed dataset and configuration |
| External validation | Performance in a distinct site, period, device, or population |
| Prospective workflow study | Effect on users, process, workload, or diagnostic performance during use |
| Comparative clinical study | Difference from a relevant comparator on prespecified endpoints |
| Regulatory authorization | The agency decision for the labeled device and intended use |
| Postdeployment monitoring | Performance and safety after implementation and change |
A diagnostic-accuracy study does not establish mortality benefit. A time-saving study does not establish note safety. A benchmark does not establish autonomous clinical practice. The task, product version, population, comparator, and endpoint must travel with every result.
Documented AI Failures Offer Critical Lessons
IBM Watson for Oncology: Investigative reporting documented unsafe and incorrect recommendations, while published evaluations largely measured concordance with multidisciplinary recommendations rather than patient outcomes (Ross & Swetlitz, 2018; Jie et al., 2021). Variable concordance and documented unsafe recommendations do not establish the safety or clinical utility of the system.
Epic Sepsis Model: An external validation of version 1 at Michigan Medicine found an AUROC of 0.63 and, at threshold 6, 33% sensitivity and 12% positive predictive value. A later prospective evaluation of version 2 across four health systems reported stronger discrimination but substantial site variation, low positive predictive values, and high alert burden (Wong et al., 2021; Wong et al., 2026). The model name alone cannot reconcile results across versions, thresholds, sites, and outcome definitions.
COVID-19 prediction research: A living systematic review found high or unclear risk of bias and concluded that none of the reviewed diagnosis or prognosis models was ready for clinical use at that stage (Wynants et al., 2020).
These examples show recurring failure modes: weak transportability, biased targets, unclear reference standards, overfitting, version drift, and deployment without evaluation of the actual human-AI workflow.
Biased targets: A population-health algorithm used healthcare cost as a proxy for health need. Replacing cost with a direct health-need measure increased the proportion of Black patients selected for additional care from 17.7% to 46.5% in the study analysis (Obermeyer et al., 2019).
Human-AI interaction: In a randomized simulated-vignette study, GPT-4 access modestly improved physician management-reasoning scores, but the study measured neither patient outcomes nor autonomous care (Goh et al., 2025). A separate simulated patient-facing triage evaluation found emergency undertriage in 33 of 64 clear-emergency responses (Ramaswamy et al., 2026). In a preregistered randomized study of 1,298 participants, models that identified conditions correctly in 94.9% of cases when tested alone supported correct identification in fewer than 34.5% of cases when actual users operated them, no better than the control group (Bean et al., 2026). Model capability measured in isolation did not predict the performance of the human-model system.
Evidence of Benefit Is Narrow and Specific
Authorized AI for specific, well-defined diagnostic tasks:
Diabetic retinopathy screening: FDA granted De Novo authorization to IDx-DR, now LumineticsCore, for a specified population and workflow (FDA DEN180001). The prospective pivotal study reported 87.2% sensitivity and 90.7% specificity for its prespecified endpoint. It did not test vision preservation or every retinal disease (Abràmoff et al., 2018).
Chest X-ray triage: FDA-authorized products exist for bounded findings and workflows. Diagnostic and notification-time results remain product-, site-, population-, and endpoint-specific.
ECG interpretation: Selected algorithms can detect atrial fibrillation in defined recordings and populations, but screening yield, downstream testing, treatment, and patient benefit require separate evidence.
Colonoscopy polyp detection: Randomized trials show increased adenoma detection with computer-aided detection in selected settings, while effects on advanced adenoma detection, serrated lesions, workload, and long-term outcomes remain distinct questions. A living guideline based on 44 randomized trials issued a weak recommendation against routine use because patient-important benefits remain uncertain and additional surveillance is a plausible burden (Foroutan et al., 2025).
Large Language Models (LLMs) for clinical tasks:
Documentation: Ambient documentation tools can transcribe encounters and generate draft notes. Randomized and observational evidence is emerging, but correction burden, note quality, privacy, workload, and downstream error require separate measurement.
Literature review: LLMs can summarize research, though accuracy requires verification
Patient communication: Draft responses to patient messages, always requiring physician review before sending
Clinical reasoning support: Controlled text-based studies can measure differential diagnosis or management reasoning. A 2026 Science study reported stronger performance than physician baselines on selected reasoning and emergency-department second-opinion tasks, but did not establish autonomous deployment readiness or patient-outcome benefit (Brodeur et al., 2026).
Autonomous AI agents represent an emerging category. Unlike a single advisory response, an agent can combine a language model with tools, memory, and multi-step actions. Simulated EHR benchmarks and early research do not establish safe autonomous execution in live clinical systems. Agent permissions, data access, action boundaries, identity, rollback, and human escalation therefore require explicit controls.
Important caveat: LLMs can generate plausible but incorrect information, including fabricated citations and unsupported clinical detail. Verification should be proportionate to the clinical consequence, and consequential output should be checked against authoritative sources and the patient record.
Specialty-Specific Evidence Varies Dramatically
Radiology: Radiology has the largest representation on FDA’s periodically updated AI-enabled medical-device list, though FDA states that the list is not comprehensive (FDA AI-Enabled Medical Devices). Evidence is strongest for selected mammography, triage, and stroke-workflow applications, not for radiology AI as a single class.
Pathology: Digital pathology includes FDA-authorized systems for bounded tasks such as adjunctive prostate-cancer detection. Product labeling must not be expanded to grading, prognosis, or other tissues without separate evidence.
Dermatology: Consumer and clinical systems have different regulatory and evidence profiles. Performance and representation concerns require evaluation across skin tones, devices, settings, lesion types, and intended users.
Cardiology: Evidence spans ECG screening, imaging, wearable signals, and workflow interventions. Results from one modality or task should not be generalized across cardiovascular AI.
Oncology: Watson illustrates the limits of broad treatment-recommendation claims. Separate evidence supports selected pathology, imaging, risk-stratification, and biomarker applications.
Primary Care: Documentation and bounded screening tools are active areas of evaluation. Prevalence, continuity, multimorbidity, referral capacity, and workflow differ from specialist settings and affect predictive value and utility.
Recommendations
For Individual Physicians
Before adopting any AI tool:
Demand evidence matched to the claim: retrospective evaluation, external validation, prospective workflow study, and comparative outcome study answer different questions
Ask about transportability to populations, prevalence, devices, sites, users, and workflows similar to the intended use
Understand the training data: what populations, imaging protocols, and clinical settings were represented
Review the FDA pathway and labeling: 510(k) clearance rests on substantial equivalence to a predicate; De Novo and PMA are different pathways, and none substitutes for local clinical evaluation.
Identify failure modes: what happens when the AI is wrong, and how will you detect errors
When using AI in practice:
Respect the intended use: An AI output supports only the task, population, user, and workflow established by its evidence and labeling.
Document consequential use: record relevant output, version, context, verification, and action according to institutional policy and clinical relevance
Monitor for drift: AI performance can degrade as patient populations change or systems update
Report failures: contribute to institutional and national learning from AI errors
For Hospital Administrators and Informaticists
Evaluation before procurement:
Require claim-matched evidence before institutional adoption
Conduct local acceptance and performance testing proportionate to the risk, novelty, transportability gap, and intended workflow
Establish performance monitoring from day one
Define success metrics beyond vendor-provided accuracy claims
Implementation:
Use silent evaluation or governed pilots when evidence gaps and risk justify staged implementation
Build feedback mechanisms for clinicians to report AI errors and near-misses
Plan for workflow integration: technical performance means nothing if clinicians will not use the tool
Allocate ongoing resources for monitoring, maintenance, incident response, vendor changes, and reevaluation
Regulatory awareness: FDA’s January 2026 Clinical Decision Support (CDS) final guidance clarifies which CDS software functions are excluded from the definition of a device under section 520(o)(1)(E) (Non-Device CDS) and notes that FDA digital health policies continue to apply to software functions that meet the definition of a device, including those intended for use by patients or caregivers (FDA CDS Guidance, January 2026). Procurement decisions should not assume that “physician in the loop” alone is sufficient to make a tool Non-Device CDS.
The ACR-SIIM Practice Parameter for Imaging AI applies governance principles to radiology through inventory, local acceptance testing, monitoring, privacy, and continuous quality improvement. It was approved in May 2026 and is scheduled to take effect October 1, 2026. Similar controls should be tailored to risk in other specialties rather than copied without validation.
For Medical Educators
Integrate AI literacy into training:
Teach critical evaluation of AI tool claims and vendor marketing
Include AI failures as case studies alongside successes
Address bias and equity: how AI can amplify health disparities
Prepare for liability: the evolving legal landscape for AI-assisted decisions
Core Principles for AI Adoption
1. Responsibility Must Be Defined
AI changes, but does not erase, professional and organizational duties. Responsibility is fact- and jurisdiction-specific and can involve clinicians, institutions, developers, and manufacturers. FDA authorization, institutional approval, and a “human in the loop” label do not by themselves allocate civil liability (Mello and Guha, 2024).
2. Performance Claims Require Skepticism
Vendor accuracy claims may come from selected datasets, thresholds, populations, and comparators. Demand the denominator, intended use, reference standard, uncertainty, subgroup results, external evidence, and workflow endpoint. Treat marketing materials as claims to verify, not as the evidence itself.
3. Bias Is Embedded, Not Eliminated
AI systems can encode bias from selection, measurement, labels, proxies, access, thresholds, and deployment. Representation in training data is important but is not the only mechanism. Evaluation should test clinically relevant subgroups and investigate how the workflow distributes errors, access, and benefit.
4. Workflow Integration Determines Adoption
The best-performing AI tool delivers no clinical value if clinicians will not use it. Clinical decision support that adds clicks, creates alert fatigue, or disrupts established workflows will be ignored or abandoned. Implementation matters as much as algorithm performance.
5. Documentation Supports Safety and Accountability
When AI materially contributes to a clinical decision, documentation should capture information relevant to care, verification, and institutional policy without copying unnecessary model output into the record. Traceable version and incident records also support quality improvement. Documentation is not a guarantee against liability.
Coverage and Scope
| Section | Focus |
|---|---|
| Front Matter | Orientation, AI agent routing, and executive summary |
| I: Foundations | AI history, fundamentals, data challenges |
| II: Clinical Specialties | AI applications across 19 specialties |
| III: Implementation | Evaluation, ethics, privacy, safety, liability |
| IV: Practical Tools | LLMs, documentation, research applications, and clinical trials |
| V: Future | Emerging tech, policy, global health |
The Bottom Line
AI tools will continue entering clinical practice regardless of individual physician preferences. Success depends on critical evaluation before adoption and careful monitoring after deployment.
The physicians who navigate this transition successfully will be those who:
Maintain healthy skepticism about performance claims
Demand rigorous evidence before adoption
Understand AI limitations and failure modes
Preserve clinical judgment as the foundation of patient care
Document carefully when AI informs decisions
This handbook provides the evidence base and practical frameworks for that critical evaluation.
Questions About the Executive Summary
Do AI tools perform as well in clinical practice as in validation studies?
Not necessarily. Performance can change with population, prevalence, acquisition, workflow, threshold, user behavior, and software version. A development study, external validation, prospective workflow study, and patient-outcome study answer different questions.
What are examples of documented AI failures in medicine?
Examples include variable and sometimes unsafe Watson for Oncology recommendations, poor external performance of one Epic Sepsis Model version at a studied threshold, biased resource-allocation targets, and high-risk-of-bias COVID-19 prediction models. Each example has a distinct design, product, version, and endpoint.
What clinical AI applications currently work well?
Evidence of benefit is narrow and task-specific. Examples include autonomous diabetic-retinopathy screening under a defined De Novo authorization and randomized evaluations of selected mammography and colonoscopy workflows. Evidence from one product or setting cannot be transferred to an entire product class.
Who is responsible when AI makes clinical errors?
Responsibility is fact- and jurisdiction-specific. Clinicians, institutions, developers, and manufacturers may each have duties based on the intended use, labeling, product design and warnings, workflow, contracts, and applicable law. FDA authorization does not allocate civil liability.
What should physicians demand before adopting AI tools?
Physicians should demand evidence matched to the intended claim, the exact product and version, a clear regulatory record when applicable, subgroup and uncertainty analysis, workflow testing, failure and fallback controls, change management, and postdeployment monitoring.
This executive summary is part of The Physician AI Handbook. For detailed analysis, evidence, citations, and specialty-specific guidance, see the full chapters.