Appendix G: The Clinical AI Morgue
Documented limitations: - Watson for Oncology: Variable concordance across retrospective studies, without patient-outcome evidence - Epic Sepsis Model: 33% sensitivity and 12% positive predictive value in one external validation - Symptom checkers: Product- and vignette-specific diagnostic and triage limitations - Population-health algorithm: A cost proxy systematically understated health need for Black patients
Warning signs that predict failure: - Internal validation only - No prospective outcome studies - “AI diagnoses everything” claims - Workflow integration ignored - No fairness testing
Key lesson: A model score is not clinical utility. Match the evidence design and endpoint to the claim, then evaluate the complete human-AI workflow.
Introduction: Learning From Documented Failure
Learning from failures is as important as celebrating successes. This appendix catalogs notable medical AI failures: what went wrong, why, and lessons for physicians evaluating future AI systems.
The scope of the evidence: - Watson for Oncology: Retrospective studies measured variable concordance, not patient outcomes - Epic Sepsis Model: One external validation found 33% sensitivity and 12% positive predictive value at the evaluated threshold - Diabetic-retinopathy deployment in Thailand: A prospective deployment study documented image-quality, workflow, connectivity, and patient-expectation problems - Symptom checkers: Independent reviews show heterogeneous diagnostic and triage performance - Population-health algorithm: A cost proxy systematically underestimated Black patients’ health needs
Financial boundary: No defensible cross-system total can be calculated from the public evidence used here. Contracts, implementation labor, remediation, and opportunity costs are inconsistently reported.
Patient-impact boundary: Model errors and inequitable allocation create plausible harm pathways, but simulated errors, missed alerts, and discordant recommendations are not automatically observed patient injuries.
Clinician-impact boundary: Alert burden, workflow friction, and trust should be measured directly rather than inferred from a model score.
Understanding these failures helps physicians recognize warning signs and demand better AI.
1. IBM Watson for Oncology (2013–2021)
What it promised: Watson for Oncology was presented as a cognitive-computing system that could organize patient information and medical literature to support treatment recommendations.
What published evidence showed: Retrospective studies mainly measured concordance between Watson recommendations and local multidisciplinary decisions. Concordance varied by cancer type, stage, and setting. In one breast-cancer cohort, concordance differed across clinical subgroups; in a separate multicancer study, performance also varied by disease and recommendation category (Jie et al., 2021; Zhou et al., 2019).
What the evidence did not establish: These studies did not randomize patients to Watson-supported versus conventional care, did not demonstrate improved survival or toxicity outcomes, and do not support one universal error rate. Public reporting about unsafe hypothetical recommendations should not be rewritten as observed patient injury.
Why the distinction matters: A concordance study asks whether two recommendation processes agree. It does not establish which recommendation is better, whether disagreement is unsafe, or whether use improves care. Local formularies, practice patterns, guideline versions, patient preferences, and the system version can all affect agreement.
Lesson: Watson for Oncology illustrates the gap between a market narrative and the endpoint actually studied. Evaluation should preserve the product version, cancer type, comparator, recommendation category, and patient-relevant endpoint.
2. Epic Sepsis Model: Version and Transportability
What it promised: An EHR-embedded score intended to identify patients at increased risk of sepsis and prompt earlier clinical evaluation.
What one external validation found: Wong and colleagues evaluated the model at Michigan Medicine across 27,697 patients and 38,455 hospitalizations. At the evaluated threshold, hospitalization-level AUROC was 0.63, sensitivity was 33%, and positive predictive value was 12%. The study’s 7% result described the subgroup of sepsis hospitalizations in which the model alerted, sepsis was not yet recognized by the clinical team, and antibiotics were administered within three hours. It was not a 7% before-onset detection rate (Wong et al., 2021).
What later evidence added: A prospective evaluation of version 2 across four health systems found stronger discrimination than the earlier version, but site variability, low positive predictive values, and substantial alert burden remained important (Wong et al., 2026).
Why apparently conflicting studies can both be informative: Model version, outcome definition, threshold, prevalence, site configuration, recipients, and workflow all change observed performance. A sepsis model is not one immutable intervention.
Lesson: Every deployment record should preserve the exact model version, threshold, site, alert recipients, workflow, and change history. Retrospective or prospective discrimination alone does not establish a patient-outcome benefit.
3. Google Flu Trends (2008–2015)
What it promised: Real-time flu outbreak prediction from search queries, faster than CDC surveillance.
What went wrong: - During the 2012–2013 influenza season, Google Flu Trends estimated more than twice the proportion of influenza-like illness visits reported by the CDC surveillance system at the peak. - Search behavior, media attention, and changes to the search platform complicated the relationship between query frequency and disease incidence. - Google shared preliminary work with CDC influenza experts during development, but the model still lacked the transparent, continuously updated epidemiologic integration needed for a dependable surveillance system.
Why it failed: - Overfitting to historical correlations that did not generalize - Search behavior changed, and algorithm updates or media coverage could produce searches unrelated to actual illness. - Confounding: Increased searches ≠ increased flu (cyberchondria, media-induced anxiety) - Performance drift: Model degraded as data distribution shifted - Lack of transparency prevented independent evaluation and correction
Lesson: A high-dimensional digital signal remains a proxy, not a case definition. Search data can complement surveillance, but performance must be monitored against epidemiologic reference systems and updated when behavior or platform dynamics change (Lazer et al., 2014).
4. COVID-19 Chest X-Ray AI (2020–2021)
What it promised: Rapid COVID-19 diagnosis from chest X-rays, supplementing scarce PCR testing.
What went wrong: - Roberts and colleagues identified 2,212 studies, screened 415 in the initial review, and included 62 papers after quality screening. - The review found no reported model likely to be a candidate for clinical translation in its reported form. - High or unclear risk of bias, weak reporting, inappropriate controls, and limited external validation were recurring problems.
Why they failed: - Training data biases: COVID-positive images from different sources/equipment than controls. AI learned image source, not pathology - Confounding: COVID patients supine (portable X-rays), controls standing (PA/lateral). AI detected positioning - Data leakage: Duplicate images in training and test sets - Lack of external validation - Biological implausibility: CXR findings in COVID-19 non-specific (similar to other viral pneumonias)
Outcome boundary: The review documented weaknesses in published evidence. It did not show that every model was deployed, withdrawn, or caused patient harm.
Lesson: Crisis-speed publication does not relax requirements for representative controls, leakage prevention, external testing, and transparent reporting (Roberts et al., 2021).
5. Symptom Checkers: Product-Specific Evidence
What the category promised: Consumer symptom checkers have been promoted as accessible tools for possible diagnoses and care-level recommendations.
What independent evidence found: In a 2015 audit of 23 publicly available symptom checkers using 45 standardized vignettes, the correct diagnosis appeared first in 34% of evaluations and appropriate triage advice was given in 57%. Triage performance ranged from 33% to 78% across products and was more accurate for emergent than self-care scenarios (Semigran et al., 2015). A later systematic review found heterogeneous designs and results, limiting broad product-to-product conclusions (Chambers et al., 2019).
Why the original Babylon framing was unsafe: Evidence about a category or an older product version cannot be transferred to one named company without an exact, contemporaneous product evaluation. Regulatory scrutiny, missed diagnoses, and inappropriate self-care claims likewise require primary records tied to the specific version and use.
Lesson: Symptom-checker evidence is product-, version-, vignette-, and endpoint-specific. Evaluation should report both diagnostic ranking and triage safety, including undertriage and overtriage, rather than treating one aggregate accuracy score as clinical utility.
6. Population-Health Algorithm: Biased Proxy Target
What it promised: AI predicting which patients need extra care coordination, targeting high-risk individuals for interventions.
What went wrong: - 2019 Science paper revealed racial bias: For same level of illness, Black patients scored as lower risk than white patients - At the same risk score, Black patients had greater illness burden than White patients. - Replacing the cost proxy with a measure of health need increased the proportion of Black patients selected for additional help from 17.7% to 46.5%.
Why it failed: - Algorithm used healthcare costs as proxy for health needs - Black patients historically received less care because of structural inequity, producing lower spending that did not mean lower need. - AI learned: “Lower costs = healthier,” when actually lower costs often meant under-treatment - Developers did not test for racial disparities before deployment
Outcome boundary: The study examined allocation through a risk threshold. It did not establish that every person excluded was denied care or measure downstream patient outcomes.
Lesson: This was a target-design failure, not evidence that protected attributes must always be removed. Before deployment, teams should test whether the prediction target represents the clinical construct and whether resource allocation differs across groups (Obermeyer et al., 2019).
7. Constructed Exercise: General Vision Model Used for Imaging
Scenario: A hospital innovation team sends radiographs to a general-purpose consumer image-recognition service and treats generic labels as pathology findings.
Why the design fails: The service has no documented intended medical use, clinically representative training evidence, validated output labels, regulatory record for the diagnostic task, or workflow-specific performance study. A technically valid API response is not a clinically valid interpretation.
Appraisal questions: - What exact intended use is documented? - Were inputs, labels, prevalence, and comparators clinically appropriate? - Is the product authorized for the proposed use where authorization is required? - Has the complete acquisition-to-action workflow been evaluated?
Lesson: General computer vision capability cannot be presumed to transfer to diagnostic imaging. Domain evidence and intended use must be established before any patient-facing use.
8. Theranos as a Non-AI Comparator
This was not an AI system. It is retained only as a governance comparator.
What it promised: Revolutionary blood testing from finger-prick samples, hundreds of tests from tiny volumes.
What went wrong: - Technology did not work as claimed - Company misled investors, regulators, patients - Founder Elizabeth Holmes convicted of fraud (2022)
Why relevant to AI: - Pattern of hype, secrecy, insufficient validation parallels some AI ventures - “Black box” technology claims without transparent evidence - Regulatory failure to demand proof before widespread use
Evidence boundary: Theranos cannot support a claim about AI accuracy, model bias, or automation. It illustrates a broader governance problem: consequential health claims should not outrun analytical validation, quality controls, regulatory obligations, or independent scrutiny.
Lesson: Proprietary status does not excuse the evidence needed for a clinical claim. Procurement should distinguish legitimate protection of intellectual property from refusal to provide the validation, intended-use, quality, and monitoring evidence required for safe use.
9. Streams: Non-AI Digital-Care and Governance Comparator
Category correction: Streams used the NHS acute-kidney-injury detection algorithm and a digitally enabled response pathway. DeepMind explicitly stated that the deployed Royal Free application did not use AI. It should not be counted as evidence that an AI model succeeded or failed.
What the service evaluation found: The pathway combined a mobile application, specialist response team, and care protocol. It improved recognition and some treatment-process measures, but did not show a significant step change in the primary renal-recovery outcome and did not establish definitive clinical-outcome benefit (Connell et al., 2019).
Governance lesson: The Royal Free data-sharing arrangement was separately investigated by the UK Information Commissioner’s Office. Data governance and clinical effectiveness are different questions, and one should not be used as a substitute for the other.
Lesson: A coupled digital workflow can improve processes without proving patient-outcome benefit, and a non-AI system should not be relabeled as AI. Technical category, governance record, process endpoint, and clinical endpoint should remain distinct.
10. Sepsis Prediction Across Versions and Workflows
Beyond Epic (general problem):
What promised: Early sepsis detection from EHR data (vital signs, labs).
What went wrong: - Model performance and alert burden vary across systems, versions, thresholds, sites, prevalence, and outcome definitions. - Retrospective discrimination does not establish that an alert changes recognition, treatment, ICU use, mortality, or antimicrobial harm. - An alert’s effect depends on who receives it, what action is expected, existing sepsis processes, competing alerts, and local capacity.
Why failing: - Sepsis definition ambiguous, varies across institutions - Look-ahead bias common (models learn from data collected after clinicians suspected sepsis) - Outcome labels noisy (billing codes, retrospective chart review) - Sepsis is a heterogeneous syndrome, so labels, onset definitions, and use cases require explicit specification.
Outcome boundary: It is inaccurate to assign one performance profile or one failure mechanism to every sepsis product. Product-specific evidence should not be transferred across vendors.
Lesson: The clinically relevant unit is the versioned alert-response pathway, not a generic category called sepsis AI. Evaluation should match the claimed benefit: transportability studies for performance, workflow studies for use, and comparative outcome studies for patient benefit.
11. LLM Benchmark Versus Reasoning Robustness (2025)
What was promised: LLMs achieving near-perfect accuracy on medical benchmarks like MedQA demonstrate clinical reasoning capability ready for deployment.
What was tested: Bedi and colleagues altered multiple-choice medical questions so that some correct responses became “None of the Other Answers.” This perturbation tested whether model performance remained stable when familiar answer patterns changed (Bedi et al., 2025).
Why it matters: - Benchmark performance may significantly overstate clinical reasoning capability - Novel clinical presentations require reasoning beyond memorized patterns - Robustness to a benchmark perturbation is informative, but it is still not a direct test of patient care.
Lesson: High benchmark accuracy does not establish robust clinical reasoning. Perturbation tests, open-ended reasoning tasks, workflow simulation, and prospective human-use evaluation answer different questions and should not be collapsed into one score.
12. LLM-User Interaction Gap (2026)
What was promised: LLMs with near-perfect medical knowledge would help patients make better health decisions.
What went wrong:
- Randomized trial (n=1,298): LLMs alone identified conditions in 94.9% of cases
- Participants using those same LLMs: fewer than 34.5% condition identification, no better than internet search
- Users provided incomplete symptoms; LLMs suggested correct conditions but users failed to recognize them
- Standard benchmarks (MedQA) and simulated patient interactions both failed to predict these real-world failures
Why it matters:
- Performance gap between LLM-alone and LLM+user cannot be closed by improving medical knowledge alone
- Safety testing with real users is required before deployment; benchmarks and simulations are insufficient proxies
Outcome: Authors recommend systematic human user testing before public healthcare deployments.
Lesson: Strong model-only performance did not transfer to strong user-plus-model performance in this controlled study. The result supports human-user testing, not a claim that every public use fails or that communication is the only bottleneck (Bean et al., 2026).
Common Themes in AI Failures
The cases support a system-level interpretation consistent with To Err Is Human (Kohn et al., 2000). A model is embedded in target selection, data production, procurement, configuration, workflow, human interpretation, escalation, and monitoring. Clinical AI failure is usually a property of that complete system, not a single bad score or a single person’s mistake.
This framing also prevents unsupported causal stories. Variable Watson concordance does not prove that one engineer produced unsafe treatment. A low-positive-predictive-value sepsis alert does not prove that clinicians ignored it or that the alert caused harm. Those are testable workflow and outcome claims, not automatic consequences of an accuracy result.
1. Validation mismatch: - The evaluated population, prevalence, site, inputs, or workflow differs from deployment. - The comparator does not represent current care. - The sample is too small or selective for the precision and subgroup claims made. - Internal and external validation are confused with evidence of clinical utility.
2. Data Problems: - Biased, non-representative training data - Confounders not recognized - Noisy labels (billing codes, retrospective annotation) - Data leakage (overlap between training and test)
3. Overfitting and limited transportability: - Models memorize training data specifics - Fail when applied to different populations, institutions, timepoints
4. Insufficient transparency: - Intended use, version, threshold, inputs, exclusions, and output meaning are unclear. - Proprietary restrictions prevent independent evaluation of a consequential claim. - Vendor claims are not linked to a study or regulatory record that supports them.
5. Hype and conflicts of interest: - Marketing exceeds evidence - Company-funded studies vs. independent evaluations - Media amplifies claims without scrutiny
6. Inadequate clinical grounding: - Developers lack domain expertise - AI solutions to non-existent problems - Technology-first, not problem-first approach
7. Regulatory and oversight gaps: - Regulatory status is missing, imprecise, or transferred from another product. - Privacy and data-governance requirements are treated as afterthoughts. - Authorization is presented as proof of effectiveness in a local workflow.
For a field-level map of where health AI breaks, Salaudeen, Ghassemi, and colleagues review common reliability failures spanning erroneous model outputs, clinically unjustified performance gaps, and deployment-time degradation across predictive and generative systems (Salaudeen et al., 2026). Use the taxonomy to structure case review and vendor diligence alongside the case-based morgue entries above. Perspective taxonomy, not new trial evidence of patient benefit or harm for a specific product.
Warning Signs and Red Flags
Red flags suggesting potential failure:
- Extraordinary claims without extraordinary evidence (“Better than doctors,” “Revolutionary”)
- Evidence that does not match the claim (an accuracy study cited for mortality benefit)
- No independent evaluation where independence matters (only marketing material for a consequential claim)
- Unexamined transportability (no evidence for the proposed population, site, or acquisition pathway)
- Missing operational transparency (version, threshold, inputs, exclusions, or intended use not disclosed)
- Conflict of interest (only company-funded studies, no independent evaluation)
- Deployment before validation (widespread use without proof of benefit)
- No discussion of limitations (every AI has failure modes. If none mentioned, red flag)
- Privacy concerns (inadequate consent, data governance)
- Regulatory ambiguity (no exact authorization record, pathway, intended use, or research-use-only boundary where applicable)
A warning sign is a reason to investigate, not an automatic verdict that a product fails. The strength and type of evidence should be proportionate to the risk and claim.
MD Anderson Oncology Expert Advisor (2013–2016)
What it promised: The Oncology Expert Advisor was intended to organize structured and unstructured patient data, support treatment and management suggestions, match clinical trials, and aid research.
Documented investment and status: A 2017 JNCI report, drawing on the University of Texas audit, reported that the project had cost $62 million and was not used on actual patients before the contract expired (Schmidt, 2017).
What went wrong: - MD Anderson launched “Oncology Expert Advisor” powered by Watson - Watson could not integrate with MD Anderson’s EHR systems - Training required manual data entry by physicians (hours per case) - The project encountered delays, procurement problems, cost overruns, and integration challenges. - It did not progress to clinical use on patients. - The audit addressed project governance and implementation; it did not establish that the system injured patients or scientifically validate one causal explanation for the program’s outcome.
Outcome: The collaboration ended without clinical deployment of the Oncology Expert Advisor. This project should not be conflated with the separately marketed Watson for Oncology product or with its retrospective concordance studies.
Why it failed: - Technical immaturity: Watson was not ready for real-world clinical deployment - Poor needs assessment: Did not solve actual physician pain points - Scope and integration risk: A broad clinical system depended on complex data and workflow integration. - Procurement and governance risk: The audit identified weaknesses in project oversight and contracting. - Evidence sequencing risk: Clinical benefit could not be established because the project did not reach clinical use.
Opportunity-cost boundary: The documented expenditure supports scrutiny of what the program delivered. It does not support invented estimates of clinician hours, staffing equivalents, or reputational damage.
Lesson: A large institution, major vendor, and substantial investment do not substitute for defined deliverables, integration evidence, and staged clinical evaluation.
13. Diabetic-Retinopathy AI in Thailand (2018–2020)
What it promised: Automated assessment of retinal photographs could support diabetic-retinopathy screening and faster referral decisions.
What went wrong: - A human-centered study across 11 clinics documented tension between the model’s image-quality requirements and images obtainable in resource-constrained settings. - Connectivity problems affected the ability to return results during the visit. - Nurses sometimes needed repeated image acquisition or pupil dilation, affecting workflow and patient expectations. - A rapid result could also create an expectation of immediate referral capacity that the surrounding system could not always meet.
Outcome boundary: The study documented implementation friction. It did not report that the entire program was discontinued, that clinics permanently returned to traditional screening, or that one fixed percentage summarized field performance.
Why it failed: - Lab-to-field gap: Did not test in real-world conditions before deployment - Inadequate workflow analysis: Did not understand clinic time constraints - Training and acquisition: Staff expertise and local acquisition conditions affected whether images met the model’s requirements. - Infrastructure assumptions: Assumed reliable internet (not available) - User research failure: Did not involve nurses in design
Clinical impact pathway: Failed acquisition, delayed output, repeat imaging, and limited referral capacity can reduce the effective value of an otherwise accurate classifier. Those pathway effects should be measured rather than inferred.
Lesson: Field performance includes acquisition, connectivity, staff behavior, referral capacity, and patient flow, not only model discrimination (Beede et al., 2020).
14. Constructed Exercise: Pathology Assistant Transfer
Scenario: A pathology service deploys a breast-cancer decision-support model developed on slides from one institution. The local laboratory uses different scanners, stains, tissue preparation, prevalence, and review conventions.
Potential problems to test: - Slide or scanner shift changes sensitivity, specificity, or failure-to-process rates. - The operating threshold produces additional review burden without a measured improvement in detection. - The interface changes attention or review order. - Subgroup and rare-morphology performance are not estimated precisely enough for the intended claim.
Required evidence: Local technical verification, appropriately external clinical validation, pathologist usability testing, workflow time, discordance review, version control, and a monitoring plan linked to scanner or laboratory change.
What a responsive vendor would do: Investigate errors with the laboratory, preserve version and scanner information, recalibrate only under an approved change process, test the revised system, and report whether the change improved the prespecified endpoint.
Lesson: A plausible pathology workflow failure should be tested as a hypothesis, not attributed to a named company without a source. Iteration can improve a product, but each changed version needs evidence appropriate to the new claim.
15. Constructed Exercise: Fabricated Chemotherapy Protocol
Scenario: A clinician asks a general-purpose LLM to reproduce a pediatric oncology regimen. The response contains plausible drug names, schedules, and monitoring language but combines elements from different protocols and omits patient- and phase-specific constraints.
Potential consequence: If copied into practice, an incorrect regimen could produce under-treatment, toxicity, or delayed therapy. The exercise does not claim that such an event occurred.
Safe response: The generated text is not used as an order or verification source. The clinician follows the current institutional protocol, pharmacist verification, oncology order set, and required independent checks.
Why it happened: - LLM hallucination: GPT-4 generates plausible but false information - Unknown source and currency: The response does not establish which protocol, version, population, or source supports it. - Confidence without competence: LLM presented wrong answer authoritatively
Lesson: A general-purpose LLM should not serve as the source of truth for chemotherapy selection, dosing, scheduling, or order verification. Medication decisions require the controlling protocol and established safety checks.
16. Constructed Exercise: Sepsis Alert Fatigue
Scenario: A hospital deploys an EHR-embedded sepsis prediction model without local calibration or a prespecified response pathway. Alert volume is higher than anticipated, and clinicians report that many alerts do not change management.
What it promised: Early sepsis detection, 6-12 hours before clinical recognition, 85% sensitivity.
What the team should measure: - Alerts per patient-day and per responsible clinician - Positive predictive value, sensitivity, time relative to clinical recognition, and actions taken - Duplicate alerts, overrides, response time, workload, antibiotic use, and escalation - Patient-relevant outcomes and unintended effects relative to an appropriate comparator
Near-miss exercise: The team should prospectively define how it would investigate a delayed response after a true-positive alert. The review should examine the complete system, including timing, recipients, competing alerts, staffing, interface, protocol, and whether earlier action was clinically indicated.
Decision options: Modify the threshold or workflow only with controlled evaluation, narrow the intended use, improve escalation, pause deployment, or terminate the system if the benefit-risk case is not met.
Why it failed: - No threshold customization: One threshold can produce an unacceptable benefit-burden tradeoff in a new population. - No local validation: Vendor’s sensitivity/specificity did not match hospital’s patient population - Alert burden unmeasured: No universal false-positive percentage defines an acceptable alert. Consequences depend on prevalence, workload, action, and benefit. - No physician input: Deployed top-down without ED/ICU physician buy-in
Lesson: Alert burden and response behavior must be measured locally; they cannot be inferred from a universal false-positive cutoff.
Part 3: Emerging AI Failure Modes (2023–2026)
17. Constructed Exercise: LLM-Generated Patient Education
Evidence label: This is a hypothetical appraisal exercise, not a reported physician case.
What happened: - Physician used ChatGPT to generate patient education handout on warfarin management - LLM output looked professional, included dietary restrictions, monitoring advice - Errors in LLM output: - Recommended avoiding all leafy greens rather than maintaining a consistent intake. - Omitted critical drug interactions (NSAIDs, antibiotics) - Included one INR target range for every patient despite indication-specific differences.
Safe workflow: A clinician compares the draft with the controlling anticoagulation guidance, updates it for the patient’s indication and medicines, checks readability, and retains responsibility for the final handout.
Lesson: Fluent patient education still requires source verification, indication-specific review, and clinical accountability.
18. Constructed Exercise: AI Triage in Urgent Care
Evidence label: This scenario is fictional. It does not describe an observed patient, vendor, verdict, or adverse event.
Vendor: Symptom checker AI for urgent care front-desk triage
What it promised: Prioritize high-acuity patients, reduce wait times.
Scenario: An older adult reports chest tightness. A symptom-based system assigns low acuity, and front-desk staff treat the score as a final decision instead of applying the organization’s chest-pain pathway.
Potential consequence: The design creates a pathway for delayed electrocardiography and clinical assessment. No outcome is asserted because the scenario is hypothetical.
Why it failed: - Unknown transportability: Age and presentation-specific performance have not been established. - Incomplete input: The system may not collect the information required for safe triage. - Authority mismatch: Staff are not given a clear override and escalation pathway.
Lesson: High-risk symptom pathways need explicit escalation rules that do not depend on an unvalidated model score.
19. Differential-Diagnosis Evaluation Across 21 Models (2026)
What was promised: LLMs achieving high benchmark scores on clinical knowledge exams would assist with differential diagnosis.
What went wrong: - All 21 LLMs tested (ChatGPT, DeepSeek, Claude, Gemini, Grok) failed to produce an appropriate differential diagnosis on more than 80% of cases when given initial patient presentation only - Even with full clinical information, final diagnosis failure rates exceeded 40% in some models - Performance varied by model and reasoning stage.
Why it matters: - Models that perform well on USMLE-style questions do not demonstrate comparable differential diagnosis reasoning - Clinical reasoning requires integration of context and probability in ways that standardized benchmarks do not capture
Important context: This cross-sectional study used 29 standardized vignettes and 16,254 scored responses, not live clinical encounters. The failure rate exceeded 0.80 for differential diagnosis across all tested models, while final-diagnosis failure rates were below 0.40 (Rao et al., 2026). This pattern coexists with evidence that an o1-series model performed strongly on controlled text-based clinical reasoning and emergency-department second-opinion tasks (Brodeur et al., 2026).
Lesson: Medical knowledge benchmarks and longitudinal reasoning evaluations measure different capabilities. Neither alone establishes safe clinical deployment.
20. NOHARM Safety Benchmark and Physician-AI Teaming (2026)
What was found: - The July 2026 version 4 preprint describes a 1,100-task benchmark across 10 specialties, with 12,747 expert annotations for 4,249 clinical-management options. - Across 20 LLMs and four retrieval-augmented clinical AI tools, direct application of recommendations had potential for severe harm in up to 24.6% of cases. - Omissions accounted for more than 80% of severe errors. - In a randomized study of 101 U.S.-licensed generalist physicians, AI assistance improved written performance compared with conventional resources, but physicians frequently omitted valuable AI-generated recommendations.
Why it matters: - Potential-harm scoring can reveal action-level omissions and commissions that knowledge benchmarks miss. - Evaluating AI by benchmark performance is insufficient for clinical deployment decisions
Evidence boundary: These were rubric-derived potential-harm scores on written consultation responses, not observed adverse events or patient outcomes. The randomized component measured written physician responses, and the study remains a preprint (Wu et al., 2026, preprint).
Lesson: Evaluation should measure completeness and severity of omissions and commissions, then test the human-AI workflow rather than assuming that either component’s standalone score transfers to practice.
Part 4: Consolidated Lessons
Pattern 1: Validation Mismatch
Validation is not one binary state. Development performance, internal validation, temporal validation, geographic external validation, prospective workflow evaluation, and comparative clinical-utility studies answer different questions.
Common root causes: - The intended population, prevalence, site, inputs, acquisition process, or care pathway differs from the evaluation setting. - The comparator is weak or absent. - The endpoint is substituted. For example, discrimination is presented as evidence of mortality benefit. - Uncertainty is omitted, especially for clinically important subgroups or rare outcomes. - A later model version is treated as if it were the version evaluated in an earlier paper.
Examples: The Epic studies show why version and site matter. The COVID-19 imaging review shows how many publications can share the same design weaknesses. Watson concordance studies show why agreement is not a patient outcome.
Red flags: - A large patient count is used to avoid discussing whether the sample is representative. - Accuracy is reported without the denominator, confidence interval, prevalence, threshold, or error consequences. - Deployment count is offered as proof of effectiveness. - A prestigious journal name is treated as more important than design and endpoint.
What to require: Evidence that matches the claim. External validation supports transportability. Prospective workflow evaluation supports claims about real-world use. Comparative evaluation with prespecified patient-relevant endpoints is needed for outcome-benefit claims. No universal number of sites or one preferred journal can substitute for design fitness.
Pattern 2: Data and Target Failures
Common root causes: - Training data do not represent the proposed population or acquisition process. - The label does not represent the clinical construct. - A convenient proxy, such as cost, encodes unequal access instead of health need. - Duplicate patients or temporally inappropriate inputs leak information across splits. - Clinical actions taken because clinicians already suspected the outcome become predictors of that outcome.
Examples: Google Flu Trends depended on a digital-behavior proxy that changed over time. COVID-19 image models could learn source or positioning differences. The population-health algorithm predicted spending rather than need. Sepsis models depend heavily on onset and label definitions.
Red flags: - Data provenance, inclusion dates, exclusions, and label construction are not disclosed. - The model uses billing codes without explaining their role and limitations. - Missingness is treated only as a technical nuisance rather than a possible signal of care processes. - The evaluation does not protect the test set from tuning or repeated inspection.
What to require: A data-lineage record, clinically justified target, leakage audit, representative error analysis, temporal evaluation, and monitoring tied to known sources of change. Billing codes and retrospective labels are not automatically invalid, but their fitness must be justified for the intended claim.
Pattern 3: Equity and Allocation Failures
Common root causes: - The target encodes unequal access, spending, documentation, or treatment. - Performance is averaged across groups with different prevalence or consequences. - The intervention generated by the score is not audited. - Protected attributes are removed without examining proxies or allocation effects.
Examples: The cost-based population-health algorithm shows how accurate prediction of the wrong target can reproduce inequity. The constructed pathology exercise shows why subgroup claims need adequate precision, but it is not evidence against any named product.
Red flags: - “The model does not use race” is offered as a complete fairness analysis. - Demographics, sample sizes, uncertainty, and missingness are not reported. - One universal parity tolerance is imposed without clinical justification. - Only model metrics are reviewed, while eligibility, referral, follow-up, and intervention receipt are ignored.
What to require: Prespecified subgroup analyses, uncertainty, clinically meaningful fairness questions, allocation audits, and mitigation tied to the use case. A fixed difference such as 10% is not a universal definition of fairness. Acceptable tradeoffs depend on the endpoint, prevalence, intervention, and harms of errors.
Pattern 4: Workflow Integration Failures
Common root causes: - The system is designed around the model instead of the care pathway. - Input acquisition, failed inputs, connectivity, review time, and referral capacity are excluded from evaluation. - The output does not arrive to the right person at the right time. - Training teaches button use but not intended use, limitations, escalation, or override. - Change management is missing.
Examples: The Thailand study documented acquisition, connectivity, staff, and patient-expectation problems that laboratory performance could not capture. The MD Anderson project illustrates integration and governance risk. Streams shows that a mobile alert plus a specialist response team is a combined intervention, not a model alone.
Red flags: - “Plug and play” is used without a site-integration plan. - The vendor cannot explain failed-input handling or downtime. - Usability testing excludes the clinicians and staff who will act on the output. - Time saved is claimed without measuring review, correction, documentation, and follow-up.
What to require: Workflow mapping, usability and human-factors evaluation, failed-input analysis, time-motion measurement where relevant, training, escalation, downtime procedures, and rollback authority. Configuration should be controlled and re-evaluated rather than casually customized.
Pattern 5: Alert Burden and Automation Failures
Common root causes: - Thresholds are chosen from model metrics without considering alert volume and clinical action. - Duplicate or low-value alerts compete with higher-priority work. - The interface encourages reflexive acceptance or dismissal. - Overrides and nonresponse are counted without investigating why they occurred. - Monitoring stops at sensitivity or AUROC instead of measuring response and effect.
Examples: The Epic validation demonstrates low positive predictive value and missed cases at one threshold. The synthetic sepsis scenario illustrates the questions a governance team should ask, but its values are not evidence.
Red flags: - Alerts per clinician, patient-day, or actionable event are not reported. - Positive predictive value is missing in a low-prevalence use case. - The organization has no definition of the action an alert should trigger. - A false-positive percentage is labeled acceptable or unacceptable without reference to benefit, workload, and consequence.
What to require: Threshold analysis linked to action, alert burden, response time, overrides, downstream testing and treatment, missed-event review, and patient-relevant outcomes. There is no universal 10% or 20% false-positive ceiling.
Part 5: The Cost of Failure
Public evidence does not support one aggregate dollar figure for the cases in this chapter. Combining a documented project expenditure with invented license prices, hospital counts, or assumed clinician hours would create false precision.
Financial Cost
Documented cost: The MD Anderson Oncology Expert Advisor project was reported at $62 million. That figure belongs to that project and should not be transferred to Watson for Oncology generally.
Costs that should be measured locally: - License, integration, interface, hardware, and infrastructure - Data mapping, validation, cybersecurity, privacy, and regulatory work - Clinician and staff training, review, correction, and support time - Downstream tests, referrals, treatment, and follow-up generated by outputs - Incident investigation, remediation, revalidation, and rollback - Opportunity cost relative to the best alternative intervention
Cost-effectiveness boundary: A lower cost per prediction is not evidence of better value. Economic evaluation needs an appropriate comparator, time horizon, perspective, uncertainty, and measured clinical consequences.
Patient-Safety Cost
The chapter contains three different evidence types that must not be merged:
- Observed allocation inequity: The population-health study measured who was selected for additional help under a biased proxy.
- Measured model or workflow limitations: Epic, Thailand, COVID imaging, symptom-checker, and LLM studies measured specified endpoints in particular designs.
- Constructed exercises: The chemotherapy, pathology, sepsis-alert, patient-education, and urgent-care scenarios illustrate plausible pathways but do not document patient injury.
A simulated error, discordant recommendation, or missed model prediction is not automatically an observed adverse event. Safety accounting should distinguish potential harm, near miss, process failure, and confirmed patient harm.
Trust and Legitimacy Cost
Trust should also be measured rather than asserted. Relevant indicators include clinician reliance, override patterns, patient understanding, complaints, perceived fairness, willingness to use the pathway, and response after an incident. A transparent correction and rollback process may protect trust better than defending an underperforming system.
The equity case adds a legitimacy question: even a statistically accurate system can lose legitimacy if its target allocates resources in a way that reproduces unequal access. Governance should examine the allocation decision, not only the prediction.
Part 6: How to Avoid These Failures
For Individual Physicians
Before relying on a new AI tool:
- Identify the exact system
- Product, version, intended use, inputs, outputs, population, and threshold
- FDA or other regulatory pathway and record where applicable
- Research-use-only status or other use limitation
- Match evidence to the claim
- Development, internal, external, prospective, or comparative design
- Endpoint and comparator
- Same product version and workflow as the proposed use
- Confidence intervals and clinically relevant error analysis
- Examine safety and equity
- Omissions, commissions, failed inputs, and escalation
- Performance and allocation across clinically relevant groups
- Potential interactions with access, documentation, and existing care disparities
- Assess workflow
- Who reviews the output and when
- What action is expected
- Work added or removed
- Override, downtime, incident, and rollback procedures
- Preserve professional judgment
- Authorization does not transfer clinical responsibility.
- Human review is not sufficient if the reviewer lacks time, information, or authority.
- Disagreement should trigger a defined resolution pathway, not silent acceptance or dismissal.
Decision rule: A warning sign should trigger further evidence review. A material unresolved risk should block use until addressed. The decision should not depend on a fixed count of publications or reference sites.
For Hospital Leaders
Before system-wide deployment:
- Define a risk-proportionate evaluation period
- Scope, population, comparator, endpoints, stopping rules, and sample-size rationale
- Duration based on event rate, workflow cycles, and the decision to be made, not a universal three- or six-month rule
- Assign accountable governance
- Clinical owner, operational owner, technical owner, privacy and security review, and executive escalation
- Authority to pause or roll back the system
- Conflict-of-interest disclosure and independent review where appropriate
- Measure the complete pathway
- Technical performance and failed inputs
- Human response, workload, overrides, and downstream actions
- Patient-relevant outcomes when benefit is claimed
- Equity, access, and unintended consequences
- Control change
- Version, prompt, threshold, interface, data pipeline, and clinical-practice changes
- Re-evaluation triggers tied to material change or performance signals
- Communicate honestly
- Describe what the system does and does not do.
- Provide staff with limitations and escalation procedures.
- Address patient notice, explanation, and choice according to law, ethics, risk, and institutional policy.
Scale decision: Expansion should occur only when prespecified evidence supports the intended benefit-risk case. Failure to meet the bar should lead to correction, narrowing, pause, or termination, not automatic escalation of commitment.
For Developers and Vendors
- Define the clinical problem and intended use before optimizing the model.
- Use data, targets, labels, and comparators that fit the clinical construct.
- Include clinicians, staff, patients, and operational teams whose work or care will change.
- Report versioned performance, uncertainty, limitations, exclusions, and failed inputs.
- Evaluate transportability and real workflow use before making outcome claims.
- Provide monitoring, incident, update, and rollback support.
- Verify every regulatory claim against the exact device record and intended use.
Responsive iteration is valuable, but a changed system is not validated merely because an earlier version was studied. Change history is part of clinical evidence.
Conclusion: Learning From the Evidence
The documented cases do not support an aggregate claim of more than $100 million wasted, thousands of physician hours lost, or thousands of patients harmed. They support a more useful conclusion: failure occurs when claims outrun the specific evidence for a versioned system in a defined workflow.
Common threads: 1. Evidence and claim are mismatched. 2. Targets, labels, or proxies do not represent the clinical objective. 3. Transportability and uncertainty are ignored. 4. Workflow, human behavior, and follow-up capacity are excluded. 5. Equity is assessed only at the model level, if at all. 6. Version, threshold, and interface change without controlled re-evaluation. 7. Marketing or institutional prestige substitutes for proof.
Principles that improve the odds of success: - Evidence before scale: Use the design needed for the claim. - Operational transparency: Preserve intended use, version, threshold, inputs, outputs, limitations, and change history. - Equity by design: Evaluate both prediction and allocation. - Clinical and patient participation: Include the people affected by the workflow. - Continuous, risk-based monitoring: Monitor when and where failure could matter, with authority to intervene. - Accurate incident language: Distinguish simulated potential harm from observed harm. - Corrective action: Rewrite false claims, repair systems, narrow use, or stop when evidence does not support continued deployment.
The strongest safeguard is not skepticism alone. It is a versioned, evidence-matched, monitored clinical system with accountable owners and a real rollback path.
Appendix: Quick Reference Red Flags
Red flags are prompts for investigation. Several unresolved high-risk findings can justify pausing procurement or deployment, but no single item proves that every product in a category fails.
Validation Red Flags
- The cited study evaluates a different product, version, population, or intended use.
- A technical endpoint is used to claim patient benefit.
- The comparator, threshold, prevalence, denominator, or uncertainty is missing.
- Deployment count is offered in place of effectiveness evidence.
- No evidence addresses the proposed site or workflow.
Data Red Flags
- Data provenance, dates, exclusions, or label construction are not disclosed.
- Leakage, duplicate patients, or temporally inappropriate inputs are not assessed.
- A proxy is used without showing that it represents the clinical construct.
- Failed inputs and missingness are excluded from performance reporting.
Fairness Red Flags
- “Race is not a feature” is presented as the complete fairness case.
- Subgroup denominators and uncertainty are omitted.
- Allocation and downstream intervention are not examined.
- One arbitrary parity threshold is applied to every endpoint.
Workflow Red Flags
- “Plug and play” or “no training needed” is claimed without site evidence.
- The recipient, timing, expected action, escalation, and downtime plan are unclear.
- Alert volume, review time, overrides, or downstream work are not measured.
- There is no rollback owner.
Product and Governance Red Flags
- The exact regulatory status and intended use cannot be verified where applicable.
- Research-use-only output is presented as authorized diagnostic use.
- The vendor will not disclose material limitations, version history, or incident process.
- Contract terms prevent required monitoring, investigation, or exit.
- Pricing and integration obligations are too vague to estimate total cost.
Decision: Do not deploy when a material safety, evidence, regulatory, privacy, or governance gap remains unresolved. Document the gap, the evidence requested, the accountable owner, and the condition for reconsideration.
What did published Watson for Oncology evaluations show?
Published studies mainly measured concordance with multidisciplinary recommendations, not patient outcomes. Concordance varied by cancer and setting, so the evidence does not support one universal failure rate or a causal claim about patient harm.
What was wrong with Epic’s Sepsis Model?
At one external site and evaluated threshold, the Epic Sepsis Model had a hospitalization-level AUROC of 0.63, 33% sensitivity, and 12% positive predictive value. The study did not establish that 7% was a before-onset detection rate or measure patient harm caused by the model.
What are warning signs of AI that will fail?
Red flags include evidence that does not match the claim, no independent evaluation, missing version or intended-use details, weak workflow testing, undisclosed subgroup performance, unavailable change history, and no monitoring or rollback plan.
Can the total cost of failed medical AI be calculated from public evidence?
No defensible cross-system total can be calculated from the public evidence used here. Contracts, implementation labor, displaced work, remediation, and opportunity costs are inconsistently reported and should not be combined with invented estimates.
What is the main lesson from medical AI failures?
A model score is not clinical utility. Match the study design and endpoint to the claim, preserve the exact version and workflow, and evaluate transportability, human use, monitoring, equity, and patient-relevant outcomes when benefit is claimed.
The Clinical AI Morgue remains a reminder: evidence, workflow, and accountability determine whether a promising model becomes a useful clinical system.