Appendix D — Case Studies in Medical AI

TL;DR

The cases below separate diagnostic accuracy, workflow effects, authorization, implementation experience, and patient outcomes.

Successes: - IDx-DR: Prospective diagnostic-accuracy evidence and De Novo authorization for a defined autonomous workflow - Viz.ai: Randomized workflow-time improvement, with no significant improvement in 90-day functional independence in one trial

Failures: - Watson for Oncology: Variable concordance across retrospective studies, with patient-outcome benefit untested - Epic Sepsis Model: 33% sensitivity and 12% positive predictive value in independent external validation - ChatGPT Health (2026): 33 of 64 responses to clear simulated emergencies were undertriaged

Ongoing pilots: - Utah/Doctronic (2026): State-authorized phased pilot that remains under licensed-practitioner review in Phase 1

Common success factors: Narrow scope, evidence matched to the claim, workflow integration, accountable ownership, monitoring, and follow-up

Common failure patterns: Overclaiming the endpoint, inadequate external validation, shortcut learning, poor workflow fit, weak monitoring, and conflating authorization with benefit

Introduction

Detailed case studies of medical AI implementations, covering both successes and failures. Each follows the structure: Background, Implementation, Outcomes, and Lessons Learned.


Case Study 1: IDx-DR and Diabetic Retinopathy

Background

Diabetic retinopathy (DR) is leading cause of preventable blindness. Early detection and treatment prevent vision loss, but screening requires ophthalmologists or trained specialists (scarce in many regions). Millions with diabetes lack regular eye exams.

Solution: IDx-DR, an autonomous AI system analyzing retinal images to detect referable DR (moderate or worse retinopathy requiring ophthalmologist referral).

Implementation

Development and pivotal evaluation: The published pivotal study prospectively enrolled 900 participants with diabetes at 10 primary-care sites. It was a single-arm diagnostic-accuracy study, not a randomized trial (Abràmoff et al., 2018).

FDA submission: The system was evaluated for an autonomous intended use in which the output did not require an eye-care specialist to interpret the retinal images.

FDA authorization: FDA granted De Novo classification in April 2018 under DEN180001. The decision summary defines the indication, eligible population, compatible cameras, output, and special controls (FDA DEN180001).

Clinical performance: The pivotal study reported 87.2% sensitivity and 90.7% specificity for more-than-mild diabetic retinopathy (Abràmoff et al., 2018). These values describe the evaluated study workflow and population.

Deployment: Primary care clinics, pharmacies, endocrinology offices. Non-ophthalmologists capture retinal photos, IDx-DR interprets immediately, result provided within minutes.

Reimbursement: CPT 92229 describes retinal imaging with autonomous point-of-care analysis. The code does not itself establish universal Medicare coverage or a fixed national payment rate; Medicare payment is contractor-priced and coverage depends on applicable requirements (CMS, 2020).

Outcomes

What is established: - Prospective diagnostic-accuracy evidence supported a defined autonomous intended use - FDA created a De Novo classification and special controls for this device type - The workflow can produce a point-of-care result without specialist image interpretation - A dedicated CPT descriptor exists for autonomous retinal image analysis

What still requires implementation evidence: - Image quality requirements (dilated pupils, specific camera) limit accessibility - Not all healthcare systems adopted (cost, workflow integration) - Follow-up care coordination remains challenge (diagnosis without treatment pathway insufficient)

Lessons Learned

1. Prospective validation essential: IDx-DR’s success partly due to rigorous prospective trial before FDA submission, not just retrospective studies.

2. Addressing real clinical need: DR screening gap was clear, well-documented problem. AI solved actual bottleneck.

3. Autonomous use requires evidence matched to risk: The performance goals and special controls applied to this device and submission. They should not be generalized as universal FDA thresholds for all autonomous or decision-support systems.

4. Reimbursement matters: A code can support billing infrastructure, but it does not guarantee coverage, payment, adoption, or completed follow-up.

5. Implementation ≠ Impact: Authorization does not guarantee adoption. Workflow integration, training, and follow-up pathways remain critical to real-world benefit.


Case Study 2: Viz.ai Stroke Triage and Notification

Background

Large vessel occlusion (LVO) strokes require emergent thrombectomy (mechanical clot removal). Time-critical: “time is brain.” But identifying LVO on CT angiography (CTA) requires expert interpretation, often delayed in emergency settings.

Solution: Viz.ai LVO, AI analyzing head CTA to detect LVO and automatically alert stroke team (neurointerventionalists, neurologists) via mobile app.

Implementation

Development: Deep learning model trained on thousands of CTA scans, validated externally across multiple centers.

FDA authorization: The Viz LVO device received 510(k) clearance for a triage and notification intended use, not autonomous diagnosis (FDA K252970). Product version and intended use should be verified against the authorization record relevant to the deployment.

Deployment: Integrated into hospital PACS (picture archiving systems). When CT performed, AI analyzes automatically, flags suspected LVO, sends mobile alerts to stroke team with images.

Clinical workflow: The software analyzes head CTA and can notify the stroke team while formal interpretation proceeds. Actual processing time, integration, escalation, and fallback depend on the configured system and site.

Outcomes

Randomized evidence: A prospective stepped-wedge cluster-randomized trial of 243 thrombectomy patients found shorter adjusted door-to-groin and CT-to-treatment times after deployment. The trial did not find a significant improvement in 90-day functional independence (Martinez-Gutierrez et al., 2023).

Observational evidence: A meta-analysis of observational Viz.ai studies found workflow improvements but no statistically significant mortality difference (Sarhan et al., 2025).

Challenges: - False positives generate unnecessary alerts (stroke team fatigue) - Requires robust IT infrastructure, integration with PACS - Cost varies by institution (subscription model)

Lessons Learned

1. Triage/notification niche: AI does not replace radiologist. It flags urgent cases for expedited review. Less regulatory burden than autonomous diagnosis.

2. Time-to-treatment is an important process endpoint: Faster workflow is clinically relevant in stroke, but a process improvement should not be relabeled as demonstrated mortality or functional-outcome benefit.

3. Mobile integration key: Sending alerts to clinicians’ phones (not just radiologist workstation) enables rapid mobilization.

4. Multi-site evaluation improves transportability evidence: Generalizability still depends on the product version, population, imaging workflow, comparator, and implementation.

5. Measure outcomes, not just accuracy: The randomized trial measured workflow and 90-day outcomes, allowing the workflow benefit and null functional-outcome result to remain separate.


Case Study 3: Epic Sepsis Model External Validation

Background

Sepsis is leading cause of hospital deaths. Early recognition and treatment (antibiotics, fluids) save lives. Hospitals sought AI to predict sepsis risk, enabling proactive intervention.

Solution: Epic (major EHR vendor) developed sepsis prediction model (Epic Sepsis Model, ESM) integrated into EHR, alerting clinicians to high-risk patients.

Implementation

Development: Trained on Epic’s multi-institutional dataset. Deployed to hundreds of hospitals using Epic EHR.

Algorithm: Predicted sepsis risk based on vital signs, labs, clinical notes. Generated alerts when risk exceeded threshold.

Intended Use: Alert clinicians 6-12 hours before sepsis onset, prompt early intervention.

What Went Wrong

Independent study at Michigan Medicine: - External validation included 27,697 patients across 38,455 hospitalizations - At the evaluated threshold, hospitalization-level sensitivity was 33% and positive predictive value was 12% - 67% of sepsis cases never triggered an alert - Only 7% of patients with sepsis who did not receive timely antibiotics had an alert during the interval when antibiotics could have become timely. This is not the same as saying the model detected only 7% of sepsis before onset (Wong et al., 2021)

Reasons for Failure: 1. Training data issues: Model trained on sepsis cases identified retrospectively (billing codes), not real-time. Billing codes imperfect (miss cases, misclassify). 2. Look-ahead bias: Model may have learned from data collected after sepsis onset (e.g., labs ordered because sepsis suspected), inflating apparent performance. 3. Definition inconsistency: “Sepsis” defined differently across institutions. Model trained on one definition performed poorly when applied to others. 4. Alert fatigue: High false positive rate. Clinicians ignored alerts. 5. Lack of external validation before deployment: Epic deployed widely before independent, prospective validation.

Outcomes

What the study establishes: - External discrimination and alert performance were substantially poorer than the developer’s reported internal values - The evaluated threshold would have generated a large alert burden with low positive predictive value - The study was an external validation; it did not randomize deployment, quantify patient harm caused by the model, or establish a legal outcome

Field-level value of the evaluation: - Highlighted need for external validation, transparent reporting - Spurred regulatory discussion (FDA scrutiny of clinical decision support) - Motivated independent research on sepsis prediction

Lessons Learned

1. External validation before widespread deployment: Models performing well internally may fail externally. Do not deploy nationally without multi-site prospective validation.

2. Training data and labels are critical: Billing codes and retrospective definitions can be noisy. Prospective adjudication may improve label quality, but no single label method is a universal gold standard.

3. Beware look-ahead bias: Ensure model uses only data available at prediction time, not future data.

4. Transparency matters: Epic initially did not disclose algorithm details, validation data, making independent evaluation difficult. Secrecy erodes trust.

5. Alert burden must be measured: Low positive predictive value can create substantial false-alert burden, but the effects on clinician response and patient outcomes require direct study.


Case Study 4: Watson for Oncology and Variable Concordance

Background

IBM Watson, famous for winning Jeopardy! in 2011, was positioned as a major healthcare AI initiative. Watson for Oncology (WFO) was designed to combine patient information, literature, guidelines, and expert input to present treatment options.

Marketing boundary: Promotional descriptions should not be treated as evidence of clinical accuracy or patient benefit. Published evaluations mainly assessed concordance with local multidisciplinary recommendations.

Implementation

Development: Trained on Memorial Sloan Kettering Cancer Center (MSKCC) cases, expert oncologist input. Deployed internationally (India, Thailand, Korea, others) via partnerships with hospitals.

Intended Use: Oncologists input patient data, WFO recommends treatment options with supporting evidence.

What the Published Evidence Showed

Concordance evidence:

  1. A meta-analysis found that concordance varied substantially by cancer type and setting. The underlying studies were largely retrospective concordance analyses, not randomized tests of patient outcomes (Jie et al., 2021).
  2. A Chinese observational study found that concordance differed across cancer types and depended on how recommendations were categorized (Zhou et al., 2019).
  3. A recommendation that differs from a local multidisciplinary team is not automatically unsafe, and concordance is not the same as correctness. Each discordant case requires clinical adjudication against the contemporaneous evidence, patient characteristics, and available treatments.
  4. The published concordance literature supports concerns about transportability. It does not justify a single universal failure rate or a causal claim that WFO harmed patients.

Implementation questions: - Whether the recommendation set reflected local formularies, resources, guidelines, and patient preferences - Whether clinicians could understand the evidence basis and limitations of each option - Whether the interface fit multidisciplinary workflow - Whether licensing and integration costs produced measured value

Corporate history: IBM later divested healthcare data and analytics assets that became Merative. Corporate restructuring is relevant context, but it is not a peer-reviewed clinical endpoint.

Outcomes

What can be concluded: - Variable concordance limited confidence that one recommendation system would transfer unchanged across cancers and settings - The evidence base did not establish patient-outcome benefit - Procurement should have required local evaluation, transparent limitations, and economic measurement before scale

Field-level value: - Sobered hype around AI. Reminder that marketing ≠ clinical reality - Reinforced need for evidence, validation, transparency

Lessons Learned

1. Hype ≠ substance: Watson’s Jeopardy! success did not translate to medicine. Natural language processing of trivia questions fundamentally different from clinical reasoning.

2. Match data to the claim: Expert-authored cases can support knowledge engineering but cannot substitute for representative clinical data and patient-outcome evaluation when clinical benefit is claimed.

3. Domain expertise and local context are required: Oncology AI needs clinicians, patients, informaticians, methodologists, and local operational expertise. The relevant question is whether the development and deployment process represented the intended use, not a categorical judgment about one company’s expertise.

4. Validate in deployment populations: Variable concordance across settings is a transportability signal, not proof that every international deployment failed.

5. Usability matters: Even accurate AI is useless if clinicians will not use it. Workflow integration, user interface critical.


Case Study 5: COVID-19 Imaging Models and Methodological Failure

Background

During the early COVID-19 pandemic, diagnostic testing was constrained in many settings. Researchers rapidly developed models intended to diagnose or predict COVID-19 outcomes from chest radiographs and CT.

Rapid publication: Reported performance was often very high, but headline accuracy did not resolve spectrum bias, source confounding, leakage, or external-validity concerns.

Implementation

Development: Trained on publicly available datasets (CXR images labeled COVID-positive or negative).

Claims: “AI can diagnose COVID-19 from CXR, supplement scarce PCR testing.”

What Went Wrong

2021 systematic review (Roberts et al., 2021): - Identified 2,212 published papers and preprints, retained 415 after initial screening, and included 62 papers after quality screening - None of the models in the review were judged likely candidates for clinical translation in their reported form - Fifty-five of the 62 papers had high risk of bias in at least one PROBAST domain; the remaining seven were unclear in at least one domain

Common Failures: 1. Training data biases: COVID-positive images often from different sources (hospitals, equipment) than COVID-negative images. AI learned to detect image source, not disease. 2. Confounding: COVID patients often supine (portable X-rays), healthy controls standing (PA/lateral). AI detected positioning, not pathology. 3. Data leakage: Duplicate images in training and test sets, inflating performance. 4. Lack of external validation: Models tested on same dataset used for development. 5. Clinical implausibility: CXR findings in COVID-19 non-specific (similar to other viral pneumonias). Unrealistic to expect perfect discrimination.

Deployment boundary: The review evaluated published model-development literature through October 3, 2020. It did not establish that every later imaging model failed, identify every deployed product, or prove that deployed systems were quietly withdrawn.

Outcomes

Documented concern: - A large body of rapid research produced little evidence ready for clinical translation - Highly optimistic performance could be misread when methodological limitations were omitted - Poor reporting and nonrepresentative datasets made independent appraisal difficult

Field-level value: - Highlighted methodological failures in AI research - Motivated better standards (CONSORT-AI, TRIPOD-AI reporting guidelines) - Taught important lessons quickly (entire arc from hype to failure in ~18 months)

Lessons Learned

1. Speed vs. rigor: Pandemic urgency led to corner-cutting. Fast publication does not mean good science. External validation, rigorous methods still essential.

2. Beware confounders: AI learns shortcuts. Always consider: “What else could explain this association?”

3. Biological and clinical plausibility matter: A model may detect signals not consistently used by unaided readers, but such claims require rigorous controls against shortcuts and prospective external evaluation. Human difficulty alone neither proves nor disproves model utility.

4. Preregistration and transparency: Prespecifying the protocol and maintaining an untouched test set can reduce selective analysis. Preregistration does not by itself prevent bias or guarantee validity.

5. Independent review is critical: Clinical, statistical, and machine-learning expertise are all needed, together with reproducibility information and access to sufficiently detailed data provenance.


Case Study 6: Utah AI Prescription-Renewal Pilot

Background

Published U.S. estimates have attributed $100–289 billion in annual costs to medication nonadherence, with substantial variation in definitions and methods (Cutler et al., 2018). This burden estimate does not establish that prescription-renewal automation will improve adherence, outcomes, or cost.

Policy intervention: Utah’s Office of Artificial Intelligence Policy authorized a phased regulatory pilot with Doctronic for AI-supported renewal of existing prescriptions. Authorization of a monitored pilot is not evidence of autonomous clinical effectiveness.

Implementation

Regulatory framework: Utah’s Office of Artificial Intelligence Policy describes the governing relief, safety controls, and phase status on its authorized-pilot record. A separate January 2026 announcement provides launch context.

Scope: Renewal of existing prescriptions within an approved formulary. The official record excludes controlled substances, new prescriptions, changes in dosage or frequency, and products outside the formulary.

Safety controls:

  • In Phase 1, every renewal requires authorization by a licensed medical practitioner
  • Progression for a medication group requires 250 completed fills and approval by the Office of Artificial Intelligence Policy
  • Patients must verify Utah residency before using the system
  • Licensed professionals oversee implementation, and pharmacists can escalate to a licensed physician

Company-reported performance: A news report attributed 99.2% concordance between AI treatment plans and physician decisions to company testing (Deseret News, January 2026). The figure is not peer-reviewed patient-outcome evidence and should not be used as proof of safety or equivalence.

Company-reported price: Secondary reporting described an initial $4 renewal price. Pricing can change and does not establish cost-effectiveness after clinician review, pharmacy work, monitoring, and downstream care are counted.

Duration: The official pilot record lists October 2025 through October 2026, with an option for renewal. Public reporting should be evaluated when results are released.

Early Observations

Potential benefits:

  • May reduce some renewal delays if eligible patients complete the workflow
  • May change clinician and pharmacist workload, which should be measured rather than assumed
  • May provide a lower transaction price, although total episode cost and downstream effects remain unknown
  • May expand access hours, subject to the pilot’s eligibility and oversight requirements

Concerns raised:

  • AMA caution: AMA CEO John Whyte, MD stated that “without physician input [AI] also poses serious risks to patients and physicians alike” (Becker’s Hospital Review, January 2026)
  • Limited to routine renewals; does not address initial prescribing or complex medication management
  • Geographic limitation (Utah only) creates regulatory fragmentation
  • Long-term safety data not yet available

Lessons (Preliminary)

1. Regulatory sandboxes can support bounded learning: Utah’s framework creates a monitored, phased pathway with defined scope and professional oversight. It does not waive the need for transparent outcome reporting.

2. Narrow scope reduces risk: Limiting to routine refills of established medications for existing patients with chronic conditions is lower risk than new prescriptions or complex cases.

3. Practitioner authorization remains required in Phase 1: The 250-fill threshold does not itself activate autonomous operation. The Office must separately approve progression for each medication group.

4. Outcome tracking must match the policy question: Reporting should distinguish technical completion, licensed-practitioner agreement, pharmacy fulfillment, adherence, adverse events, access, equity, and total cost. Until results are available, these remain evaluation endpoints rather than demonstrated benefits.

5. This is a pilot, not a precedent (yet): Whether this model scales depends on outcomes data from the 12-month demonstration. Other states (Arizona, Texas, Delaware) have sandbox frameworks but have not yet approved similar programs.

Status: The official record describes an ongoing pilot and Phase 1 practitioner review. See Healthcare Policy and AI Governance for regulatory context and verify the official pilot page for current status.


Case Study 8: ChatGPT Health Emergency Triage in Simulation

Background

Product: ChatGPT Health (OpenAI consumer health feature, launched January 2026)

Premise: The feature allows users to ask health questions and may influence decisions about whether and when to seek care. That makes emergency escalation an important safety endpoint even when the product is not marketed as a formal triage system.

Study: Investigators evaluated 60 clinician-authored simulated vignettes across 21 clinical domains under 16 factorial conditions, producing 960 responses (Ramaswamy et al., 2026).

What Went Wrong

Emergency undertriage: In the clear-emergency condition, 33 of 64 responses were undertriaged (51.6%). This is a simulated-vignette safety result, not an observed rate of delayed treatment or patient harm.

Variation across conditions: The factorial design tested how response behavior changed across clinical and contextual conditions. It supports direct testing of escalation behavior rather than assuming a general-purpose model will respond consistently.

Mechanism remains unproven: The study measured outputs, not the model’s training objective or internal cause of each undertriage error. Proposed explanations such as conversational helpfulness, prompt sensitivity, policy tuning, or missing structured triage logic require separate testing.

Outcomes

  • The study documented a reproducible simulated-evaluation design and a substantial emergency-undertriage signal
  • It did not observe real patients, actual care-seeking behavior, delayed treatment, injury, or death
  • It did not determine the product’s FDA regulatory status, establish that a particular enforcement action was required, or evaluate a withdrawal decision
  • The findings support targeted post-deployment surveillance and transparent escalation testing for patient-facing systems

Lessons Learned

1. Consumer health AI is a distinct risk category. Unlike clinical AI deployed under physician oversight, consumer triage tools interface directly with patients making high-stakes decisions alone. The harm pathway is direct and unmediated.

2. General-purpose health assistance is not validated emergency triage. The study did not establish the system’s complete training corpus or prove whether it used named triage protocols. It did establish that emergency-escalation performance must be measured directly.

3. Regulatory status follows intended use and claims. Some consumer health functions fall outside device oversight, while others may meet the device definition. Absence from an FDA authorization list is neither proof of safety nor proof that a product is unlawfully marketed.

4. Escalation behavior requires an explicit objective and test plan. A patient-facing system should be evaluated for undertriage, overtriage, clarity, consistency, and safe handling of uncertainty rather than assumed to inherit these properties from general conversational quality.

5. Crisis and emergency pathways require dedicated safeguards. Appropriate safeguards may include validated routing logic, refusal or escalation policies, local emergency resources, human review, logging, monitoring, and clear limitations. The exact design should be tested for the intended population and setting.

See also: LLMs in Clinical Practice for expanded coverage of LLM safety limitations in clinical settings.


Case Study 9: Penda Health LLM-Supported Outpatient Care

Background

Many clinical-AI evaluations measure model accuracy in retrospectively assembled data. The Penda Health trial tested an LLM-based clinical decision-support system within routine outpatient care in Kenya, making it an important example of pragmatic evaluation outside a high-income academic health system.

Implementation and Design

A pragmatic cluster-randomized trial across 103 clinician clusters in a private outpatient network in Nairobi and Kiambu included 9,702 eligible encounters. The intervention supplied AI-supported clinical recommendations within the care workflow. The study prespecified clinical and process outcomes and reported competing interests and OpenAI in-kind support (Agweyu et al., 2026).

Outcomes

Fourteen-day treatment failure was 2.0% with usual care and 2.2% with AI-supported care (adjusted odds ratio 0.77, 95% CI 0.55–1.08; P=.13). The primary clinical endpoint did not significantly improve. Some process and documentation outcomes improved.

Lessons Learned

  1. Pragmatic trials are feasible: LLM support can be evaluated in routine care using prespecified clinical and process endpoints.
  2. Keep the primary result visible: Improved documentation or process measures should not be rewritten as improved clinical outcomes.
  3. Report context and interests: One private outpatient network does not establish transportability to other countries, public systems, inpatient care, or different model versions.
  4. Measure the full human-AI intervention: The result belongs to the model, interface, clinicians, training, escalation process, and local workflow together.

Questions Clinicians Ask

What is the Utah AI prescription pilot?

Utah authorized a phased pilot for AI-supported renewal of existing prescriptions. The pilot remains in Phase 1, in which every renewal requires authorization by a licensed medical practitioner; progression by medication group requires 250 fills and Office of Artificial Intelligence Policy approval.

Why did IBM Watson for Oncology fail?

Published Watson for Oncology studies mainly measured concordance with multidisciplinary recommendations, not patient outcomes. Concordance varied by cancer and setting, so the evidence does not support a single universal failure rate or causal claim about patient harm.

What makes medical AI implementations succeed?

Common success factors include a narrow intended use, evidence matched to the claim, external and prospective evaluation when appropriate, workflow integration, accountable clinical ownership, monitoring, and a functioning follow-up pathway.

How did IDx-DR receive FDA authorization?

IDx-DR underwent a prospective, single-arm diagnostic-accuracy study at 10 primary-care sites in 900 participants, reporting 87.2% sensitivity and 90.7% specificity for more-than-mild diabetic retinopathy before FDA De Novo authorization.

What lessons do medical AI failures teach?

Key lessons: internal validation alone is insufficient, high AUC does not equal clinical utility, workflow integration is critical, and marketing claims must be verified with outcome data.


Conclusion

These case studies reveal common themes:

Successes share: - Rigorous prospective validation - Addressing clear clinical needs - Appropriate deployment (triage/support vs. autonomous) - Transparency and external evaluation - Attention to workflow, usability, reimbursement

Failures share: - Inadequate validation (retrospective only, single-site, no external) - Confounders and biases in training data - Overfitting and poor generalization - Lack of transparency - Hype exceeding evidence

The difference between success and failure often is not algorithm sophistication. It’s methodological rigor, clinical grounding, and humility about limitations. These lessons apply to current and future medical AI development.