Emergency Medicine

Emergency departments present simultaneously favorable and difficult conditions for AI. Large volumes of structured data support time-sensitive detection, while time pressure, missing data, heterogeneous presentations, and immediate consequences make errors difficult to contain. Stroke imaging AI has randomized and observational workflow evidence, but patient-outcome benefit remains unproven for isolated imaging-alert systems. Sepsis AI remains oversold despite widespread deployment.

Learning Objectives

After reading this chapter, you will be able to:

  • Identify validated AI applications for emergency departments
  • Understand AI tools for ICU monitoring and early warning
  • Evaluate sepsis prediction systems critically
  • Assess stroke, trauma, and cardiac emergency AI tools
  • Navigate implementation challenges specific to acute care settings
  • Recognize failure modes in high-stakes environments
  • Balance speed, accuracy, and safety in AI-assisted acute care

The Clinical Context: Emergency and critical care present ideal and terrible conditions for AI simultaneously. Ideal conditions include large volumes of structured data (vitals, labs), clear time-sensitive outcomes (mortality, deterioration), and potential for AI to detect subtle patterns humans miss. Terrible conditions include time pressure that precludes verification, common missing data, extreme heterogeneity, false positives creating alert fatigue, and consequences of errors that are immediate and severe.

What Works Well:

Application Systems Evidence Level Key Benefit
LVO Stroke Detection Viz.ai trial and observational workflows Product-specific Shorter selected process intervals; functional-outcome benefit not established
ICH Detection Detection and worklist prioritization Product-specific Can accelerate review; sensitivity varies by study and product
PE Detection Detection and worklist prioritization Product-specific Standalone accuracy and treatment effects require separate evidence
Pneumothorax Detection and notification workflows Product-specific Some studies show strong detection and selected workflow effects
Trauma Imaging Fracture and injury detection Emerging and heterogeneous Verify anatomy, modality, age range, and intended use

What’s Problematic:

Application Concern Reality
Epic Sepsis Model 33% sensitivity (missed 67% of cases), 12% PPV Widely deployed despite poor validation (Wong et al., 2021)
Autonomous Triage Insufficient validation Too many edge cases for unsupervised use
Generic Early Warning Variable Alert burden and clinical utility depend on threshold and response pathway
ED-ICU delirium multi-agent prediction Abstract-only accuracies 0.749 / 0.731 / 0.670 Retrospective reports, not a CAM-ICU substitute (Shang et al., 2026)

Key Implementation Principles:

  1. Time is everything - AI must be faster than current workflow or provide substantial value
  2. Alert fatigue can defeat adoption - Measure alert burden per shift and define risk-based action thresholds locally
  3. Local validation required - Academic center performance ≠ your ED/ICU
  4. Clinical judgment irreplaceable - AI assists, never replaces physician assessment
  5. Responsibility must be defined - FDA authorization does not allocate liability; institutional policy, product labeling, professional duties, contracts, and jurisdiction all matter

The Bottom Line: Stroke imaging AI has the strongest emergency workflow evidence, but results remain product-specific and do not prove functional benefit. Sepsis and deterioration alerts require local validation, an actionable response pathway, and measured alert burden. Integration into the standard workflow is essential. A 2026 systematic review of open-ended LLM diagnosis and triage found extreme heterogeneity and almost no prospective diagnostic-accuracy studies; assisted use, not independent diagnosis, is the supported claim (Chen et al., 2026).


Introduction

Emergency medicine operates under constraints that define which AI tools can succeed: decisions are time-critical, patient presentations are often undifferentiated, and workflow interruptions directly impact care. These constraints create both opportunities and barriers for AI adoption.

The strongest evidence in emergency AI comes from selected stroke-imaging workflows, where a stepped-wedge cluster-randomized Viz.ai trial shortened defined treatment intervals without significantly improving 90-day functional independence (Martinez-Gutierrez et al., 2023). Other applications, particularly sepsis prediction, remain oversold despite widespread deployment.

High-Impact AI Applications in Emergency and Critical Care

1. Stroke Detection and Triage

Large Vessel Occlusion (LVO) Stroke Detection:

Systems: Viz.ai, RapidAI, Brainomix Regulatory status: Multiple product-specific authorizations; verify the exact device, indication, version, and user in the primary FDA record Function: Detect LVO on head CT angiography, alert stroke team immediately Viz.ai randomized evidence: A prospective stepped-wedge cluster-randomized trial at four stroke centers included 243 patients treated with endovascular thrombectomy. AI activation reduced adjusted door-to-groin time by 11.2 minutes and CT-to-treatment time by 9.8 minutes, without a significant improvement in 90-day functional independence (Martinez-Gutierrez et al., 2023) Viz.ai observational evidence: An 82-patient pre/post study found no statistically significant overall door-to-groin difference, while the off-hours subgroup improved from 157 to 95 minutes (Figurelle et al., 2023) Evidence boundary: A meta-analysis of 12 observational Viz.ai studies reported improved workflow intervals without a statistically significant mortality difference (Sarhan et al., 2025). RapidAI has separate observational evidence. These citations do not establish equivalent effects for Brainomix.

The randomized and meta-analytic evidence in this section is Viz.ai-specific.

How it works: 1. Patient gets head CTA 2. AI analyzes immediately 3. Positive LVO → Automated alerts to stroke team, neuroIR, neurosurgery 4. Team mobilized before radiologist reads 5. Faster door-to-groin time

Performance: No class-wide sensitivity should be assumed. Confirm the operating point and population for the exact product and version.

Limitations: - False positives (vessel anatomy mimics occlusion) - Small vessel occlusions may be missed - Requires CT angiography (not plain CT)

Intracranial Hemorrhage (ICH) Detection:

Systems: Aidoc, Viz.ai, RapidAI Regulatory status: Product-specific authorizations exist; the primary record controls Function: Flag ICH on non-contrast head CT, prioritize worklist Evidence: Worklist reprioritization cut median time to radiologist diagnosis of ICH on routine outpatient head CT from 512 to 19 minutes (Arbabshirani et al., 2018); sensitivity in that specific study was 73% (specificity 80%, AUC 0.85) Evidence boundary: The 73% sensitivity above belongs to one retrospective implementation study and should not be replaced by a class-wide claim without a defined systematic review, product, and operating point.

Use cases: - ED triage (identify critical studies) - ICU monitoring (post-procedure surveillance) - Trauma (rapid identification)

Limitations: - Small subarachnoid hemorrhages sometimes missed - Calcifications, artifacts can cause false positives - Subtle bleeds in posterior fossa challenging

2. Pulmonary Embolism (PE) Detection

CT Pulmonary Angiography AI:

Systems: Aidoc, Avicenna.AI Regulatory status: Product-specific authorizations exist; verify the exact detection or triage function Function: Detect PE on CTPA, triage positive studies Performance: No class-wide 90–95% sensitivity applies. Performance depends on embolus location, threshold, acquisition, population, and reference standard.

Potential workflow effects: - Earlier review of selected flagged studies - Worklist prioritization in busy EDs - Faster treatment requires a complete notification and response pathway and should not be inferred from standalone accuracy

Limitations: - Subsegmental PE (clinical significance debated, hard to detect) - Motion artifacts reduce accuracy - Chronic vs. acute PE differentiation imperfect

3. Sepsis Prediction and Early Warning

Epic Sepsis Model (Controversial):

Most widely deployed sepsis AI Evidence: MIXED AND CONCERNING

External validation (Michigan Medicine, Wong et al. 2021): - 33% sensitivity (missed 67% of sepsis cases) (Wong et al., 2021) - 12% PPV (88% of alerts were false positives) - Alert fatigue documented - Clinical benefit unproven

Why it struggles: - Sepsis definition ambiguous (clinical judgment, not algorithmic) - Confounding by treatment (sepsis suspicion → antibiotics → AI detects antibiotics, not sepsis) - Missing data patterns (vital signs checked more frequently in sick patients → AI learns correlation) - Real-time prediction harder than retrospective

Deployment reality: - Widely used despite limited validation - User trust low (many ignore alerts) - Some institutions disabled after poor performance

Alternative approaches:

SOFA/qSOFA scores enhanced with ML: - Continuous monitoring models - Performance differs by cohort, endpoint, and implementation; no universal superiority should be assumed - Still imperfect

Vital sign trajectory analysis: - AI detects subtle trends preceding deterioration - Promising but requires prospective validation

Clinical bottom line on sepsis AI: Require local validation, report positive predictive value and alerts per shift at the intended threshold, define the response pathway, and preserve clinician assessment. A risk alert is not a sepsis diagnosis.

Randomized dyskalemia-alert evidence: A pragmatic randomized trial of AI-enabled dyskalemia alerts had null coprimary treatment outcomes. Earlier hyperkalemia treatment appeared only in an AI-positive subgroup analysis and should not replace the overall null result (Lin et al., 2026). The trial was null on both coprimary outcomes.

Sepsis audit-and-feedback evidence: In a two-site cluster-randomized quality-improvement study, AI-supported abstraction and feedback increased SEP-1 compliance, largely through a fluid-related process and documentation component. ICU admission and mortality did not improve, and the study was retrospectively registered (Boussina et al., 2026). Process compliance and patient outcomes are separate endpoints.

4. Cardiac Emergency AI

ECG Interpretation AI:

Atrial Fibrillation Detection: - Apple Watch, Kardia, others - PPV approximately 84% for irregular pulse notifications confirmed as AFib (Perez et al., 2019) - ED challenge: High volume of patient-reported AFib alerts, many in asymptomatic individuals

The Apple Heart Study enrolled self-selected smartwatch users and evaluated a defined notification pathway. Its positive predictive value should not be generalized to every wearable, patient population, or ED presentation.

STEMI Detection: - AI analysis of 12-lead ECG can support recognition and communication, but evidence remains workflow-specific - In the ARISE cluster-randomized trial, the primary endpoint was door-to-balloon time. The AI workflow shortened treatment time. Cardiac death was a prespecified secondary signal (85 vs. 116; odds ratio 0.73; P=0.029), while all-cause mortality was not different (odds ratio 1.02; P=0.568) (Lin et al., 2025) - ARISE does not establish that AI reduced infarct size or all-cause mortality, and it should not be generalized to every ECG product or autonomous catheterization-laboratory activation.

Hidden MI patterns: - AI detects subtle STEMI equivalents - Posterior MI, Wellens syndrome - Comparative performance depends on the model, reader, case mix, and endpoint

Evidence: Improving rapidly, deployment expanding

Cardiac Arrest Prediction:

ICU deterioration models: - Estimate risk over a defined prediction horizon using available clinical data - A prediction has value only when it triggers an effective, timely, and proportionate response - Prospective evidence is mixed, and predictive accuracy alone does not establish fewer arrests or lower mortality

5. Trauma and Critical Injuries

Rib Fracture Detection (CT):

Systems: Aidoc, others Function: Detect rib fractures on chest CT Use case: Trauma, elderly falls Potential role: Flag suspected fractures for review. Evidence that this changes analgesia, complications, or disposition requires product-specific clinical evaluation.

C-Spine Fracture Detection:

Systems: Aidoc Function: Detect cervical spine fractures Use case: Trauma workup. The exact authorized anatomy, acquisition, age range, and intended use must be verified before deployment.

Pneumothorax Detection:

Systems: Oxipit, Lunit, Aidoc Function:** Detect PTX on chest X-ray or CT Use case: Trauma, post-procedure, ICU monitoring

Clinical impact: Some product-specific workflows can accelerate review or notification. Faster decompression and fewer complications require separate clinical evidence.

Intracranial injury triage (TBI):

Research aim: Estimate which patients may need neurosurgical evaluation or transfer. Predictive performance does not establish that the model safely directs transfer decisions.

Abdominal CT Multi-Condition Triage:

System: BriefCase-Triage: CARE Multi-triage CT Body (K252970, FDA cleared January 7, 2026) Function: Flags 11 specified findings on adult contrast or noncontrast chest, abdominal, or pelvic CT for concurrent triage and notification (FDA K252970 record) Use case: ED crowding and imaging backlogs where urgent abdominal findings wait hours in first-in, first-out reading queues Evidence boundary: The labeling reports indication-specific performance. One mean sensitivity or specificity should not be treated as the performance of every finding or local population. Prospective clinical utility remains unestablished. Workflow constraint: The system operates in parallel with standard review, does not remove or deprioritize studies, and its preview is not intended for diagnosis. See the Radiology chapter for the labeled condition list and limitations.

K252970 covers 11 specified triage findings, not 14 autonomous diagnoses.

Prehospital injury severity (PHISE):

ISS and NISS require diagnostic imaging that is not available on scene. A 7 August 2026 npj Digital Medicine derivation-and-external-validation study mapped AIS descriptions to an eight-region, four-level scene score (PHISE) in 397,864 NTDB patients and 14,136 TraumaRegister DGU patients (Sigle et al., 2026). Standalone PHISE correlated with ISS/NISS (Spearman ρ = 0.61 / 0.79 internally) but underperformed them for in-hospital mortality (AUC 0.71 vs 0.83 / 0.82 internally; 0.63 vs 0.81 / 0.84 externally) while approaching them for red-cell transfusion (0.77 vs 0.83 / 0.81 internally; 0.76 vs 0.78 / 0.77 externally). Embedding PHISE with prehospital vitals, mechanism, and age in XGBoost narrowed that gap on the paper’s reported train/test tables (mortality test AUC 0.89 vs ISS-ML 0.95; transfusion 0.82 vs 0.88), but this is registry derivation with LLM-assisted imaging-need labels, not a prospective EMS trial, and does not license replacing CT-based ISS or deploying PHISE as a standalone field-triage tool.

6. ICU Early Warning Systems

Patient Deterioration Prediction:

Continuous monitoring AI: - Analyzes vital signs, labs, medications in real-time - Predicts deterioration 6-24 hours ahead - Use cases: Ward→ICU transfer decisions, ICU resource allocation

Evidence: - Some RCTs show reduced code blues, ICU transfers - Other studies show no benefit (alert fatigue) - Variable results depend on implementation

Systems: - Epic Deterioration Index - WAVE Clinical Platform (ExcelMedical) - Various hospital-developed models

Challenges: - Rare outcomes can produce low positive predictive value even when discrimination appears strong - Clinicians ignore frequent alerts - Lack of actionable interventions for many alerts

Mechanical Ventilation AI:

Weaning prediction: - AI can estimate readiness for extubation - Effects on ventilator days and reintubation require prospective comparative evidence - Personalized recommendations require a validated, clinician-controlled protocol

Lung-protective ventilation: - Research systems can recommend PEEP or tidal-volume strategies - Recommendation quality and clinical benefit must be tested against protocolized care

Status: Research stage mostly, some commercial systems emerging

Acute Kidney Injury (AKI) Prediction:

Real-time AKI risk scores: - Estimate AKI risk over a defined prediction horizon - A useful alert must connect to safe, evidence-based review of fluids, hemodynamics, and nephrotoxic exposure

Evidence: Improving, prospective trials ongoing

ED-ICU delirium prediction (LLM multi-agent):

A 13 August 2026 Cell Reports Medicine study introduced DeLiriuMAgents, an LLM-driven multi-agent system that combines a machine-learning risk score with virtual emergency, neurology, and psychiatry reasoning plus retrieval-augmented generation to predict delirium in emergency critically ill patients (Shang et al., 2026). Published-abstract accuracy/sensitivity/specificity was 0.749/0.762/0.747 on MIMIC-IV internal validation, 0.731/0.708/0.736 on a two-hospital Peking University external cohort, and 0.670/0.708/0.665 on eICU-CRD. This is retrospective multi-cohort prediction with generated reports, not a prospective ED outcome trial and not a substitute for CAM-ICU or bedside assessment; cohort sizes are not in the published abstract, and eICU-CRD accuracy of 0.670 is modest.

7. ED Triage and Workflow Optimization

ESI (Emergency Severity Index) Augmentation:

Prospective evidence (NEJM AI, 2025):

A multisite quality-improvement study included 83,404 preintervention and 91,244 postintervention visits across three emergency departments. An AI-informed triage decision-support tool increased identification of critical-care patients from 78.8% to 83.1% and reduced median time from arrival to initial care by 33% (Taylor et al., 2025). The study was not randomized, so it does not establish a class-wide causal effect.

AI-assisted triage (general): - Predicts acuity, resource needs - Reduces undertriage - Variable evidence for benefit

A 2025 systematic review limited to prospective ED triage studies found only seven eligible studies from 1,633 screened records. Reported model accuracy ranged from 80.5% to 99.1%, but the studies varied in triage systems, model type, quality, and outcome measurement (Yi et al., 2025). The review supports AI as a triage aid, not autonomous triage, and reinforces that EDs should demand prospective local evaluation rather than relying on retrospective accuracy claims.

Algorithm-guided LLM versus routine triage:

A retrospective study of 1,960 adult ED visits compared algorithm-guided ChatGPT 5.2 (standardized prompt; off-the-shelf, not fine-tuned) with routine 5-level ESI-based triage and found 71.3% five-level accuracy, quadratic weighted Cohen’s κ of 0.824, and urgent-versus-non-urgent AUC 0.768 with sensitivity 0.630 and specificity 0.906 (Halici et al., 2026). Lower-acuity discordance versus routine triage (9.8%; higher-acuity 18.9%) was more frequent in older adults, trauma presentations, and diabetes; infectious presentations had the highest concordance (Halici et al., 2026). Substantial weighted κ is not clinical safety: 71% five-level agreement and 63% urgent sensitivity show the gap. Routine triage was an operational comparator, not a gold standard, so agreement with local practice is not accuracy, safety, or outcomes. Prospective outcome validation is required, and this study does not support autonomous LLM triage.

Pre-vital queue prioritisation:

Halici compared a single LLM with routine ESI on real visits; Sharma et al. add a different lesson those visits cannot teach: pre-vital queue prioritisation (a provisional ESI-scaled signal from symptoms alone), then a criterion-linked ESI once vitals exist. ED-Triage-Agent (five LangGraph agents; ESI Implementation Handbook v4 RAG) was calibrated on 30 Practice Cases and evaluated on Competency Cases plus the external TRIAGEAGENT benchmark, with Phase 1 exact-match 76.39% (κw = 0.8787; 95.52% high-priority sensitivity) and Phase 2 exact-match 87.04% (κw = 0.9090; 0.00% significant over-triage; 0.46% significant under-triage; 97.22% within ±1); a chain-of-thought baseline had 10.00% significant under-triage on Competency Cases versus 0.00% for ETA (Sharma et al., 2026). These figures are handbook-case and TRIAGEAGENT scores, not prospective ED outcomes, and they are not clinical safety or autonomy. The authors state that clinical utility, safety, and workflow still require prospective validation with real ED data.

LLM second-opinion differentials:

Brodeur et al. tested OpenAI o1 and GPT-4o against two attending physicians on 79 Beth Israel Deaconess emergency department cases at three predefined touchpoints: initial triage, emergency physician encounter, and admission or ICU transfer. The task was blinded second-opinion differential diagnosis generation from EHR text, not autonomous triage. o1 identified the exact or very close diagnosis in 65.8% of triage cases, 69.6% after the emergency physician encounter, and 79.7% at admission or ICU transfer, exceeding both physician comparators at each stage (Brodeur et al., 2026).

Clinical interpretation: This is a strong signal for physician-facing second-opinion support, especially when information is sparse. It should not be read as proof that LLMs can manage ED triage, disposition, procedures, or treatment without clinician oversight.

Leibovitch et al. ran a DECIDE-AI stage 1 evaluation of SHAKED, a multi-LLM clinical decision support system, across 1,138 patients in two parallel units of one live tertiary emergency department over 4 weeks. Adoption among eligible cases fell from 68% to 30%, with lower use later in each shift (odds ratio 0.72 per shift hour, 95% CI 0.62–0.83). Expert review rated 99 of 100 sampled outputs clinically appropriate, and length of stay was 4.9 hours in both wings (Leibovitch et al., 2026). Sustained clinician engagement and workload, not algorithmic accuracy, were the barrier; these findings inform randomized trial design and do not justify clinical deployment of AI CDSS at this stage.

A 2026 systematic review and meta-analysis of 50 studies (25 LLMs; open diagnostic generation only) reported standalone LLM versus clinician relative top-1 diagnostic accuracy 0.89 (95% CI 0.79–1.00), with I² 95–98%, and LLM-assisted clinicians versus clinicians alone 1.13 (1.00–1.27); triage accuracy was similar (1.01, 0.94–1.09) (Chen et al., 2026). Only 3 of 50 studies were prospective, and those were triage-only; standalone models were comparable to junior clinicians (0.95) and worse than senior clinicians (0.77) on top-1. This is a heterogeneous review funded in part by Tencent, not a warrant for independent LLM diagnosis.

Workflow benchmark update: ER-Reason was designed to test LLMs across linked ED workflow stages rather than isolated diagnostic questions. The benchmark includes 3,437 patients across 3,984 ER encounters, 25,174 de-identified longitudinal clinical notes, and 72 physician-authored rationales covering triage intake, EHR review, treatment planning, final diagnosis, and disposition. In baseline testing, o3-mini performed best on several tasks but still compressed acuity predictions toward “Urgent,” underestimated discharge decisions, and overestimated admissions. This supports the same operational conclusion: LLM output may help structure second opinions, but ED triage and disposition require local calibration and clinician control (Mehandru et al., 2025, preprint).

Consumer AI Triage Evaluation:

A structured stress test of ChatGPT Health used 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions, producing 960 responses. Among gold-standard emergencies, 52% were undertriaged. The most dangerous errors concentrated at the clinical extremes, and family or friends minimizing symptoms shifted some recommendations toward less urgent care (Ramaswamy et al., 2026). These were simulated vignettes, not observed patient outcomes.

Chest Pain Risk Stratification:

AI predicts 30-day MACE (major adverse cardiac events): - Models can estimate short-term risk, but comparison with HEART depends on the population, endpoint, missing data, and validation design - Reduced admissions and safe discharge are clinical-utility claims requiring prospective comparative evidence

Wait Time Prediction:

AI forecasts ED volume, wait times: - Potential input for staffing and patient communication - Forecast accuracy should be assessed at the operational horizon and during unusual surges

Disposition Prediction:

AI predicts admission vs. discharge: - Potential input for bed management and transfer coordination - A disposition prediction should not become a self-fulfilling clinical decision or encode historical access inequities

What Does NOT Work Well:

Autonomous triage without physician oversight: Too many edge cases and liability concerns

Consumer AI tools for patient triage guidance: ChatGPT Health undertriaged 52% of gold-standard emergency cases, downgrading them from “emergency” to “urgent,” with sociodemographic variation across race and sex subgroups (Ramaswamy et al., Nature Medicine, 2026)

Sepsis AI as standalone diagnostic: High false positives, missed cases, and need for clinical judgment

Alert systems without actionability: Warnings without clear intervention pathways create alert fatigue

Black-box predictions without explanation: Clinicians need to understand WHY a patient was flagged

One-size-fits-all thresholds: Optimal operating points vary by institution and patient population

Implementation Challenges Specific to Emergency/Critical Care:

1. Time Pressure: - AI must be FASTER than current workflow or provide substantial value - No time for complex interactions - Alerts must be actionable immediately

2. Incomplete Data: - ED presentations often lack history, prior records - Missing data common - AI must handle missingness gracefully

3. Heterogeneity: - Extreme patient diversity (age, acuity, comorbidities) - Undifferentiated presentations - AI trained on specific populations may fail

4. Alert Fatigue: - ICUs already have alarm overload - Adding AI alerts risks desensitization - Critical: Optimize thresholds for YOUR false positive tolerance

5. Workflow Integration: - Busy clinicians cannot switch to separate systems - Must integrate into EHR and monitor displays - Mobile alerts must reach the right people at the right time

6. Liability in High-Stakes Settings: - Missed diagnoses in ED/ICU carry high malpractice risk - Over-reliance on AI vs. under-utilization both risky - Documentation of AI use and overrides essential

7. Shift Work and Handoffs: - Multiple clinicians per patient - AI alerts must persist across handoffs - Continuity challenges

Deployment Practices for Emergency and Critical Care AI:

Implementation Checklist for Acute Care AI

Pre-Deployment: - LOCAL retrospective validation (YOUR patients, YOUR data) - Prospective silent-mode testing for a duration and sample size justified by volume, prevalence, uncertainty, and the consequence of error - False positive rate assessment (calculate alerts per shift) - Workflow mapping (where alerts go, who responds, what actions) - Clinical champion identification (ED/ICU physician leader)

Deployment: - Gradual rollout (pilot unit → full ED/ICU) - Threshold optimization (balance sensitivity vs. alert burden) - Mobile alert systems (push notifications to responsible clinicians) - Clear escalation pathways (what to do when AI flags patient) - Override mechanisms (clinicians can dismiss with documentation)

Post-Deployment: - High-frequency early monitoring, with cadence based on how rapidly harm could accumulate - User feedback collection (alert fatigue assessment) - False positive/negative tracking - Clinical outcome monitoring (does it improve patient outcomes?) - Continuous threshold adjustment (refine based on real-world performance)

Red Flags to Stop/Revise: - Alert response rates below the locally defined safety threshold - False-positive burden above the locally defined tolerance for the receiving team - User complaints escalating - Adverse events potentially related to AI (missed cases, over-reliance) - Performance drift detected (accuracy declining)

Evidence-Based Assessment by Application:

Relatively Mature Workflow Evidence (Still Requires Local Evaluation):

LVO stroke detection: Viz.ai has randomized and observational evidence for shorter selected workflow intervals, without established functional-outcome benefit. RapidAI and Brainomix require their own product-specific evidence.

ICH detection: Selected systems have detection and notification evidence. False negatives and queue displacement prevent treating the downside as inherently low.

PE detection: Product-specific standalone and workflow evidence exists, but treatment and outcome effects cannot be assumed.

Moderate Evidence (Deploy with caution, monitor closely):

Sepsis prediction: Mixed evidence, variable alert burden, and a need for local validation

Deterioration prediction: Variable results, implementation-dependent

Chest pain risk stratification: Promising but needs more validation

Weak Evidence (Pilot only, research stage):

Automated triage: Insufficient validation for autonomous use

Ventilator management: Early stage, more research needed

Most “AI-enhanced” early warning systems: Incremental benefit unclear

Special Considerations:

Pediatric Emergency/Critical Care:

Challenge: Most AI trained on adults Problem: Pediatric physiology, vital sign ranges, disease patterns differ Need: Pediatric-specific AI validation Current state: Limited pediatric AI available

Mass Casualty and Disaster:

Potential: AI-assisted triage, resource allocation Reality: Insufficient validation in disaster scenarios Concern: Undertriage of salvageable patients

Rural/Community EDs:

Challenge: Different patient populations, resources, workflows than academic centers where AI trained Need: External validation in community settings Transfer decisions: AI may help identify patients needing transfer to higher-level care

The Clinical Bottom Line:

Key Takeaways for Emergency and Critical Care
  1. Stroke AI has the strongest workflow evidence: Selected LVO and ICH systems can shorten defined intervals; patient-outcome benefit remains unproven for isolated alerts

  2. Sepsis AI is oversold: Epic model has major limitations, do not trust blindly

  3. Time savings matter most: AI must be faster or substantially better to justify use

  4. Alert fatigue is real: Optimize thresholds carefully, monitor response rates

  5. Local validation essential: Academic medical center performance ≠ your ED/ICU

  6. False positives are costly: Translate false-positive rate into alerts per shift using local prevalence and volume

  7. Clinical judgment irreplaceable: AI assists but does not replace physician assessment

  8. Integration is everything: Standalone systems will not be used in fast-paced environments

  9. Responsibility must be defined: Authorization alone does not allocate legal duties or civil liability

  10. Continuous monitoring required: Performance drifts, vigilance essential

  11. Communication matters: Alert right person, right time, right information

  12. Evidence hierarchy: Prospective trials > retrospective studies > vendor claims

Questions About Emergency Medicine AI

How much time does stroke AI save in the ED?

Viz.ai studies have reported shorter selected stroke workflow intervals, including an 11.2-minute adjusted reduction in door-to-groin time in a stepped-wedge cluster-randomized trial. That trial did not show a significant improvement in 90-day functional independence, and its results should not be transferred to other products.

What is the sensitivity of AI for detecting pulmonary embolism?

There is no class-wide sensitivity for pulmonary-embolism AI. Performance depends on the device, embolus location, threshold, acquisition, population, and reference standard. Standalone detection, worklist triage, faster review, and faster treatment are separate claims.

Should emergency departments use autonomous AI triage?

Current evidence supports clinician-facing triage assistance in defined workflows, not unsupervised replacement of emergency assessment. Local evaluation should measure undertriage, overtriage, time to care, subgroup performance, queue displacement, and patient outcomes.

What is the false positive rate for generic ED early warning systems?

There is no universal false-positive rate for emergency early-warning systems. Alert burden depends on prevalence, endpoint, threshold, data availability, workflow, and how repeated alerts are counted. Measure positive predictive value and alerts per clinician shift locally.

Future Directions:

The following horizons are planning scenarios, not forecasts.

Near-term scenario (1–3 years): - More stroke applications (wake-up stroke, hemorrhagic conversion prediction) - Better sepsis models (lower false positives, earlier prediction) - Expanded trauma AI (solid organ injury grading, hemorrhage prediction) - Real-time clinical decision support (integrated into EHR workflows)

Medium-term scenario (3–7 years): - Multimodal AI (vitals + labs + imaging + notes integrated) - Continuous learning systems (improve from local data) - Personalized risk prediction (accounting for individual patient factors) - Closed-loop systems (AI suggests intervention, monitors response)

Long-term scenario (7+ years): - AI copilots for emergency and critical care decision support - Autonomous monitoring systems (with human oversight) - Predictive resource allocation (anticipate surges, optimize staffing)

Across all horizons: Human expertise, clinical judgment, and defined institutional accountability remain central.

Next Chapter: We’ll explore AI in Internal Medicine and Hospital Medicine, where longitudinal data and chronic disease management create different opportunities and challenges.


Professional Society Guidance on AI in Emergency and Critical Care

Emergency Medicine Consensus and ACEP Resources

The 2026 multisociety emergency-medicine Statement of Principles on Artificial Intelligence addresses physician leadership, evidence, equity, transparency, privacy, education, and accountability (official PDF). It is a consensus statement, not a clinical practice guideline for a specific product.

The American College of Emergency Physicians has developed AI resources through its Research Committee’s Artificial Intelligence Subcommittee:

JACEP Open Primer (2025): “Artificial Intelligence in Emergency Medicine: A Primer for the Nonexpert” provides foundational guidance on AI applications in emergency medicine (Smith et al., 2025), including:

  • Triage system enhancement
  • Disease-specific risk prediction
  • Staffing needs estimation
  • Patient decompensation forecasting
  • Imaging interpretation assistance

AIIPEM Program: The Artificial Intelligence to Improve Performance in Emergency Medicine (AIIPEM) program focuses on optimizing the ED intake process through AI applications.

Key Guidance: ACEP emphasizes that AI integration should enhance emergency physician decision-making without replacing clinical judgment, particularly for time-critical conditions where AI false negatives could be catastrophic.

Society of Critical Care Medicine

SCCM has published educational content exploring AI applications in the ICU:

Current Applications Under Investigation:

  • Replacement of traditional monitoring systems with multidimensional pattern recognition
  • Enhanced clinical risk assessment tools
  • Efficient extraction and interpretation of clinical information
  • Predictive modeling for patient deterioration

Design Considerations:

The SCCM 2024 Guidelines on Adult ICU Design address the infrastructure in which digital monitoring and remote-care capabilities operate. Design guidance is not evidence that a particular AI system improves outcomes. The document notes that earlier ICU design eras did not envision: - Remote manipulation of ventilator settings - Remote infusion pump adjustments - AI-integrated monitoring systems

These capabilities are now being incorporated into modern ICU design standards.

Surviving Sepsis Campaign

The Surviving Sepsis Campaign 2021 guideline, endorsed by SCCM and ESICM, governs sepsis care rather than validating a specific AI alert. Its recommendations have implications for AI:

  • Sepsis screening algorithms should be validated locally
  • Electronic alert systems must balance sensitivity with specificity
  • AI predictions should support, not replace, clinical assessment of sepsis
  • Time-to-treatment metrics remain the focus, regardless of detection method

Hypothetical Decision Exercises

Fictional teaching cases

The three cases below are fictional decision exercises. Patient histories, product outputs, clinician actions, performance metrics, outcomes, policies, and legal arguments are illustrative unless a source is linked in the same sentence. They are not reported cases, predictions of litigation, or representations of a named product’s real-world behavior.

Scenario 1: Sepsis AI Alert Fatigue and Missed Diagnosis

Assume an emergency physician works a busy night shift at a 400-bed academic medical center whose EHR includes a generic sepsis-risk alert.

Your experience with the system: - First month: 3-5 sepsis alerts per shift - Majority are false positives (patients not septic) - Common false triggers: Febrile URI, dehydration, COPD exacerbation - You and your colleagues increasingly ignore the alerts

11:45 PM - New patient arrival: - Patient: 67-year-old woman - Chief complaint: “Weakness and confusion × 2 days” - Triage vitals: BP 108/62, HR 98, RR 18, T 37.8°C (100.0°F), SpO2 96% on RA - Triage note: Alert and oriented × 3, no acute distress

12:10 AM - Sepsis-risk alert fires: “HIGH RISK for sepsis. Consider sepsis workup and antibiotics.”

Your assessment: Patient looks okay, vitals not alarming. Probably another false positive. You’re managing 2 critical patients (STEMI, respiratory failure). You acknowledge the alert and plan to see patient when freed up.

1:30 AM - Nurse pages you: “Room 14 (the weakness patient) now BP 88/50, HR 115, more confused. Family says she’s not acting right.”

You reassess immediately: - Vitals: BP 85/48, HR 118, RR 24, T 38.2°C (100.8°F) - Exam: Confused, lethargic, poor skin turgor, no obvious source - Family: “She had UTI symptoms last week, didn’t want to see doctor”

Your workup: - Labs: WBC 18.5, lactate 4.2 mmol/L, Cr 2.1 (baseline 0.9) - Urinalysis: 100+ WBC, nitrite positive, bacteria - Diagnosis: Urosepsis with septic shock

You initiate sepsis bundle: - Fluid resuscitation - Blood cultures - Broad-spectrum antibiotics (2 hours after ED arrival)

ICU course: - Required vasopressors × 48 hours - AKI requiring temporary dialysis - Prolonged ICU stay (7 days) - Eventual recovery but new baseline kidney dysfunction

Fictional M&M review findings: - The sepsis-risk alert fired at 12:10 AM based on: - Elevated heart rate (98 vs. patient’s baseline 70s) - Subtle temp elevation - Elevated lactate (2.1 mmol/L on triage labs, before you saw patient) - Confusion (documented by triage nurse)

Committee question: “The AI correctly identified sepsis 1 hour 20 minutes before you initiated treatment. Why was the alert ignored?”

Question 1: What factors led to the delayed sepsis recognition despite AI alert?

Root causes of AI alert being ignored:

1. Alert fatigue from low local positive predictive value - A published external validation of one Epic Sepsis Model version reported 33% sensitivity and 12% positive predictive value (Wong et al., 2021). Those figures are background evidence, not the measured performance of the fictional alert. - Clinician experience: 3-5 alerts/shift, most not sepsis - Pattern learned: “Sepsis alerts are usually wrong” - Cognitive bias: Automation complacency → ignore frequent incorrect alerts

2. Non-specific presentation - Triage vitals borderline (not meeting classic SIRS criteria) - Initial lactate 2.1 (elevated but not dramatically) - Temperature initially normal-range (100.0°F) - Patient looked “okay” on initial assessment

3. Competing priorities - 2 critical patients requiring immediate attention (STEMI, respiratory failure) - Sepsis alert for stable-appearing patient deprioritized - Time pressure in busy ED

4. Alert design issues - Alert did not convey urgency effectively - No clear action pathway (“Consider sepsis workup” too vague) - Alert easily dismissed without forcing reassessment

5. Lack of trust in AI system - Previous false positives eroded confidence - No explanation provided for WHY patient flagged - “Black box” prediction without clinical reasoning

Question 2: Are you liable for the delayed sepsis treatment?

Decision and legal analysis:

No breach, causation finding, settlement, or verdict can be predicted from a fictional case. Legal duties are jurisdiction-specific and depend on the clinical record, hospital policy, expert evidence, and applicable law. The following are competing arguments to examine, not legal conclusions.

Sepsis guidance: - The Surviving Sepsis Campaign recommends immediate antimicrobials, ideally within one hour, for possible septic shock or a high likelihood of sepsis. For possible sepsis without shock, it recommends rapid assessment and administration within three hours if concern persists (Evans et al., 2021). - A guideline does not determine negligence or causation in an individual case.

Plaintiff’s argument:

  • “The AI system correctly identified sepsis at 12:10 AM”
  • “Dr. Smith ignored the AI alert for 1 hour 20 minutes”
  • “If antibiotics had been given at 12:10 AM instead of 2:00 AM, patient would not have required dialysis”
  • “Hospital implemented AI system but physician failed to act on alerts”
  • Damages: AKI requiring dialysis, prolonged ICU stay, permanent kidney dysfunction, pain and suffering

Defense arguments:

1. Standard of care is clinical judgment, not AI compliance: - AI is decision support, not diagnostic certainty - Physician must assess patient, not blindly follow algorithm - Initial presentation did not meet clinical criteria for sepsis (SIRS, qSOFA)

2. Recognition at reassessment (1:30 AM) was appropriate: - Initial vitals borderline, patient stable-appearing - When clinical deterioration occurred (hypotension, worsening mental status), sepsis recognized immediately - Treatment initiated within 30 minutes of deterioration

3. Competing priorities justified: - STEMI and respiratory failure patients were higher acuity - Resource allocation appropriate for ED triage

4. Causation uncertain: - AKI may have been present on arrival (Cr already elevated) - Earlier antibiotics may not have prevented dialysis need - Sepsis progression can be rapid despite treatment

Plaintiff’s rebuttal:

Hospital policy and implementation are relevant: - Hospital spent millions implementing system - Training emphasized following AI recommendations - Other EDs using system successfully - Hospital policy may state: “Respond to all sepsis alerts”

Competing priorities do not excuse delayed assessment: - Could have delegated initial assessment to resident, PA, or advanced practice provider - Could have re-triaged patient higher after alert - 1 hour 20 minutes too long to defer assessment

Questions that would influence review:

  • If hospital policy REQUIRES response to sepsis alerts: Stronger plaintiff case (policy violation)
  • If policy states alerts are “advisory only”: Stronger defense, clinical judgment prevails
  • Key factor: Was initial assessment reasonable given presentation?
    • Vitals not meeting sepsis criteria → defense stronger
    • Lactate 2.1 visible on chart → should have prompted earlier assessment

The adverse outcome and documented delay would warrant clinical and legal review, but neither establishes that an earlier response would have prevented kidney injury.

Lessons for risk management: - Acknowledge AND assess all high-risk alerts within defined timeframe - Document rationale if disagreeing with AI (e.g., “Assessed patient, does not meet sepsis criteria, will monitor closely”) - Re-triage patients when AI flags high-risk conditions

Question 3: How should sepsis AI be implemented to prevent this scenario?

Illustrative deployment options for sepsis AI:

These options require adaptation to local policy, staffing, evidence, and risk. They are not universal clinical or legal mandates.

1. Pre-Implementation Validation

LOCAL performance assessment (mandatory):

Run sepsis AI in silent mode for a duration and sample size justified by local volume, prevalence, uncertainty, and the consequence of error.

The following table is an illustrative governance worksheet, not a universal acceptance standard:

Metric Target Unacceptable
Sensitivity Locally justified target Locally defined stop rule
Specificity Locally justified target Locally defined stop rule
Positive Predictive Value Sufficient for the response burden Too low for safe sustained response
Alert rate Compatible with reliable action Exceeds receiving-team capacity

If performance unacceptable: Do not deploy OR adjust thresholds

2. Alert Design to Reduce Fatigue

Tiered alert system:

Illustrative high-priority tier: - Septic shock criteria (SBP <90 + 2 SIRS criteria + suspected infection) - Lactate >4 mmol/L - Alert: Page physician immediately, cannot dismiss without assessment

Illustrative moderate-priority tier: - 2 SIRS criteria + lactate 2-4 - Suspected infection + organ dysfunction - Alert: Task in EHR, reminder if not addressed

Illustrative lower-priority tier: - 1 SIRS criterion - Borderline labs - Alert: Passive flag in chart, no interruption

3. Actionable Guidance (Not Just Warning)

Poor alert: “HIGH RISK for sepsis. Consider sepsis workup.”

Better alert:

SEPSIS ALERT - Patient meets predictive criteria

Risk Score: 78% probability of sepsis
Key factors: Lactate 2.1, HR 98 (baseline 70), confusion, suspected UTI

RECOMMENDED ACTIONS:
☐ Reassess patient within 30 minutes
☐ Order sepsis labs if not done: CBC, CMP, lactate, blood cultures
☐ Consider empiric antibiotics if sepsis confirmed
☐ Acknowledge alert and document assessment

4. Clinical Decision Support Integration

Order set auto-population: - If sepsis alert fires, pre-populate sepsis workup orders (pending physician review) - One-click order placement (do not make physician manually enter 10 orders)

Documentation template: - Auto-generated sepsis assessment template in chart - Forces structured evaluation

5. Feedback Loop for Learning

Alert outcome tracking:

Every sepsis alert should be reviewed: - Was patient septic? (gold standard: physician diagnosis + antibiotics given) - If yes, was treatment timely? - If no, why false positive?

Share performance data with clinicians: - “Last month: 45 sepsis alerts, 18 true positives (PPV 40%)” - “Top false positive triggers: COPD exacerbation, dehydration” - Goal: Help clinicians calibrate trust in system

6. Threshold Optimization

Adjustable sensitivity:

Different EDs may prefer different operating points:

High-volume academic ED: Lower sensitivity (fewer alerts), higher PPV to reduce fatigue Community ED with limited backup: Higher sensitivity (catch more cases), accept lower PPV

Allow customization based on local performance and preferences

7. Escalation Pathway for Ignored Alerts

If an alert is not acknowledged within the locally defined response interval: - Escalate to charge nurse - Charge nurse assesses patient or ensures physician has seen - Prevents alerts from being lost

8. User Training (Essential)

All clinicians must understand: - How sepsis AI works (what inputs, what it’s predicting) - What to do when alert fires (assessment, workup, documentation) - AI is adjunct, not diagnostic truth - Clinical judgment can override AI (with documentation) - PPV expectations (e.g., “30% of alerts will be true sepsis”)

Key message: “AI helps you NOT MISS sepsis, but YOU decide if patient is septic

9. Audit and Accountability

Monthly review: - Sepsis cases missed by AI (false negatives) → Why? - Sepsis alerts ignored that were true sepsis → Why? - Trends in alert response rates

Individual feedback: - If physician repeatedly ignores alerts without documentation → coaching - If physician has better sepsis recognition than AI → learn from their practice

10. Vendor Accountability Questions

Before purchasing sepsis AI:

MUST ANSWER: 1. “What is PPV at 10% sepsis prevalence in ED population?” (not just AUC) 2. “Provide data from 3+ external validation sites (not just your development site)” 3. “What is alert rate per 100 ED patients?” 4. “How often do clinicians dismiss alerts at your deployment sites?” 5. “Provide prospective trial data showing improved outcomes (not just retrospective prediction)” 6. “What incident-response, investigation, documentation, and contractual support is available after a suspected miss?”

RED FLAGS: - Vendor cannot provide external validation data - Only reports AUC, not PPV/alert rate - Claims “90%+ accuracy” without defining what that means - Resists local validation period - No mechanism for threshold adjustment

Lesson: Sepsis AI with a low positive predictive value can create alert fatigue and unreliable response. Implementation should include local validation, actionable guidance, threshold selection, and continuous monitoring of both system performance and clinician response. Documentation should follow clinical relevance, product labeling, and institutional policy; it is not a guaranteed liability shield.

Scenario 2: Stroke AI False Positive and Unnecessary Thrombectomy

Assume an emergency physician works at a stroke center that uses a generic LVO notification system on CT angiography.

System track record: - Deployed 18 months ago - Generally excellent performance - Reduced door-to-groin time by 40 minutes on average - High staff satisfaction

2:30 AM - Patient arrival: - Patient: 58-year-old man - EMS report: Found by wife at 11 PM with slurred speech, right arm weakness - Last known well: 10 PM (4.5 hours ago) - NIHSS: 6 (moderate stroke severity)

2:35 AM - Head CT non-contrast: No hemorrhage

2:40 AM - CTA head and neck ordered

2:43 AM - LVO alert fires:

LARGE VESSEL OCCLUSION DETECTED
Vessel: Left M1 MCA occlusion
Confidence: HIGH
IMMEDIATE THROMBECTOMY CANDIDATE

Alert simultaneously sent to: - You (ED physician) - Stroke neurologist (Dr. Lopez) - Neurointerventional radiologist (Dr. Chen) - OR team

2:45 AM - Stroke team assembles

Dr. Lopez (neurologist) reviews patient: - NIHSS now 5 (mild improvement) - Right arm drift, mild dysarthria - Alert and cooperative

Dr. Lopez: “The alert says M1 occlusion. Let’s get him to the angio suite.”

You: “Should we wait for official radiology read?”

Dr. Lopez: “The system has performed well in our recent cases. Time is brain. Let’s go.”

Dr. Chen (neuroIR) reviews CTA images on mobile device: “I see the M1 cutoff. Looks like LVO. Let’s take him.”

3:00 AM - Patient to angio suite

3:15 AM - Groin access, catheter advanced

3:25 AM - Dr. Chen performs angiogram:

Finding: Left M1 appears patent on angiogram. No occlusion visible.

Dr. Chen: “This is strange. The CTA definitely showed cutoff, but angiogram shows flow. Maybe it recanalized?”

Dr. Chen performs thrombectomy attempt anyway (already committed, patient under anesthesia):

Result: No clot retrieved. Vessel appears normal.

3:45 AM - Procedure concluded

4:00 AM - Overnight neuroradiologist (Dr. Patel) reads CTA (official report):

IMPRESSION:
1. No large vessel occlusion identified
2. Left M1 segment demonstrates atherosclerotic narrowing but patent
3. Apparent "cutoff" on CTA likely artifact from patient motion + atherosclerotic calcification
4. Small lacunar infarct left corona radiata (chronic, not acute)

CONCLUSION: No acute LVO. CTA findings likely motion artifact.

Patient outcome: - Thrombectomy complications: Groin hematoma requiring compression, contrast-induced AKI (Cr 1.1 → 2.4) - NIHSS improved to 2 by morning (likely TIA or minor stroke, not LVO) - Discharged day 3 with residual mild weakness - Follow-up: Angry about “unnecessary procedure,” considering legal action

Question 1: What went wrong in this case?

Root causes of the fictional false-positive pathway:

1. AI misclassification - CTA artifact: Patient motion + atherosclerotic calcification mimicked M1 occlusion - Training and testing distribution: Motion artifacts can challenge imaging algorithms, but the fictional system’s development data are unspecified - False positive: System flagged stenosis + artifact as complete occlusion

2. Over-reliance on AI without independent verification - Misuse of an accuracy summary: No single percentage captures sensitivity, specificity, predictive value, or performance in an artifact-rich case - No independent radiology confirmation before thrombectomy - Neuroradiologist read would have identified artifact (did identify, but after procedure)

3. Cognitive biases - Automation bias: Trusting AI over clinical judgment - Confirmation bias: Dr. Chen “saw” occlusion on CTA because AI said it was there - Sunk cost fallacy: Once in angio suite, proceeded with thrombectomy despite normal angiogram

4. Time pressure overriding verification - “Time is brain” urgency led to skipping official radiology read - Valid concern for true LVOs, but prevented error detection

5. Lack of protocol for AI-physician discordance - What if angiogram does not match CTA? No clear pathway for this scenario - Should have aborted thrombectomy when angiogram showed patent vessel

Question 2: Who is liable for the unnecessary thrombectomy?

Decision and legal analysis:

No breach, liability allocation, settlement, or verdict can be predicted from this fictional case. The following arguments identify issues for review; they are not statements of governing law.

Standard of care for LVO stroke: - Mechanical thrombectomy proven for LVO strokes up to 24 hours (DAWN, DEFUSE-3 trials) - CTA is standard imaging for LVO detection - Thrombectomy should be performed rapidly when LVO confirmed

Plaintiff’s argument:

  • “Doctors performed invasive procedure based solely on AI, without radiologist confirmation”
  • “Angiogram showed NO occlusion, yet they attempted thrombectomy anyway”
  • “If they had waited 15 minutes for official read, would have avoided unnecessary procedure”
  • “Resulted in groin hematoma, kidney injury, unnecessary anesthesia risk”
  • Damages: Procedural complications, AKI, emotional distress, medical bills

Defendants:

  1. Dr. Lopez (neurologist) - Ordered thrombectomy based on AI
  2. Dr. Chen (neuroIR) - Performed procedure, continued despite normal angiogram
  3. You (ED physician) - Raised concern but deferred to specialists
  4. Hospital - Implemented AI system, protocols

Defense arguments:

1. Defense position on rapid intervention: - “Time is brain.” Every minute delay causes more infarction - Waiting for official read would delay treatment - CTA showed apparent occlusion (artifact mimicked occlusion convincingly)

2. Fictional local track record: - The system had performed well in the center’s prior cases, but this does not establish the standard of care - Prior 30 cases at this hospital were all correct - Reasonable to trust system

3. Angiogram discordance addressed appropriately: - When angiogram showed no occlusion, Dr. Chen did NOT force stent retriever - “Thrombectomy attempt” was diagnostic angiography - Procedure aborted when no clot found

4. Complications minor and resolved: - Groin hematoma treated conservatively - AKI resolved (Cr returned to normal) - No permanent harm

Plaintiff’s rebuttal:

AI is adjunct, not diagnostic gold standard: - The plaintiff could argue that a qualified physician should interpret the CTA before an irreversible intervention - Whether a particular delay is reasonable depends on the local workflow and clinical facts

Angiogram showed no occlusion, yet the procedure continued: - The plaintiff could argue that further intervention after a normal angiogram lacked justification - The defense would need to explain what was performed and why

Informed consent inadequate: - Patient not told “AI detected occlusion but not confirmed by radiologist” - Patient consent assumed confirmed diagnosis

Questions that would influence review:

  • Whether the institution required separate imaging confirmation
  • What occurred after the initial angiogram and whether it was clinically justified
  • Whether the complications were caused by an avoidable intervention

Relevant facts include: - Hospital’s AI protocols (do they require radiology confirmation?) - Severity of AKI and whether permanent - Patient’s residual deficits from stroke itself

Expert evidence would address whether the interpretation, consent, and procedure met the applicable standard of care.

Question 3: How should stroke AI be implemented to prevent false positive procedures?

Deployment safeguards for LVO detection AI:

1. Understand AI as Triage Tool, Not Diagnostic Certainty

Notification-system role: - Triage: Prioritize worklist, mobilize team - Notification: Alert stroke team rapidly - Time-saving: Reduce door-to-groin time

A notification system is not: - Diagnostic confirmation - Replacement for radiologist interpretation - 100% accurate

Performance interpretation: - Use sensitivity, specificity, and predictive values from the exact device, version, population, and operating point - Predictive value changes with LVO prevalence and referral selection - Do not convert a sensitivity estimate into a presumed false-positive fraction

Clinical implication: The local team should know how many alerts are confirmed, how many are false positives, and which artifacts recur.

2. Verification Protocol Before Thrombectomy

Recommended workflow:

STEP 1: LVO notification fires
↓
STEP 2: Stroke team mobilizes (appropriate, saves time for true LVOs)
↓
STEP 3: While patient moved to angio suite, SIMULTANEOUS:
  - Neurologist examines patient
  - A qualified imaging physician reviews the CTA while the team prepares
  - Anesthesia preps patient
↓
STEP 4: CONFIRMATION REQUIRED before groin puncture:
  ☐ A qualified physician confirms the imaging finding under the institution's stroke protocol
  ☐ Neurologist confirms clinical syndrome consistent
  ☐ Time window appropriate (within 24 hours for confirmed LVO)
↓
STEP 5: Proceed with thrombectomy

Key principle: Mobilization based on AI, but INTERVENTION based on physician confirmation

3. Neuroradiology Confirmation Protocol

Rapid physician review for stroke: - Route the alert immediately to the qualified imaging and stroke clinicians named in the local protocol - Define and audit a turnaround target that preserves timely treatment - Complete physician image review before an irreversible intervention - Parallel transport and preparation can reduce delay

Confirmation checklist:

Neuroradiologist must confirm:
☐ Large vessel occlusion present (not artifact)
☐ Vessel identity correct (M1 vs M2 vs ICA vs basilar)
☐ No contraindications visible (hemorrhage, mass, old infarct)
☐ Collateral flow assessment
☐ Clot burden estimation

Neuroradiologist signs off: "Confirmed LVO, safe to proceed"

4. Handling AI-Angiogram Discordance

Protocol for discrepant findings:

If CTA (confirmed by radiologist) shows LVO, but angiogram shows patent vessel:

Possible explanations: 1. Spontaneous recanalization (happens in ~20% of LVOs) 2. CTA artifact (false positive) 3. Technical issue with angiogram

Illustrative decision logic:

The treating stroke and neurointerventional team must apply the actual protocol and procedural findings. The diagram is a teaching aid, not a treatment order.

Angiogram shows NO occlusion:

→ If neurologic improvement (NIHSS decreased): STOP procedure
   - Likely spontaneous recanalization or false positive
   - No benefit to thrombectomy if vessel already open

→ If neurologic stable/worsening: Repeat angiography, different angles
   - May be technical miss
   - If still no occlusion visible: STOP procedure

→ NEVER perform thrombectomy on angiographically patent vessel

5. Informed Consent and Diagnostic Uncertainty

Consent discussion should include:

“Imaging shows what appears to be a blocked blood vessel in your brain. An AI system detected this and alerted our team. Our radiologist is reviewing the images now to confirm. If confirmed, we recommend a procedure to remove the clot. This can significantly improve outcomes, but carries risks including bleeding, stroke, and groin complications. Do you have questions?”

Key elements: - Explain the imaging finding, uncertainty, proposed intervention, alternatives, and time sensitivity - Procedure risks explained - Time-sensitive decision

Whether the use of AI itself requires disclosure is jurisdiction- and context-specific; the chapter does not impose a universal script.

6. Audit and Feedback

Track all LVO alerts: The values below are fictional examples of an audit table.

Month Alerts Confirmed LVO False Positives PPV Thrombectomies Clot Retrieved
Jan 12 9 3 75% 9 8 (89%)
Feb 10 7 3 70% 7 7 (100%)

Review false positives: - Why did AI misclassify? - Common artifacts causing false positives - Share with team to improve recognition

Review false negatives (missed LVOs): - Were there clinical clues AI missed? - Should have been escalated despite negative AI?

7. Vendor Accountability

Questions for any LVO AI vendor:

  1. “What is PPV at 10% LVO prevalence?” (not just sensitivity)
  2. “What percentage of your alerts at other sites are false positives?”
  3. “What are most common causes of false positives?” (motion artifact, atherosclerosis, etc.)
  4. “Do you recommend radiologist confirmation before thrombectomy, or is AI alone sufficient?”
  5. “What incident-response, investigation, documentation, and contractual support is available after a suspected false positive?”
  6. “Provide data on thrombectomies performed that retrieved no clot (suggests false positive)”

RED FLAGS: - Vendor claims “no need for radiologist confirmation” - Vendor cannot provide PPV data - Vendor dismisses false positives as “rare” - Vendor resists post-market surveillance audits

8. Team Training

All stroke team members must understand:

The notification system is a triage aid, not diagnostic certainty Local predictive value and recurring failure modes must be measured Radiologist confirmation required before thrombectomy Angiogram overrides CTA if discordant Clinical judgment can override AI (document reasoning)

Scenario-based training: - “Viz.ai says M1 occlusion, radiologist sees artifact. What do you do?” - “Angiogram shows patent M1 despite CTA occlusion. Proceed or stop?”

Lesson: Stroke-notification AI can shorten selected workflow intervals, but false positives occur and product-specific outcome benefit remains unproven. The local protocol should require qualified physician review before an irreversible intervention. When angiography contradicts CTA, the team should reassess the imaging, clinical course, and possibility of recanalization or artifact rather than treating the alert as diagnostic truth.

Scenario 3: ICU Early Warning System and Code Blue

Assume an intensivist works in a 24-bed medical ICU that deployed a generic early-warning system six months earlier.

System description: - Analyzes vital signs, labs, medications, nursing assessments in real-time - Generates risk score 0-100 (higher = greater risk) - Alerts when score crosses thresholds: 50 (moderate), 70 (high), 90 (critical)

Your experience: - 10-15 alerts per shift (24-bed ICU) - Most alerts are patients you’re already managing (already in ICU, on pressors, etc.) - Rarely actionable (patient already receiving maximum care) - You and ICU team have learned to mostly ignore alerts

2:00 PM - You’re managing: - 3 post-op patients on ventilators - 2 septic shock patients on 3 pressors each - 1 ARDS patient on ECMO - Multiple floor patients awaiting ICU transfer (no beds available)

2:15 PM - Deterioration alert:

PATIENT: Jackson, Robert (ICU Bed 12)
AGE: 72
DETERIORATION INDEX: 78 (HIGH RISK)
PREDICTED RISK: Cardiac arrest within 6 hours
RECOMMENDATION: Assess patient urgently

You review chart: - Patient: 72-year-old man, post-op day 3 after colectomy for colon cancer - Current status: Extubated yesterday, doing well, off pressors - Vitals (last 2 hours): BP 118/70, HR 88-95, RR 16-20, SpO2 96-98% on 2L NC - Labs (this morning): WBC 11.5, Hgb 9.2 (stable post-op), Cr 1.1, K 3.8 - Nurse note (1 hour ago): “Patient comfortable, tolerating clear liquids, ambulated to chair”

Your assessment: Looks fine, probably false positive. Patient clearly improving, not deteriorating.

You acknowledge alert, no action taken.

4:45 PM - You’re in family meeting for ECMO patient

4:50 PM - Overhead page: “CODE BLUE, ICU BED 12”

You run to Bed 12:

Finding: Mr. Jackson unresponsive, pulseless

Nurse: “I was checking on him, he looked fine 10 minutes ago. Then alarms went off. V-fib on monitor!”

Code Blue team initiates ACLS: - CPR started - Defibrillation × 2 - Epinephrine, amiodarone given

5:05 PM - ROSC achieved (return of spontaneous circulation)

Post-code workup: - Stat labs: K 6.9 mmol/L, Mg 1.2, pH 7.18, lactate 8.2 - ECG: Peaked T waves (hyperkalemia) - Review of vitals trend: - 2:00 PM: HR 88 - 2:30 PM: HR 92 - 3:00 PM: HR 95 - 3:30 PM: HR 102 - 4:00 PM: HR 108 - 4:30 PM: HR 118 - 4:45 PM: V-fib arrest

Cause identified: Hyperkalemic cardiac arrest

Root cause investigation: - Patient has chronic kidney disease (baseline Cr 1.4) - Post-op, placed on IV fluids containing potassium - Morning labs: K 3.8 (low-normal) - Replacement order: KCl 40 mEq IV × 2 doses (given at 8 AM, 12 PM) - Renal function declined: Post-op AKI (Cr 1.1 → 1.8 by afternoon, not yet resulted in chart) - Potassium accumulated: K 3.8 → 6.9 over 6 hours

Fictional retrospective analysis of the alert: - Why did AI flag patient at 2:15 PM? - Subtle upward trend in heart rate (88 → 95) - Decreased urine output (30 mL/hr last 2 hours) - Potassium replacement orders in chart - Post-op patient with CKD - AI predicted deterioration 2.5 hours before cardiac arrest

Patient outcome: - Survived cardiac arrest - Post-arrest care in ICU - Anoxic brain injury (prolonged downtime before code called) - Neurologic prognosis uncertain - Family considering withdrawal of care

M&M Committee Review:

Committee: “The AI correctly predicted cardiac arrest 2.5 hours in advance. Why was the alert ignored? If potassium had been checked at 2:15 PM when alert fired, hyperkalemia would have been detected and treated, preventing the arrest.”

Question 1: What factors led to the ignored early warning alert and subsequent cardiac arrest?

Root causes:

1. Alert fatigue from poor PPV - 10-15 deterioration alerts per shift in 24-bed ICU - Most alerts for patients already critically ill (already maximal care) - Alert system “crying wolf.” Clinicians habituated to ignore

2. Alert timing and context - Alert fired for post-op patient who appeared stable - Recent vitals reassuring (BP 118/70, SpO2 96%) - Nurse assessment 1 hour ago: “comfortable, improving” - Cognitive dissonance: Alert says “high risk,” eyes say “patient looks fine”

3. Lack of actionable guidance - Alert said “Assess patient urgently” but did not suggest WHAT to assess - No specific recommendation (e.g., “Check potassium level”) - Unclear what intervention would address “cardiac arrest risk”

4. Competing priorities - ICU at capacity, multiple critical patients requiring attention - Family meeting in progress when alert fired - Stable-appearing patient deprioritized

5. System limitations not well understood - Clinicians did not understand WHY patient flagged - “Black box” prediction without explanation - If alert had said “Risk factors: K replacement + declining UOP + CKD → check K level,” might have prompted action

Question 2: Are you liable for failing to act on the AI alert?

Decision and legal analysis:

No breach, causation finding, settlement, or verdict can be predicted from this fictional case. The following are arguments for review, not statements of governing law.

Standard of care for ICU monitoring: - Intensivists must monitor for patient deterioration - Timely response to changes in clinical status - Electrolyte monitoring for at-risk patients (CKD, K replacement)

Plaintiff’s argument:

  • “Hospital deployed AI early warning system to prevent exactly this type of event”
  • “AI correctly predicted cardiac arrest 2.5 hours early”
  • “Dr. Anderson acknowledged alert but took no action”
  • “If potassium level had been checked at 2:15 PM, hyperkalemia would have been identified and treated”
  • “Cardiac arrest and brain injury were preventable”
  • Damages: Anoxic brain injury, prolonged ICU stay, likely death or severe disability, pain and suffering

Defense arguments:

1. Clinical assessment at 2:15 PM was reasonable: - Patient appeared stable (normal vitals, comfortable, improving post-op course) - No clinical signs of hyperkalemia at that time - Physician assessed risk and determined patient stable

2. AI early warning systems have high false positive rates: - Most alerts do not result in deterioration - Physician must exercise clinical judgment, not blindly follow algorithm - Standard of care is clinical assessment, not AI compliance

3. Hyperkalemia was unpredictable: - Morning K level was low-normal (3.8), appropriately repleted - Renal function decline not yet evident (afternoon Cr not resulted) - Rapid K accumulation unusual

4. Resuscitation response: - Code team responded appropriately - ROSC achieved within 15 minutes - The adequacy of the response would require review of the full record

Plaintiff’s rebuttal:

Hospital chose to deploy this AI system: - Hospital invested in the fictional deterioration index for early intervention - Training emphasized following AI recommendations - If physician routinely ignores alerts, why have system?

AI identified specific risk 2.5 hours early: - Alert was NOT for already-critical patient (patient was stable post-op) - AI detected subtle pattern (trending HR, UOP decline, K replacement orders) - Reasonable physician would have checked K level in this context

Monitoring after potassium replacement: - The plaintiff could argue that kidney disease, replacement dosing, and urine-output change required earlier electrolyte reassessment - The defense could contest when the renal decline was knowable and what monitoring interval was reasonable

Questions that would influence review:

  • Whether the arrest was preventable and whether the alert contained clinically actionable information
  • The severity and cause of the neurologic injury
  • Hospital policy question: Did hospital policy require response to high deterioration alerts?
    • If YES → stronger plaintiff case (policy violation)
    • If NO → stronger defense (alerts advisory only)

Expert evidence would address whether and when electrolyte reassessment was required given the patient’s renal function, replacement dosing, urine output, and evolving clinical information.

The severity of injury would make the case consequential, but the vignette cannot establish that the alert identified hyperkalemia, that a particular response was required, or that earlier testing would have prevented the arrest.

Question 3: How should ICU early warning systems be implemented to be useful, not just noisy?

Illustrative deployment options for ICU deterioration AI:

These options require local validation and governance; their thresholds and response times are not universal standards.

1. Optimize Thresholds to Reduce Alert Fatigue

Problem: A default threshold can generate more alerts than the receiving team can reliably evaluate

Solution: Site-specific threshold tuning

Run AI in silent mode for a locally justified period and sample size. The table below uses fictional values to illustrate threshold tradeoffs:

Threshold Alerts/Day True Deteriorations PPV Alert Fatigue Risk
Score >50 45 8 18% VERY HIGH
Score >70 18 7 39% HIGH
Score >85 6 5 83% MODERATE

Choose threshold that balances: - Sensitivity (catch deteriorations) - PPV (avoid alert fatigue)

For ICUs: The appropriate threshold depends on the endpoint, prevalence, monitoring environment, response burden, and consequences of false negatives and false positives.

2. Actionable, Specific Alerts (Not Generic Warnings)

Poor alert:

DETERIORATION INDEX: 78
Cardiac arrest risk high
Assess patient urgently

Better alert:

DETERIORATION INDEX: 78
Cardiac arrest risk: 15% within 6 hours

KEY RISK FACTORS:
• Heart rate trending up (88 → 102 over 2 hours)
• Urine output declining (30 mL/hr × 2 hours)
• Potassium replacement orders + CKD history
• Post-op day 3 (risk period)

SUGGESTED ASSESSMENTS:
☐ Check stat basic metabolic panel (K, Cr, Mg)
☐ Review fluid balance and UOP trend
☐ Assess for occult bleeding (post-op)
☐ Consider EKG if electrolyte abnormalities

Key improvements: - Quantified risk (15% not just “high”) - Explanation (why flagged) - Specific actions (check K level, not just “assess”)

3. Integrate Alerts into Workflow (Not Separate System)

Alert delivery: - In-basket task in EHR (not just pop-up that can be dismissed) - Cannot be cleared without documentation: “Assessed patient, [findings], [plan]” - Escalation: If not addressed in 1 hour, alert charge nurse

Order set integration: - Alert includes one-click order for suggested workup - Example: “Order stat BMP for Deterioration Alert” (pre-populated order)

4. Contextualize Alerts (Filter Out Already-Managed Patients)

Avoid alerting for: - Patients already on maximum ICU care (3 pressors, ECMO, etc.). You already know they’re high-risk - Patients with comfort-measures-only status - Patients actively being managed for deterioration

DO alert for: - Stable-appearing patients with subtle trends - Post-op/post-procedure patients (often lower acuity but can deteriorate suddenly) - Patients on general ICU monitoring (not already high-intensity care)

5. Feedback Loop and Continuous Learning

Track all alerts:

Patient Alert Time Score Assessed? Action Taken Outcome
Jackson, R 2:15 PM 78 No None Arrest 4:50 PM
Smith, J 2:30 PM 72 Yes Checked labs, normal No deterioration
Lee, K 3:00 PM 81 Yes Transfused, transferred Stabilized

Monthly review: - True positives: Alerts that preceded deterioration → Learn what worked - False positives: Alerts that did not lead to deterioration → Adjust thresholds - False negatives: Deteriorations not predicted → Improve model

Share with team: - “Last month: 42 alerts, 18 true deteriorations (PPV 43%)” - “Top reasons for true positives: post-op AKI, sepsis, arrhythmia” - Goal: Help clinicians calibrate when to trust vs. question alerts

6. User Training and Expectations

All ICU clinicians must understand:

What deterioration AI predicts (arrest, transfer, mortality) How it works (what inputs, what patterns) What to do when alert fires (specific assessments, not just “look at patient”) Expected PPV (e.g., “40% of alerts will be true deteriorations, 60% false”) Physician judgment overrides AI (but must document rationale)

Key principle: “AI helps you catch SUBTLE trends you might miss, but YOU decide what to do

7. Protocol for High-Risk Alerts

Illustrative protocol for a locally defined high-risk threshold:

Possible actions, with timing set by local policy and clinical urgency: 1. Bedside assessment by physician or advanced practice provider 2. Vital signs recheck 3. Review I/O, medications, recent labs 4. Stat labs if risk factors suggest (e.g., K replacement + CKD → check K) 5. EKG if cardiac arrest risk 6. Document findings and plan in chart

If alert seems inappropriate: - Document why (e.g., “Patient extubated, ambulating, tolerating diet. Alert appears false positive, will monitor”) - DO NOT simply dismiss without assessment

8. Vendor Accountability

Questions for any deterioration AI vendor:

  1. “What is PPV for cardiac arrest prediction at 1% base rate?” (not just AUC)
  2. “What alert rate per 100 ICU patient-days do you recommend?”
  3. “How many sites have reported alert fatigue and stopped using the system?”
  4. “Provide prospective trial data showing reduced code blues or mortality” (not just retrospective prediction)
  5. “Can thresholds be customized per institution?”
  6. “What explanations does system provide for WHY patient flagged?”

RED FLAGS: - Vendor cannot provide site-level PPV data - One-size-fits-all thresholds (no customization) - No prospective outcome trials - Black-box predictions without explanations - Alert rates that exceed the locally defined response capacity

Lesson: Early-warning AI may detect subtle deterioration patterns, but low positive predictive value or excessive alert burden can undermine response. Effective implementation requires locally justified thresholds, actionable content, workflow integration, user training on expected predictive value, and a defined response protocol. Explanations can help users investigate an alert, but an explanation does not make the prediction correct.