Cardiology and Cardiothoracic Surgery
Cardiology has used automated signal interpretation for decades. Newer models can detect ECG patterns associated with low ejection fraction, hyperkalemia, and other phenotypes, but detection performance and clinical utility remain separate claims. Heart-failure prediction and wearable AF detection create management questions that are not answered by model discrimination alone.
After reading this chapter, you will be able to:
- Evaluate AI systems for ECG interpretation and arrhythmia detection, including FDA-cleared algorithms and emerging applications
- Critically assess AI applications in echocardiography, cardiac MRI, and coronary CT angiography, including opportunistic extra-cardiac screening claims
- Understand heart failure prediction models and their clinical limitations, including false positive rates
- Analyze wearable device AI for atrial fibrillation detection and cardiovascular monitoring
- Recognize opportunistic retinal imaging as an AF-risk marker rather than a stand-alone AF test
- Recognize major failures in cardiovascular AI, including IBM Watson for Oncology’s cardiology applications
- Apply evidence-based frameworks for evaluating cardiology AI tools before clinical adoption
- Navigate medico-legal implications of AI-assisted cardiovascular decision-making
Introduction: Cardiovascular AI’s Promise and Pitfalls
Cardiology generates more structured data than perhaps any other medical specialty. Every heartbeat produces electrical signals. Every cardiac cycle can be imaged with ultrasound, MRI, or CT. Decades of epidemiologic studies have linked cardiovascular biomarkers to outcomes in millions of patients.
This data richness makes cardiology theoretically well suited to AI applications. Automated ECG interpretation predates the current deep-learning era, while newer systems learn representations from ECG waveforms, imaging, longitudinal records, and wearable signals.
The evidence is uneven. Some systems have product-specific regulatory records and prospective evaluations. Others remain retrospective research models or proprietary risk scores without a tested response pathway. Wearable AF detection also creates management questions that detection studies alone cannot answer.
Understanding which tools have earned clinical trust, which have failed, and how to evaluate new offerings requires examining the evidence across each application domain.
Part 1: ECG Interpretation AI
The 30-Year History You Did Not Know About
If you’ve ordered an ECG in the past two decades, AI has already interpreted it. The automated interpretation printed at the top of every ECG (“Sinus rhythm,” “Acute anterior STEMI,” “Left ventricular hypertrophy”) comes from algorithms developed in the 1980s-1990s and refined over millions of ECGs.
These are not new AI tools. They’re three-decade-old expert systems and pattern recognition algorithms that have become so ubiquitous that we forget they’re algorithmic at all.
Performance of traditional ECG algorithms: - Automated interpretation performance varies by rhythm, abnormality, population, acquisition quality, and algorithm. A systematic review describes persistent error and the need for clinician over-reading rather than a single class-wide accuracy figure (Schläpfer & Wellens, 2017). - STEMI, atrial fibrillation, and left-ventricular-hypertrophy outputs should therefore be evaluated as separate tasks.
These systems are familiar and widely integrated, but ubiquity is not proof that every output or implementation is reliable.
But they have limitations: - Sensitivity-specificity tradeoffs: STEMI algorithms optimized for sensitivity (to avoid missing MIs) produce false positives that experienced clinicians routinely override - Population-specific performance: Algorithms trained on predominantly white populations show worse performance in Black patients for LVH detection (Sokolow-Lyon criteria) - Over-reading and under-reading: Automated interpretations sometimes flag “abnormal ECG” for clinically insignificant findings while missing subtle ST changes that experienced cardiologists detect
The clinical lesson is durable: automated ECG interpretations require review in the context of the tracing, symptoms, prior studies, and clinical urgency.
Implementation Reality: Why Accurate Algorithms Still Miss STEMIs
ECG algorithms achieve >90% sensitivity for STEMI detection. So why do hospitals still miss STEMIs?
The implementation failures: 1. Alert fatigue: False positive STEMI alerts (especially in patients with old infarcts, LBBB, LVH) cause providers to ignore or delay response to true positives 2. EHR integration problems: STEMI alerts buried in EHR notifications alongside medication warnings and “patient census updated” messages 3. Workflow design failures: ECG interpretation printed on paper that does not trigger emergency response protocols 4. Over-reliance on AI: Providers skip careful ECG review because “the computer would have caught it”
A multicenter study of STEMI detection across 7 U.S. emergency departments found: - 12.8% of STEMIs were missed by ED screening criteria and did not receive an ECG within 15 minutes of arrival - Performance varied dramatically by site: 3.4% to 32.6% miss rates across facilities - Cause: Variation in triage protocols and ECG screening criteria implementation (Yiadom et al., 2017)
The lesson: Implementation > accuracy. A 95% accurate algorithm is worthless if clinicians do not see or trust its alerts.
ARISE Trial: AI-ECG Improved Treatment Time
The ARISE cluster-randomized trial evaluated an AI-ECG alert workflow for suspected STEMI. Its primary endpoint was door-to-balloon time, not mortality (Lin et al., 2024).
Design: Cluster RCT, 43,234 patients at Tri-Service General Hospital, Taipei, Taiwan; randomized 1:1 to AI-ECG-assisted STEMI detection versus standard care.
Results: - Median door-to-balloon time: 82.0 vs 96.0 minutes (P = 0.002) (14-minute reduction) - ECG-to-balloon time: 78.0 vs 83.6 minutes (P = 0.011) - Cardiac deaths: 85 vs 116 cases (OR 0.73, P = 0.029), a prespecified secondary signal - All-cause mortality: OR 1.02 (P = 0.568), a null result
Why this matters: Most AI-ECG studies report diagnostic performance. ARISE tested a clinical workflow and improved the prespecified treatment-time endpoint. The cardiac-death result is important but should not be promoted to a proven mortality benefit when all-cause mortality was null and post hoc ejection-fraction and biomarker outcomes were not different.
Limitations: Single-center cluster RCT in Taiwan; generalizability to U.S. emergency department workflows requires confirmation. The AI algorithm used (BioSignal) is not currently FDA-cleared for STEMI identification in the U.S.
Part 2: Cardiac Imaging AI
Echocardiography: Where AI Actually Helps
Echocardiography is operator-dependent, and measurement variability can affect serial interpretation. AI systems target acquisition guidance, view classification, segmentation, measurement, and study-level interpretation.
AI-assisted echocardiography addresses these problems:
Research evidence and authorized products must be separated: - EchoNet-Dynamic is a published research model for video-based ejection-fraction estimation and cardiomyopathy assessment, not an FDA clearance record (Ouyang et al., 2020). - Commercial acquisition and analysis products have different intended uses. Each retained product claim should be linked to its exact FDA record rather than inferred from a research paper or another product in the class.
Intraprocedural AI guidance (March 2026): Philips’ EchoNavigator R5.0 received FDA 510(k) clearance (K253614) as automated radiological image processing software for structural heart procedures (FDA 510(k) database). The DeviceGuide feature uses AI to track and visualize Edwards PASCAL Ace mitral transcatheter edge-to-edge repair (M-TEER) devices in real time by combining live echocardiography with fluoroscopy in a single procedural view (Philips, March 2026). This represents cardiology AI moving beyond offline interpretation into intraprocedural guidance for technically demanding structural heart interventions. The physician remains in control; the AI is an assistive visualization layer.
EchoNet-Dynamic in a randomized interpretation workflow: In a blinded randomized trial of 3,495 transthoracic echocardiograms, cardiologists made substantial changes to the initial LVEF assessment in 16.8% of the sonographer-initial studies and 27.2% of the AI-initial studies, meeting the prespecified noninferiority endpoint and supporting an assisted interpretation-workflow claim (He et al., 2023). The trial did not test patient outcomes or authorize autonomous interpretation.
The associated EchoNet-Dynamic dataset includes 10,030 labeled echocardiogram videos, which supports reproducible research and external method comparison. Dataset availability does not eliminate differences in acquisition protocols, populations, labels, or clinical workflow.
Opportunistic AI can also read kidney disease risk from heart ultrasounds. Yuan, Ouyang, and colleagues trained a deep learning model on more than 300,000 parasternal long-axis echocardiogram videos and detected chronic kidney disease with external validation at Stanford and Kaiser Permanente Northern California (Yuan et al., 2026). Screening discrimination in echo-referred cohorts: not a replacement for eGFR or proof that echo-AI CKD alerts improve outcomes.
Echocardiography Foundation Models
Single-task echo AI is giving way to foundation models that interpret complete studies across multiple views, closer to how cardiologists read echocardiograms. Three models define the current landscape.
EchoPrime was the first large-scale vision-language foundation model for echocardiography, published in Nature (Vukadinovic et al., 2025).
Training scale: EchoPrime was trained on 12,124,168 echocardiogram videos paired with cardiologist reports from 275,442 studies across 108,913 patients at Cedars-Sinai Medical Center. This is over 10 times the data used to train previous echocardiography foundation models like EchoCLIP.
Multi-site validation: Performance was validated across four external health systems: Stanford Healthcare, Beth Israel Deaconess (MIMIC), Chang Gung Memorial Hospital (Taiwan), and Kaiser Permanente. Mean AUC ranged from 0.85 to 0.92 across 17 classification tasks.
Key innovations: - Multi-view synthesis: Unlike single-view models, EchoPrime integrates information from all echocardiogram videos in a study using an anatomical attention module that weights views based on the cardiac structure being assessed - Retrieval-augmented interpretation: Generates comprehensive reports by matching input videos to similar historical cases, enabling prediction of hundreds of pathologies across 15 anatomical sections - View classification: Identifies 58 standard echocardiographic views with AUC of 0.997
Performance vs. task-specific models: - LVEF estimation: MAE 4.79% (vs. 4.1% for EchoNet-Dynamic), with 2% improvement in R² score - Aortic regurgitation detection: AUC 0.88 (vs. 0.68 for EchoCLIP) - Mitral regurgitation: 4% improvement over EchoNet-MR - Agreement with cardiologists (balanced accuracy 0.89) comparable to inter-cardiologist agreement (0.82)
Potential clinical role: EchoPrime may support study-level representation and preliminary assessment, but retrospective diagnostic performance does not establish safe autonomous reporting, improved access, or patient benefit.
Limitations: - Excludes spectral Doppler and M-mode images (relies on video encoder) - All validation was retrospective; prospective clinical trials comparing EchoPrime reports to cardiologist interpretations are ongoing - Performance in point-of-care ultrasound settings (lower image quality) requires further study
Open resources: Code and model weights are publicly available at github.com/echonet/EchoPrime, enabling external validation and research use.
PanEcho takes a different approach: supervised multitask learning across 39 reporting tasks, trained on over 1 million echocardiographic videos at Yale-New Haven Health System. Published in JAMA, PanEcho is view-agnostic (processes any echo video without requiring view identification) and achieves median AUC 0.91 across 18 classification tasks with LVEF MAE 4.2% on internal data, 4.5% externally (Holste et al., 2025). Code available at github.com/CarDS-Yale/PanEcho.
EchoJEPA challenges a fundamental assumption in echo AI: that models should learn by reconstructing masked pixels. In ultrasound, pixel-level reconstruction forces models to memorize stochastic speckle noise, because faithful reconstruction requires reproducing it. EchoJEPA instead predicts embeddings (learned representations) of masked video regions from visible context, using a slowly-updating teacher network that reinforces only temporally coherent structures like chamber geometry and wall motion, not random speckle. This latent prediction approach, built on Meta’s V-JEPA2 architecture (Assran et al., 2025), is pretrained on 18 million echocardiogram videos across 300,000 patients (Munim, Fallahpour, Szasz et al., 2026, preprint).
The controlled comparison is the strongest evidence: when EchoJEPA and a pixel-reconstruction baseline (VideoMAE) are trained with identical architecture, data, augmentations, and compute, changing only the training objective, latent prediction reduces LVEF error by 27% (MAE 5.97 vs. 8.15) and improves view classification by 45 percentage points. The full model (ViT-Giant, 1.1B parameters) achieves LVEF MAE 4.26 on internal data and 3.97 on the public EchoNet-Dynamic dataset.
Two results have technical implications. Under physics-informed perturbations simulating depth attenuation and acoustic shadows, the preprint reports 2.3% degradation for EchoJEPA versus 16.8% for EchoPrime. It also reports 79% view-classification accuracy with 1% of labeled data versus 42% for the comparison baseline at 100% (Munim, Fallahpour, Szasz et al., 2026, preprint). These are benchmark results, not evidence that the model improves care for patients with genuinely poor acoustic windows or can be deployed with minimal local annotation.
Caveats: EchoJEPA is a preprint; EchoPrime (Nature) and PanEcho (JAMA) are peer-reviewed. The strongest results come from a proprietary 18M-video corpus; the public model (EchoJEPA-L, 525K MIMIC-IV-Echo videos) shows smaller advantages. Robustness was tested with synthetic perturbations, not prospective data from patients with genuinely poor acoustic windows. The evaluation uses EchoJEPA’s standardized frozen-backbone protocol across all models; head-to-head comparisons using each model’s own evaluation paradigm have not been performed.
What AI does well in echo: 1. Automated endocardial border detection: Traces LV cavity more consistently than manual tracing 2. View optimization: Guides sonographer to acquire standard views correctly 3. Quantitative measurements: Chamber volumes, wall thickness, valve areas more reproducibly than manual calipers 4. Strain analysis: Automated global longitudinal strain calculation (time-consuming manually)
What AI does not do well yet: - Complex valve pathology: AI struggles with multiple jets, eccentric regurgitation, prosthetic valves - Technically difficult studies: Poor acoustic windows, obesity, COPD remain challenging for most echo AI. Latent prediction architectures (EchoJEPA) show improved robustness to acoustic degradation in synthetic testing, but prospective validation in genuinely difficult patients is pending - Novel findings: AI detects what it was trained to detect; will not identify rare pathology
Implementation boundary: Automated measurements and decision support require product-specific evidence, review of the authorized use when applicable, cardiologist oversight, and local subgroup monitoring. Pooled performance should not substitute for evaluation in the population and acquisition protocols where the tool will be used.
AI-Augmented ATTR-CM Detection: ATTRACTnet
Transthyretin amyloid cardiomyopathy (ATTR-CM) is substantially underdiagnosed. Effective disease-modifying therapies (tafamidis, acoramidis) have changed the prognosis for diagnosed patients, making earlier detection a clinical priority. Yet ATTR-CM is often missed because its presentation overlaps with hypertensive heart disease and other causes of left ventricular hypertrophy.
The ATTRACTnet model integrates four data sources available in the EHR: ECG waveforms, echocardiographic measurements, patient demographics, and ICD diagnosis codes for orthopedic manifestations of amyloidosis (carpal tunnel syndrome, spinal stenosis, biceps tendon rupture). Published in JAMA Cardiology in February 2026 (Jain et al., 2026), this is among the first multisite prospective clinical trials of an AI screening program for a specific cardiomyopathy diagnosis.
Study design:
The trial was conducted as a nonrandomized, single-system, multisite, single-arm, open-label prospective study. ATTRACTnet was trained at a large academic ATTR-CM referral center and externally validated at a second academic site. Patients with LV wall thickness 12 mm or more and an ATTRACTnet score of 0.5 or higher were eligible for the intervention. Exclusions included prior ATTR-CM testing, hypertrophic cardiomyopathy, advanced dementia, and nursing home residence.
Model performance:
- Area under the ROC curve: 0.85 (internal set, 5-fold cross-validation range 0.77–0.85); 0.82 (external test set, 95% CI 0.81–0.83)
- Performance was similar across Hispanic, non-Hispanic Black, and non-Hispanic White patients, addressing a key limitation of many cardiac AI models
Prospective screening results:
During the study period, 1,471 patients had positive ATTRACTnet scores (0.5 or above). Of 256 patients who met eligibility criteria, 50 underwent nuclear scintigraphy and monoclonal protein testing after physician and patient agreement. 24 of 50 tested patients (48%) were diagnosed with ATTR-CM, and 21 of those 24 (88%) initiated treatment within 3 months. The positivity rate was more than 2.8 times higher than historical controls (15.3%; 95% CI, 13.1%–17.9%; P < .001), with an 18% relative increase in new ATTR-CM diagnoses compared with the prior year (Jain et al., 2026).
Limitations:
- Nonrandomized design; no control group receiving usual care in parallel
- Testing uptake was low: only 50 of 256 eligible patients (20%) proceeded to testing, reflecting physician gatekeeping and patient hesitancy
- No randomized trial yet to determine whether AI-augmented screening improves survival or hospitalization outcomes compared to usual care
- Training data from a specialized academic referral center may limit generalizability to community practice
Clinical implications:
ATTRACTnet identified confirmed ATTR-CM among selected AI-positive patients who completed testing, and most diagnosed patients initiated treatment. Without a concurrent randomized control, the study does not establish how many would otherwise have remained undiagnosed or whether the workflow improves clinical outcomes. The trial’s editorialists noted that prospective randomized evidence is needed before the approach becomes standard of care (JAMA Cardiology editorial, November 2025).
For cardiologists evaluating AI screening programs for ATTR-CM, the key questions remain: what screening intervals are appropriate, how to handle patients with positive AI scores whose physicians decline testing, and whether the approach scales to community echocardiography laboratories without specialist referral centers.
Cardiac MRI and CT: Technical Excellence, Clinical Validation Pending
AI for cardiac MRI and CT shows impressive technical performance but lacks the decades of clinical validation that ECG algorithms have.
Applications: - Automated segmentation: LV/RV/atrial volume calculation from cine MRI - Perfusion defect detection: Stress MRI ischemia analysis - Coronary CT angiography analysis: Stenosis grading, plaque characterization, FFR-CT
Incidental coronary calcium: In an observational cohort, coronary artery calcium quantified by deep learning on routine nongated chest CT was associated with mortality and adverse cardiovascular outcomes. The study did not establish that AI-guided preventive treatment improves outcomes (Peng et al., 2023).
Prognostic evidence: Perivascular fat attenuation index (FAI) uses CCTA attenuation in pericoronary fat as an imaging biomarker of coronary inflammation. In the ORFAN multicentre longitudinal cohort of 40,091 patients undergoing clinically indicated CCTA, an FAI-based AI-Risk model was externally validated for cardiovascular risk stratification in patients with and without obstructive coronary artery disease. The observational study did not test whether FAI-guided care reduces cardiovascular events (Chan et al., 2024).
Evidence boundary: Technical performance depends on the task, reference standard, and dataset. It should not be summarized as a class-wide correlation with expert readers.
Problems: 1. Clinical utility remains unproven: Observational CCTA cohorts provide prognostic information, but whether management based on these outputs improves patient outcomes remains unknown. 2. Vendor lock-in: Most algorithms proprietary, embedded in scanner software, cannot be independently validated 3. Overdiagnosis risk: Highly sensitive algorithms may detect “abnormalities” of uncertain clinical significance 4. Value uncertainty: Procurement cost and clinical benefit over standard CCTA are product- and setting-specific.
Clinical bottom line: Use AI-assisted cardiac MRI/CT for efficiency (automated measurements save radiologist time), but do not change clinical management based on AI findings without expert review.
Part 3: Heart Failure Prediction and the Threshold Problem
Heart failure readmission prediction is a classic AI overpromise story.
The pitch: “Our proprietary machine learning algorithm predicts 30-day HF readmission with AUC 0.85! Identify high-risk patients for intensive case management!”
The reality: An AUC of 0.85 does not determine the positive predictive value or false-positive burden at the intended operating point. Those depend on prevalence, sensitivity, specificity, threshold selection, and the population in which the model is used.
Illustrative threshold calculation, not a measured model result: - Population: 1,000 HF discharges - Actual readmissions: 200 (20% rate) - Algorithm at 80% sensitivity: Detects 160/200 true positives - Assume for illustration that the selected threshold also flags 600 of 800 patients who will not be readmitted - Result: 760 patients flagged, of whom 160 (21%) actually readmit
Why this matters: If a local intervention were assumed to cost $500–$1,000 per patient, treating 760 flagged patients would cost $380,000–$760,000 before accounting for implementation. Those figures are scenario assumptions, not published cost-effectiveness results, and the number of preventable readmissions would still have to be measured.
Many models use readily available demographic, comorbidity, laboratory, and utilization variables. Angraal et al. developed a HFpEF model in a cohort that was 88% White and identified limited diversity as a generalizability concern (Angraal et al., 2020). This does not prove that every heart-failure model relies on the same proxies or adds no information.
A complex model should be compared with a simpler, interpretable baseline in the same cohort. Complexity is not evidence of incremental clinical value.
CardioMEMS demonstrates clinical utility where EHR-based prediction models do not. CardioMEMS (implantable pulmonary artery pressure sensor) reduced HF hospitalizations by 39% over a mean 15-month follow-up in the CHAMPION RCT (Abraham et al., 2011, Lancet). But this is a device that enables early intervention based on hemodynamic data, not a prediction algorithm based on EHR data.
Clinical bottom line: Be skeptical of HF readmission prediction algorithms. Ask: 1. “What is the false positive rate at the sensitivity threshold you recommend?” 2. “What interventions will we apply to algorithm-flagged patients, and what’s the evidence those interventions prevent readmissions?” 3. “How does this algorithm perform in our specific patient population?” (Most are validated only in development cohort)
HFpEF sacubitril phenomapping (PARAGUIDE):
PARAGON-HF did not show an overall benefit of sacubitril/valsartan versus valsartan on cardiovascular death or total HF hospitalizations (RR 0.90, 95% CI 0.77–1.06). A 12 August 2026 npj Digital Medicine phenomapping analysis of 4,161 PARAGON-HF patients used 53 baseline characteristics and each patient’s 20% phenotypic neighborhood to compute a phenotype-specific response score (median -0.62, IQR -1.21 to 0.05); 70.8% of patients had PRS ≤ 0 and that selected subset showed an 18% risk reduction (RR 0.82, 95% CI 0.69–0.98), including 57.6% of men and 70.2% of patients with LVEF > 57% (Yoon et al., 2026). The authors distilled a 16-variable PARAGUIDE tool with repeated 4-fold internal cross-validation and stated external validation in PARADIGM-HF, but this is a post-hoc heterogeneous-treatment-effect model on a trial that was not positive overall, PARADIGM-HF is an HFrEF population so it is not HFpEF transportability, and the analysis was funded by Novartis; it does not license routine sacubitril prescription in HFpEF or replace guideline-level subgroup evidence.
Part 4: Wearable Device AI and the Asymptomatic AFib Dilemma
Apple Heart Study: 450,000 Participants, 84% PPV, Massive Clinical Uncertainty
The Apple Heart Study (Perez et al., 2019) was the largest prospective study of wearable AFib detection Perez et al., 2019.
Study design: - 419,297 participants wore Apple Watch with photoplethysmography (PPG)-based irregular pulse detection - When algorithm detected irregular pulse, participant received ECG patch to confirm AFib - Primary outcome: PPV of algorithm (what percentage of alerts were true AFib)
Results: - 2,161 participants (0.52%) received irregular pulse notifications - 450 returned ECG patches - 153 of those patches showed AFib - PPV: 84% (better than expected for screening test)
But here’s the clinical problem: - 84% of 0.52% = 0.44% of participants had confirmed AFib - Most were asymptomatic - Most had paroxysmal AFib (brief episodes) - Clinical question: Should asymptomatic paroxysmal AFib detected by smartwatch be treated with anticoagulation?
CHADS-VASc does not answer this: CHADS-VASc was developed for AFib detected clinically or on ECG, not for asymptomatic device-detected episodes. Stroke risk for smartwatch-detected paroxysmal AFib is uncertain.
Key trial results:
- GUARD-AF (n=11,905, adults ≥70): 14-day ECG patch screening increased AFib diagnoses (5.0% vs 3.3%) and anticoagulant initiation (4.2% vs 2.8%), but screening did not reduce stroke hospitalizations (0.7% vs 0.6% over 15 months). Trial halted early due to COVID (enrolled 11,905 of planned 52,000), limiting statistical power (Lopes et al., 2024)
- EQUAL trial (van Steijn et al., JACC, 2026): RCT, 437 participants ≥65 with CHA₂DS₂-VASc ≥2 and no prior AF, randomized to 6-month Apple Watch monitoring versus standard care. New AF detected: 9.6% vs 2.3% (HR 4.40, 95% CI 1.66-11.66); most episodes (57.1%) were asymptomatic (van Steijn et al., 2026)
- HEARTLINE: Published design and engagement reports do not substitute for the trial’s prespecified cardiovascular outcome results (Gibson et al., 2023; Spertus et al., 2025).
Apple Watch Hypertension Notification (FDA-cleared September 2025)
FDA cleared the Hypertension Notification Feature under K250507 on September 11, 2025. The software analyzes opportunistically collected PPG data over multiple days and may notify adults age 22 years or older who have not been diagnosed with hypertension. It is not intended to replace diagnosis, monitor treatment, or provide blood-pressure surveillance (FDA K250507 summary).
Clinical guidance: The notification is a screening signal, not a diagnosis. Patients receiving hypertension notifications require office blood pressure measurement before any treatment decisions. The tool has population-level potential (hundreds of millions of Apple Watch users) but performance in diverse populations requires monitoring.
Current clinical management: An irregular-pulse notification requires diagnostic confirmation and individualized stroke-risk assessment. Detection studies alone do not establish a universal anticoagulation threshold, and management should follow current atrial-fibrillation guidance for the confirmed rhythm, episode duration, risk profile, and patient preferences.
The lesson: Detecting more disease does not automatically improve outcomes. GUARD-AF found more AFib and led to more anticoagulation, yet stroke rates remained unchanged.
Beyond wearables, multimodal retinal imaging (“oculomics”) also carries AF signal: across AlzEye and UK Biobank, thinner macular ganglion-cell layers associated with prevalent AF and with higher hazard of incident AF over follow-up, supporting opportunistic eye images as a possible early-risk marker rather than a stand-alone AF test (Huemer et al., 2026). Two-cohort association: not an AF diagnostic pivotal; hospital-coded AF undercounts silent AF.
AI-Guided Catheter Ablation: TAILORED-AF
The TAILORED-AF trial is the first large transatlantic RCT to show that AI-guided targeting of ablation beyond pulmonary vein isolation improves outcomes in persistent AF (Deisenhofer et al., 2025, Nature Medicine).
Design: Multicenter, double-blind, superiority RCT; 370 patients with drug-refractory persistent AF across 26 centers in 5 countries (US and Europe); randomized to AI-guided tailored ablation (PVI + targeting of spatio-temporal dispersion areas identified by Volta AF-Xplorer) versus conventional PVI alone.
Results: - AF-free at 12 months: 88% tailored vs 70% PVI-only (P < 0.0001) - Single-procedure AF freedom: 62% vs 48% (P = 0.04) - Periprocedural AF termination: 66% vs 15% (P < 0.001) - Safety: No difference in complications; procedure time approximately doubled in the tailored arm
What the AI does: The AF-Xplorer algorithm analyzes intracardiac electrograms to identify areas of spatio-temporal dispersion (sequential activation spanning the entire AF cycle length across 3+ adjacent electrodes). These areas are additional ablation targets beyond the standard PVI.
Clinical implications: For electrophysiologists treating drug-refractory persistent AF (one of the most challenging arrhythmias), this provides the first RCT evidence that AI-guided identification of non-PV targets improves outcomes. The doubled procedure time is a real limitation; patient selection for the additional ablation burden matters.
FDA status: FDA cleared Volta AF-Xplorer under K243812 on May 9, 2025 as a programmable diagnostic computer that assists operators in annotating atrial electrograms exhibiting spatiotemporal dispersion (FDA K243812). The regulatory intended use and the TAILORED-AF trial endpoint should be reviewed separately.
Part 5: Lessons From IBM Watson
What Went Wrong With Watson for Oncology (and Its Cardiology Applications)
IBM Watson for Oncology promised treatment recommendations based on literature and clinical information. Its history is relevant to cardiology as a warning against transferring brand reputation, concordance studies, or retrospective performance into claims of patient benefit.
What the published evidence supports: - Peer-reviewed studies primarily evaluated concordance between Watson for Oncology and multidisciplinary recommendations, with variable results across cancers and settings (Tian et al., 2021; Somashekhar et al., 2018). - These studies did not establish improved patient outcomes, and concordance with a local recommendation is not equivalent to clinical benefit. - Investigative reporting described unsafe or incorrect recommendations, but those reports should be identified as reporting rather than randomized evidence (Ross and Swetlitz, 2018).
Why did hospitals buy it? Aggressive marketing, partnerships with major academic centers (Memorial Sloan Kettering), and promises of “AI-assisted decision-making” that sounded impressive to hospital executives.
Why did it fail? 1. Training data problem: Watson was trained on expert preferences (what MSK oncologists recommended) not evidence (what RCTs showed worked) 2. No clinical validation: Deployed without prospective trials showing benefit 3. Black box: Physicians could not understand why Watson made recommendations, eroding trust 4. Overpromising: Marketed as “thinking like a doctor” when it was really “echoing MSK treatment patterns”
Cardiology boundary: Evidence about Watson for Oncology should not be presented as direct evaluation of a cardiology product. The transferable lesson is methodological: identify the exact system, population, comparator, endpoint, and version before drawing conclusions.
The lessons for cardiovascular AI: 1. Demand RCT evidence: If a vendor cannot show published outcomes studies, do not deploy their tool 2. Beware proprietary algorithms: Black boxes hide methodological flaws 3. Marketing ≠ evidence: Partnerships with prestigious institutions do not prove clinical benefit 4. Physician judgment remains essential: No AI should make autonomous treatment recommendations
Part 6: Equity in Cardiovascular AI
The Pulse Oximetry Problem Extends to ECG and Echo
In 2020, researchers discovered that pulse oximeters systematically overestimate oxygen saturation in Black patients, leading to delayed treatment for hypoxemia Sjoding et al., 2020.
Potential equity problems require direct measurement in cardiovascular AI:
ECG algorithms: - Sokolow-Lyon criteria for LVH (embedded in automated ECG algorithms) have lower sensitivity in Black patients - QTc prolongation thresholds do not account for race-specific differences in baseline QT intervals - Most ECG AI trained on predominantly white populations from tertiary care centers
Echo AI: - Development cohorts, disease prevalence, equipment, and acquisition protocols may differ from the intended deployment population. - Subgroup estimates should be reported with sample sizes and uncertainty rather than inferred from pooled performance.
Wearable detection: - Optical-sensor performance can vary with device fit, motion, perfusion, and skin characteristics. - Demographic representation and subgroup performance should be checked in the exact validation study and device labeling.
What can you do? 1. Ask vendors for race-stratified performance metrics: If they cannot provide them, the algorithm was not validated equitably 2. Validate locally: Test algorithm performance in your patient population before widespread deployment 3. Monitor outcomes by race: Track algorithm errors, false positives, false negatives by race/ethnicity 4. Maintain clinical skepticism: AI is a tool, not truth. Your clinical judgment remains essential.
Part 6b: Digital and AI-Assisted Heart-Failure Care
ADMINISTER Trial: Digital Consults, Not Machine Learning
The ADMINISTER trial evaluated a multifaceted digital-consult strategy for guideline-directed medical therapy optimization in HFrEF. The intervention was not described as an artificial-intelligence or machine-learning system (Man et al., 2024).
Design: 150 HFrEF patients across multiple European centers, randomized 1:1 to digital consults versus usual care for 12 weeks. Digital consults combined: patient digital data sharing (home vitals, KCCQ questionnaires), patient e-learning modules, and guideline recommendations pushed to clinicians.
Result: GDMT score improved by a median of 1.19 units in the digital consult arm versus 0.08 in usual care (P < 0.001). This represents a meaningful increase in the proportion of patients receiving all four foundational drug classes.
Why this matters: ADMINISTER is evidence for a digital-care workflow, not evidence for AI. Its results should not be transferred to machine-learning systems that predict heart-failure risk or recommend treatment.
ASSIST-HF: Supervised Medication Optimization
In the ASSIST-HF randomized pilot, a generative-AI workflow delivered by nonmedical staff supported medication optimization in heart failure with reduced ejection fraction, with cardiologist approval of treatment decisions and prescriptions. The comparison evaluated the combined workflow, including visits every two weeks, against usual care and did not isolate the model’s contribution (Navarese et al., 2026).
Part 7: Implementation Framework
Before Adopting Cardiovascular AI Tools
Questions to ask vendors:
- “What peer-reviewed evidence supports the exact claim and product version?”
- Internal validation can support development, but not transportability or patient benefit
- Judge the design and endpoint, not only the journal name
- “What is the algorithm’s performance across clinically relevant subgroups?”
- Missing subgroup estimates limit assessment of equity and transportability
- Request denominators and uncertainty, not only point estimates
- “What is the false-positive rate at your recommended sensitivity threshold?”
- AUC alone is insufficient for an implementation decision
- Need to understand PPV/NPV at clinically relevant operating points
- “How does the algorithm integrate with our EHR and clinical workflow?”
- Poor integration causes alert fatigue and missed diagnoses
- Demand live demonstrations in your specific EHR environment
- “What happens when the algorithm fails? What are the failure modes?”
- All algorithms fail sometimes
- You need to understand when and how to recognize failures
- “Can performance and workflow be evaluated in our intended population before broad deployment?”
- The evaluation should reflect local prevalence, equipment, workflow, and patient mix
- “What is the full pathway cost, and what product-specific economic evidence exists?”
- Include licensing, integration, confirmation, downstream testing, monitoring, and staff response
- Compare with the actual local alternative
- “How are responsibilities allocated when the system fails or is overridden?”
- Review labeling, institutional policy, vendor terms, and jurisdiction-specific advice
- Do not assume a universal allocation of liability
- “Can you provide references from cardiologists at hospitals similar to ours who use this tool?”
- Talk to actual users, not marketing testimonials
- Ask about problems, workflow disruptions, false positives
- “Is the algorithm FDA-cleared? If yes, through what pathway?”
- 510(k) clearance ≠ clinical validation
- 510(k) requires only “substantial equivalence” to existing device
Red Flags Requiring Resolution Before Adoption
- Vendor refuses to share peer-reviewed publications (“Our algorithm is proprietary”)
- No external validation studies (validated only on development cohort)
- Performance metrics not stratified by demographics (equity not assessed)
- No defined failure, override, monitoring, or update process
- A superiority claim without a same-cohort comparator and claim-matched endpoint
Part 8: Cost-Benefit Reality
What Does Cardiovascular AI Actually Cost?
The relevant cost is the entire clinical pathway, not a stale list price. It may include licensing, implementation, interfaces, staff training, confirmatory tests, downstream referrals, alert response, monitoring, model updates, and the opportunity cost of false-positive and false-negative results.
Commercial prices and contract terms change. A procurement review should use the current written quote, identify every variable and fixed cost, and distinguish a consumer-device purchase from a medical diagnostic pathway.
Do These Tools Save Money?
Plausible mechanisms: - Earlier identification may enable confirmatory testing and treatment. - Automated measurement may reduce interpretation time or variability. - A pathway can still add cost if it produces low-value confirmation, duplicate work, or unmanageable alerts.
In practice: Uncertain: - Economic evidence is product-, comparator-, population-, and time-horizon-specific. - The Attia research study did not establish whether earlier low-EF identification reduces downstream costs. - An economic analysis of FFRCT illustrates how conclusions depend on the diagnostic strategy being compared (Hlatky et al., 2015).
Implementation costs to measure: - IT integration and security review - Workflow redesign: Cardiologist/administrator time - Training: Sonographer/tech/physician education - Maintenance: Software updates, troubleshooting - Alert management: Triaging false positives
Illustrative local model, not a published estimate: - Assume an annual software cost of $50,000. - Assume 10 minutes saved on each of 5,000 studies, or 833 hours. - Assume the released time is valued at $200 per hour, producing a modeled gross value of $166,600 before implementation and downstream costs. - Test each assumption with observed local data, including whether time is truly released, whether quality changes, and whether additional confirmation is required.
The question: Are those assumptions valid? We do not have good data.
Part 9: The Future of Cardiovascular AI
What’s Coming in the Next 5 Years
Demonstrated research directions with uneven clinical maturity: 1. Expanded ECG AI applications: Detection of pulmonary hypertension, aortic stenosis, HCM from ECG 2. Wearable integration with EHR: Smartwatch data flowing into medical records with clinical decision support 3. Automated echo AI in primary care: Point-of-care echo by non-cardiologists with AI guidance 4. Predictive models for sudden cardiac death: Risk stratification for ICD placement beyond EF alone
Promising but uncertain: 1. AI-guided medication optimization: Automated titration of GDMT in HF 2. Real-time procedural guidance: AI-assisted PCI, ablation, structural interventions 3. Precision medicine for CAD: Genetic + imaging + clinical data to personalize revascularization decisions
Beyond the evidence reviewed here: 1. Autonomous cardiovascular diagnosis without defined oversight: The chapter does not identify evidence supporting unrestricted replacement of clinical judgment. 2. Smartwatch-only AF management: A screening notification does not establish an anticoagulation indication without rhythm confirmation and risk assessment.
The rate-limiting step: Not algorithmic accuracy alone. The required evidence depends on the claim, ranging from external validation to comparative studies with patient-relevant endpoints.
Most cardiovascular AI has impressive technical performance. What we lack is evidence that deploying these tools actually helps patients live longer or better.
Professional Society Guidelines on AI in Cardiology
The most useful guidance is direct and claim-specific. An AHA scientific statement addresses value assessment for AI in cardiovascular imaging (Hanneman et al., 2024). A later AHA advisory provides a risk-proportionate framework for predeployment evaluation, implementation monitoring, and postdeployment surveillance across health AI (Jain et al., 2025). A joint EHRA/HRS/ESC statement addresses AI in clinical electrophysiology (Svennberg et al., 2025).
Society guidance is a framework for evaluation and governance. It is not authorization of every model discussed in this chapter.
AHA Science Advisory on AI Evaluation (November 2025)
The American Heart Association published its first formal framework document on health AI deployment: “Pragmatic Approaches to the Evaluation and Monitoring of Artificial Intelligence in Health Care” (Circulation, November 2025).
The advisory proposes a three-phase evaluation framework: predeployment validation, implementation monitoring, and postdeployment surveillance. It emphasizes that evaluation should be proportionate to risk and continue after deployment (Jain et al., 2025).
Practical implication: If your institution is deploying any AI-assisted cardiovascular tool, the AHA now provides a specific framework for evaluating it.
EHRA/HRS/ESC Scientific Statement on AI in Electrophysiology (May 2025)
The joint EHRA/HRS/ESC working group published a review and scientific statement on AI in clinical electrophysiology, including atrial-fibrillation management, sudden-cardiac-death prediction, and electrophysiology laboratory applications (Svennberg et al., 2025).
A standardized 29-item AI reporting checklist was applied: no single checklist item appeared in more than 55% of papers. Trial registration, participant demographics, and training metrics were all reported in fewer than 20% of studies. The statement establishes standardized reporting requirements but stops short of clinical recommendations, because the evidence base is not yet sufficient.
The bottom line for cardiologists: Reporting limitations across the reviewed literature support careful study appraisal. They do not establish that every study or every checklist item is inadequate.
Cardiovascular Imaging Value Assessment
The AHA imaging statement recommends evaluating the clinical pathway, comparator, outcomes, costs, and distributional consequences rather than treating technical accuracy as value (Hanneman et al., 2024). This framework is directly relevant to the ECG, echo, CCTA, and wearable examples above.
Endorsed Risk Calculators
The ACC/AHA endorse several validated risk calculators that incorporate statistical modeling:
- ASCVD Risk Estimator Plus: 10-year and lifetime cardiovascular risk
- Pooled Cohort Equations: Primary prevention statin therapy decisions
- CHA2DS2-VASc: Stroke risk in atrial fibrillation
- HAS-BLED: Bleeding risk with anticoagulation
- TIMI Risk Score for STEMI: 30-day mortality risk at presentation (Morrow et al., 2000)
- TIMI Risk Score for UA/NSTEMI: Risk stratification for unstable angina and non-ST elevation MI (Antman et al., 2000)
- GRACE Score: Hospital mortality and 6-month post-discharge mortality in acute coronary syndromes (Granger et al., 2003)
These represent validated, guideline-integrated predictive tools that precede modern AI but establish the framework for algorithmic clinical decision support.
Electrophysiology Guidance
The joint EHRA/HRS/ESC statement is the direct route for electrophysiology-specific AI evaluation. It should be consulted for its actual scope and recommendations rather than replaced by uncited summaries attributed separately to each society (Svennberg et al., 2025).
Key Takeaways
10 Principles for Cardiovascular AI
ECG performance is task-specific: Conventional and machine-learning outputs require review against the tracing and clinical context.
Hidden ECG phenotypes are measurable: The Attia model established diagnostic performance for low EF, not improved outcomes for every implementation.
Echo AI evidence spans different claims: Acquisition, measurement, study interpretation, foundation-model benchmarks, authorization, and patient outcomes must not be conflated.
Wearable AFib detection creates clinical dilemmas: High sensitivity, but asymptomatic paroxysmal AFib management uncertain
HF alert burden is threshold-dependent: AUC does not determine the false-positive rate or positive predictive value at the proposed operating point.
Match evidence to the claim: Technical accuracy, workflow performance, and patient benefit require different study designs.
Measure equity directly: Request subgroup denominators, performance estimates, uncertainty, and local monitoring.
Implementation > accuracy: Poor EHR integration causes missed diagnoses despite accurate algorithms
Do not transfer evidence across products or specialties: A brand, institution, or related model is not a substitute for evidence about the exact system.
Responsibility is shared and context-specific: Define clinical, institutional, vendor, monitoring, and escalation duties before deployment.
Hypothetical Vendor Evaluation
The following vendor and financial details are fictional. They are included to practice evidence appraisal and procurement reasoning, not to describe a real product, contract, legal outcome, or measured return on investment.
Scenario: Your Cardiology Department Is Considering Purchasing an AI Tool
The pitch: A vendor demonstrates an AI tool that predicts 30-day cardiovascular mortality risk for hospitalized cardiology patients. They show you: - AUC 0.92 in internal validation - “Outperforms traditional risk scores” - Integration with your EHR - Cost: $150,000/year
The department chair asks for your recommendation.
Questions to Ask Before Recommending Purchase:
- “What peer-reviewed publications support this algorithm?”
- Look for Circulation, JACC, JAMA Cardiology publications
- Internal validation white papers are insufficient
- “What is the positive predictive value at clinically useful sensitivity thresholds?”
- If 30-day mortality is 3%, even 92% AUC may yield terrible PPV
- Ask for sensitivity/specificity table at multiple thresholds
- “How does this algorithm perform in patient populations similar to ours?”
- Algorithm validated at academic medical center may fail at community hospital
- Request performance stratified by age, race, sex, comorbidities
- “What interventions will we apply to high-risk patients identified by this algorithm?”
- If answer is “closer monitoring,” what’s the evidence that prevents deaths?
- Many high-risk patients die despite optimal care
- “What are this algorithm’s failure modes?”
- Does it underestimate risk in young patients? Overestimate in elderly?
- What clinical situations does it handle poorly?
- “Can we run a prespecified local evaluation before broad deployment?”
- Define sample size, endpoints, thresholds, subgroups, and stop rules before seeing results
- Compare predictions and workflow effects with an appropriate local comparator
- “How are responsibilities allocated if an output is wrong, delayed, unavailable, or ignored?”
- Review the product labeling, institutional policy, contract, and jurisdiction-specific advice
- “What is the value compared with existing risk stratification?”
- Identify the actual comparator, full pathway cost, and prespecified outcomes
- Ask whether any economic analysis uses the same population and workflow
- “How will this integrate with nursing workflow? Who triages the high-risk alerts?”
- Implementation costs often exceed purchase price
- Alert fatigue is real
- “Can independent users at comparable hospitals describe their experience?”
- Get real user experiences, not marketing testimonials
Red Flags in This Scenario:
AUC 0.92 reported without PPV/NPV: Insufficient to determine alert burden or clinical value
“Outperforms traditional risk scores”: Were comparisons done in same patient cohort? Published?
No mention of prospective evaluation: Retrospective performance does not establish prospective workflow or outcome benefit
High annual cost without economic evidence: Treat the stated price as a scenario assumption and require a full local value model
No documented failure modes or monitoring plan: An explanation display cannot substitute for safety controls, but absent failure analysis is a serious gap
Hypothetical Decision Exercises
These cases are fictional educational exercises. They do not describe actual patients, products, contracts, outcomes, or legal determinations. The purpose is to practice identifying evidence boundaries and missing information.
Scenario 1: The AI-Detected Low EF
Clinical situation: A 58-year-old woman with hypertension presents to primary care for annual physical. ECG ordered as part of routine screening shows normal sinus rhythm, normal intervals, no ST/T changes. However, the ECG machine’s AI algorithm flags: “Low ejection fraction predicted. Recommend echocardiogram.”
The patient is asymptomatic. There is no dyspnea, edema, or chest pain, and the physical examination is normal. The clinician has not previously encountered this AI alert.
Question 1: Do you order the echocardiogram based on this AI prediction?
Click to reveal answer
Answer: The output is not sufficient to answer without identifying the system and its evidence. Echocardiography may still be reasonable after clinical assessment, but the rationale should not be that every low-EF AI alert is validated.
Reasoning: The Attia research model reported 86.3% sensitivity and 85.7% specificity for detecting EF ≤35% in its test set (Attia et al., 2019). That paper does not establish that the local ECG machine uses the same model, version, threshold, or intended population.
Key points: - Identify the product, version, intended use, regulatory status, and local policy. - Review whether the patient has an independent clinical indication for echocardiography. - Determine the expected positive predictive value at the local prevalence and threshold. - Treat an alert as a screening signal until a diagnostic test confirms ventricular function.
However: - Explain that this is a screening signal and may be a false positive - Explain that AI detected subtle ECG patterns suggesting possible heart dysfunction - Do not alarm patient unnecessarily before echo confirms
If echo confirms reduced EF: Evaluate etiology and apply current heart-failure guidance to the individual patient.
If echo normal: Reassure patient, document that AI alert was false positive
Bottom line: Evidence for one research model cannot be transferred automatically to an unidentified machine output. A low-risk confirmatory test may be reasonable, but the decision requires product verification and clinical context.
Scenario 2: The Apple Watch AFib Alert
Clinical situation: A 68-year-old man with no listed CHA₂DS₂-VASc risk factor other than age 65–74 has a score of 1 and presents with irregular-pulse notifications from an Apple Watch. He received three alerts over the past week while asymptomatic, with no palpitations, dyspnea, or dizziness.
You order ECG in office: Normal sinus rhythm. You order 24-hour Holter: Shows 2-hour episode of atrial fibrillation at 3 AM (patient asleep, asymptomatic).
Question 2: Do you start anticoagulation for asymptomatic, device-detected paroxysmal AFib?
Click to reveal answer
Answer: Unclear. This is a genuine clinical gray zone.
Arguments FOR anticoagulation: - Confirmed AF and an age-related score of 1 warrant individualized assessment under current guideline thresholds and shared decision-making. - Symptoms are not required for thromboembolic risk, but episode duration and the method of detection affect the evidence base. - Subclinical AFib detected by pacemakers has been associated with increased stroke risk (though episodes were typically >24 hours) - The Apple Heart Study’s 84% positive predictive value applied to concurrent irregular-pulse notifications and analyzable simultaneous ECG patches. In this hypothetical case, the Holter tracing, not the watch alert alone, confirms AF.
Arguments AGAINST anticoagulation: - CHADS-VASc was derived from symptomatic AFib populations; applicability to device-detected asymptomatic AFib uncertain - Paroxysmal AFib (2-hour episodes) may carry lower stroke risk than persistent AFib - Bleeding risk and net clinical benefit depend on the individual patient and chosen therapy; no universal annual rate should decide this case. - GUARD-AF (2024) found no stroke reduction from screening: increased AFib detection (5.0% vs 3.3%) and anticoagulation (4.2% vs 2.8%), but stroke rates 0.7% vs 0.6% (trial underpowered due to early COVID termination) (Lopes et al., 2024)
Trial status: - GUARD-AF: Completed. Screening increased detection and treatment but did not reduce stroke (underpowered) - EQUAL trial: Completed. Smartwatch detected 4x more new AF than standard care (9.6% vs 2.3%); stroke outcomes not powered (van Steijn et al., 2026) - HEARTLINE: Published design and engagement papers should not be described as cardiovascular outcome results (Gibson et al., 2023; Spertus et al., 2025).
Decision process: - Confirm the rhythm and characterize AF burden with a clinically appropriate monitor. - Calculate CHA₂DS₂-VASc correctly and assess bleeding risk, comorbidities, and preferences. - Apply the current AF guideline to the confirmed episode duration and risk profile. Do not invent a universal 6–24-hour treatment threshold. - Use shared decision-making when benefit is uncertain.
Example discussion: - “You have real AFib, detected by your watch and confirmed on Holter monitor” - “The stroke risk is uncertain because you’re asymptomatic and episodes are brief” - “The score and episode duration inform the decision, but neither alone settles it in every patient” - “The notification required confirmation, and the confirmed rhythm now needs guideline-based risk assessment” - “The largest screening trial (GUARD-AF) found more AFib but no fewer strokes, though it was cut short by COVID” - “Options depend on the confirmed burden, stroke risk, bleeding risk, and preferences”
Bottom line: GUARD-AF showed that its screening strategy increased detection and treatment without a demonstrated reduction in stroke hospitalizations, although early termination limited power. That result does not make every treatment option equally appropriate. Document the confirmed rhythm, risk calculation, evidence boundary, shared decision, and follow-up plan.
Scenario 3: The Proprietary HF Readmission Model
Clinical situation: A fictional hospital’s population-health team is considering a proprietary AI tool that predicts 30-day heart-failure readmission risk. For the exercise, assume the vendor reports an AUC of 0.87 and says the proposed threshold flags 35% of discharges as high risk.
The case-management team asks whether intensive post-discharge interventions should be offered to every algorithm-flagged patient.
Assume, only for this exercise, that the intervention costs $800 per patient and the hospital has 400 heart-failure discharges per year.
Question 3: Do you implement the algorithm-driven intervention program?
Click to reveal answer
Answer: No, not without further analysis.
Problems with this scenario:
1. AUC does not determine the threshold confusion matrix: - Assume for the exercise that the local readmission prevalence is 20%, producing 80 readmissions among 400 discharges. - A threshold that flags 35% would flag 140 patients, but the number of true and false positives cannot be derived from AUC alone. - The vendor must provide sensitivity, specificity, positive predictive value, and uncertainty at the exact proposed threshold in a comparable cohort.
2. The economic result depends on unmeasured assumptions: - Under the exercise assumptions, 140 patients multiplied by $800 produces a direct annual intervention cost of $112,000. - A value model still needs the intervention effect, implementation cost, downstream utilization, quality outcomes, and the relevant local cost perspective. - It is not valid to insert an uncited average readmission cost or assume a fixed proportion is preventable.
3. The intervention needs its own evidence: - Prediction performance does not establish that the proposed home-visit and telephone pathway reduces readmissions. - The intervention, intensity, comparator, adherence, and endpoint should match the evidence used for the implementation decision.
4. Algorithm transparency is absent: - What features is the algorithm using? - Compare the model with a simpler, interpretable baseline using the same cohort and endpoint. - Incremental performance must be connected to a meaningful difference in decisions or outcomes.
What to do instead:
- Request vendor provide:
- Peer-reviewed publication of algorithm validation
- Performance stratified by demographics
- PPV/NPV at multiple sensitivity thresholds
- Feature importance (what is the algorithm learning?)
- Prespecified local evaluation:
- Define the cohort, comparator, sample-size rationale, threshold, endpoints, subgroups, and stop rules in advance.
- Measure calibration, predictive values, alert volume, workflow burden, and whether recommendations change.
- Evaluate intervention evidence:
- Systematic review of transitional care interventions for HF
- Select interventions supported for the intended population and delivery setting.
- Compare alternative allocation strategies:
- Compare model-based allocation with current care, an interpretable rule, and broader eligibility when clinically plausible.
- Do not claim cost-effectiveness from an assumed effect.
Bottom line: Proprietary algorithms with impressive AUCs may or may not add value over existing practice. Demand evidence of incremental decision value, workflow feasibility, and economic consequences before implementation.
The algorithm is not necessarily wrong. It’s just not clear it adds value over existing approaches.
Questions About Cardiology AI
Can AI read an ECG?
Automated ECG interpretation and newer machine-learning models can support rhythm classification and case finding. Performance is task-specific, and physician review remains necessary. A low-ejection-fraction model reported an AUC of 0.93, but diagnostic performance does not by itself establish better patient outcomes.
How accurate is Apple Watch for AFib detection?
The Apple Heart Study reported 84% positive predictive value for concurrent atrial fibrillation among participants with an irregular-pulse notification and analyzable simultaneous ECG patch data. The estimate does not mean that 84% of all watch alerts in every population are atrial fibrillation, and notifications require clinical confirmation.
Can AI predict heart failure?
Models can estimate heart-failure risk or detect signals associated with low ejection fraction, but no universal false-positive rate applies. Clinical value depends on prevalence, threshold, alert burden, confirmatory testing, and whether a tested response pathway improves care.
What are the leading echocardiography foundation models?
EchoPrime and PanEcho are peer-reviewed echocardiography foundation-model studies; EchoJEPA is a preprint. Their architectures, datasets, tasks, and validation designs differ, and none should be treated as proof of improved patient outcomes or unrestricted clinical use.