Diagnostic Imaging, Radiology, and Nuclear Medicine

Radiology has the largest regulatory footprint for medical AI, but regulatory volume is not evidence of clinical benefit. FDA’s periodically updated AI-enabled medical-device list continues to show radiology as the dominant specialty category, while FDA cautions that the list is not comprehensive and depends on identifiable AI-related terms in public authorization records (FDA AI-Enabled Medical Devices). Randomized evidence now shows both benefit and limitation: a stroke workflow trial shortened selected treatment intervals without improving 90-day functional independence, while a lung-cancer pathway trial found no reduction in time to CT or diagnosis.

Learning Objectives

After reading this chapter, you will be able to:

  • Understand the types of AI applications in radiology (CAD, triage, quantification, report generation)
  • Evaluate evidence for specific imaging AI tools across modalities
  • Recognize FDA-cleared devices and their validation evidence
  • Understand workflow integration challenges specific to radiology
  • Assess the “Will AI replace radiologists?” debate with evidence
  • Identify failure modes and limitations of imaging AI
  • Recognize cognitive de-skilling risks in AI-assisted training
  • Mitigate AI anchoring bias through workflow design
  • Navigate liability and reimbursement considerations

The Clinical Context: Radiology leads medical AI adoption, with the FDA’s AI-enabled medical device list consistently dominated by radiology entries (FDA AI-Enabled Medical Devices). Deep learning has achieved radiologist-level performance for specific tasks, but significant gaps remain between retrospective accuracy and prospective clinical utility.

AI Application Categories:

Category Purpose Key Examples
Detection (CAD) Flag suspicious findings Lung nodules, breast masses, ICH, PE
Triage Prioritize critical studies Stroke, pneumothorax, hemorrhage
Quantification Automated measurements Tumor volumes, cardiac function, ASPECTS
Diagnosis Classify findings Diabetic retinopathy, breast density
Report Generation Draft findings Normal CXR reports (emerging)

What Works Well:

Application Evidence Key Benefit
ICH and LVO triage Randomized and observational workflow evidence Shorter selected workflow intervals; patient-outcome benefit remains unproven for isolated alerts
Chest radiograph nodule detection Randomized detection evidence Increased actionable nodule detection in one defined screening workflow
Mammography AI Randomized and prospective program evidence Higher sensitivity or detection and lower reading workload in defined workflows, with recall tradeoffs in AITIC (Lång et al., 2023; Gommers et al., 2026; Elías-Cabot et al., 2026)
Report drafting Prospective observational evidence Lower documentation time in one 12-hospital system, without a randomized patient-outcome comparison

What’s Problematic:

Application Concern
COVID-19 CXR AI Most learned confounders, not pathology DeGrave et al., 2021
First-gen mammography CAD Increased recalls without improving detection
Autonomous diagnosis Liability concerns, misses incidental findings
CE-CT liver-malignancy additional readers Large single-arm Chinese trial evidence; not FDA-cleared and not K252970 triage (Zhang et al., 2026)
NC-MRI FLL assistance (PRISM) Multicenter Chinese diagnostic study; not FDA-cleared and not a claim that contrast MRI can be omitted (Dong et al., 2026)
Consumer LMMs on neuroradiology Open-ended primary accuracy 51.6–62.7% on 252 RSNA cases; multiple-choice overstates generative performance (Acar et al., 2026)

Will AI Replace Radiologists? No. AI excels at narrow tasks but struggles with: complex reasoning across modalities, incidental findings, clinical integration, edge cases, and liability/accountability. Future is human-AI partnership.

Key Implementation Principles:

  1. Local validation essential - Performance varies by scanner, population, workflow
  2. Alert fatigue is real - Balance sensitivity vs. specificity carefully
  3. PACS integration required - Standalone systems fail
  4. Responsibility must be defined locally - Liability is fact-specific and jurisdiction-specific; authorization is not a legal shield
  5. Continuous monitoring - Performance drifts with scanner changes

The Bottom Line: Radiology AI has credible evidence for selected detection, screening, and workflow tasks, but products and endpoints cannot be treated as interchangeable. Demand local validation, monitor continuously, preserve the standard queue for unflagged studies, and define responsibility before deployment.


Introduction

Radiology has more FDA-authorized AI devices than any other medical specialty, based on the FDA’s periodically updated AI-enabled medical-device list (FDA AI-Enabled Medical Devices). This maturity brings both opportunities and cautionary lessons. In a stepped-wedge cluster-randomized stroke trial, AI activation reduced adjusted door-to-groin time by 11.2 minutes and CT-to-endovascular-treatment time by 9.8 minutes, but did not significantly improve 90-day functional independence (Martinez-Gutierrez et al., 2023). Other applications, like early mammography CAD systems, increased false positives without improving cancer detection.

The core question remains: does this tool improve patient outcomes in my clinical context?

Evidence should be read in descending order of clinical proximity:

  1. Randomized clinical-utility studies: Compare complete care or screening workflows and measure downstream outcomes.
  2. Prospective implementation studies: Evaluate a deployed workflow without randomization.
  3. Prospective silent studies: Run the model on live cases without exposing outputs to clinicians.
  4. Retrospective external testing: Test stored data from institutions outside model development.
  5. Reader studies: Measure performance under controlled interpretation conditions.
  6. Regulatory evidence: Establishes that a device met the requirements of its pathway and intended use.
  7. Vendor or benchmark evidence: Generates hypotheses but does not establish clinical utility.

AUC, sensitivity, and specificity do not establish that a deployed workflow benefits patients. The endpoint must match the claim. A triage product should be evaluated for time to interpretation or action, a screening product for detection, recall, interval disease, and overdiagnosis, and a report generator for final-report accuracy, editing burden, and consequential errors.

AI Applications in Radiology

1. Computer-Aided Detection (CAD)

  • Purpose: Flag suspicious findings for radiologist review
  • Applications: Lung nodules, breast masses, intracranial hemorrhage, pulmonary embolism, fractures
  • Status: Widely deployed, mixed evidence for clinical benefit
  • Limitation: High false positive rates → alert fatigue

2. Triage and Worklist Prioritization

  • Purpose: Identify critical findings requiring urgent attention (e.g., ICH, pneumothorax)
  • Applications: ED triage, ICU studies, stroke code activation
  • Evidence: Product-specific studies show that selected systems can shorten defined workflow intervals; downstream patient benefit remains inconsistent or unproven
  • Caveat: Mis-triage risks: false negatives delay care, false positives waste resources

3. Quantification and Measurement

  • Purpose: Automated measurements (tumor volumes, cardiac function, bone density)
  • Applications: Oncology treatment response, cardiac MRI quantification, stroke ASPECTS scores
  • Advantage: Standardized, reproducible measurements
  • Limitation: Segmentation errors propagate to measurements

4. Diagnosis and Classification

  • Purpose: Classify findings (benign vs. malignant, normal vs. abnormal)
  • Applications: Diabetic retinopathy grading, breast density assessment, prostate MRI PI-RADS scoring
  • Status: Most imaging AI is assistive. Any autonomous use must be supported by the exact authorization and intended-use record.
  • Critical consideration: Liability when algorithm makes final diagnosis without radiologist confirmation

5. Report Generation

  • Purpose: Auto-generate structured reports or draft findings
  • Applications: Normal chest X-ray reports, standardized measurements
  • Status: Emerging, limited deployment
  • Risk: Hallucinations, missing incidental findings

In a prospective cohort across a 12-hospital academic system, documentation time was 15.5% lower for 11,980 model-assisted radiograph interpretations than for 11,980 matched baseline interpretations. A blinded peer-review sample of 800 reports found no detected difference in clinical accuracy or textual quality (Huang et al., 2025). This was observational evidence from one system, not a randomized demonstration of better patient outcomes.

FDA January 2026 CDS Guidance: what counts as Non-Device CDS

In January 2026, FDA issued its final guidance on Clinical Decision Support (CDS) software (issued January 29, 2026). The guidance clarifies which CDS software functions are excluded from the definition of a device under section 520(o)(1)(E) (Non-Device CDS criteria) and provides examples; it also notes that FDA’s existing digital health policies continue to apply to software functions that meet the definition of a device, including those intended for use by patients or caregivers (FDA CDS Guidance, January 2026; PDF; FDA CDS FAQs).

For radiology, the practical distinction is intended use: software that acquires, processes, or analyzes medical images generally does not meet the Non-Device CDS criteria, while text-focused tools that reformat or summarize a radiologist’s own interpretation may fall outside FDA device regulation depending on how they are marketed and used.


Clinical Applications by Modality

Chest X-Ray AI

Pneumothorax Detection: Multiple commercial and FDA-authorized systems exist; the exact device record and intended use must be checked - Standalone evidence: A retrospective study of 1,000 chest radiographs from four US hospitals reported 94.3% sensitivity and 92.0% specificity for any pneumothorax at the selected operating point (Hillis et al., 2022) - Implementation evidence: A difference-in-differences study of 603,028 chest radiographs found that adding electronic notifications to CAD shortened time to oxygen supplementation by 143.8 minutes, but did not significantly change time to aspiration or tube thoracostomy or surgical consultation (Oh et al., 2025) - Limitation: False positives on chest tubes, skin folds, positioning artifacts

Pulmonary Nodule Detection: - Performance: Depends on nodule size, acquisition, threshold, population, reference standard, and reader workflow - Potential benefit: A randomized screening study below found more actionable nodule detections, but long-term outcome benefit was not measured - Challenge: Interpreting sub-6mm nodules, distinguishing nodules from vessels/artifacts

Pneumonia Detection (COVID-19 and general): - Hype peak: 2020 COVID-19 pandemic saw flood of AI models - Reality check: Most never deployed, many learned confounders not pathology DeGrave et al., 2021 - Example failure: AI detected “portable” in metadata (sicker patients) not lung findings - Current status: Some deployed for workflow triage, not diagnostic

Common pitfalls: - Sensitivity to positioning, technique, patient factors - Confounding by clinical context (portable vs. PA/lateral) - Overfitting to single-institution data - Underdiagnosis of under-served groups on chest X-ray classifiers is documented separately as an ethics case (Seyyed-Kalantari et al., 2021; see Ethics Case 5) - Training-set sex imbalance can lower thoracic CAD performance for underrepresented sexes (Larrazabal et al., 2020; see Ethics)

Randomized, prospective, and pathway evidence:

  • A pragmatic, open-label, single-center randomized trial of 10,476 health-screening participants found that AI assistance increased detection of actionable lung nodules from 0.25% to 0.59% without a statistically significant increase in positive reports or false referrals (Nam et al., 2023). The endpoint was detection, not mortality or long-term outcomes.
  • A prospective multicenter silent study at five NHS sites analyzed 63,083 adult chest radiographs. The model classified 20% as normal, achieved 97% sensitivity and 35% specificity for abnormal examinations, and had 31 clinically significant misses after discrepant-case review (Storey et al., 2026). A silent trial cannot establish the safety of autonomous exclusion from radiologist review.
  • The multicenter LungIMPACT randomized trial analyzed 93,326 chest radiographs requested by UK primary care and found that AI worklist prioritization did not shorten median time to CT or lung-cancer diagnosis (Woznitza et al., 2026). A changed worklist does not necessarily change the patient pathway.

Region-of-Interest CADe: Qure.ai qXR-Detect

FDA cleared qXR-Detect (Qure.Ai Technologies) through the 510(k) pathway on January 16, 2026 as substantially equivalent, under regulation 892.2070 and product code MYN (FDA K251934 record). Its Indications for Use describe computer-assisted detection that highlights suspicious regions of interest on adult chest radiographs and sorts each into six categories: lung, pleura, bone, mediastinum and hila and heart, hardware, and other. The labeling states that the system is an adjunct concurrent-reading aid, not a replacement for clinician review or judgment (FDA K251934 summary).

Those six categories are regions and object classes, not six named diagnoses. This distinction matters for both clinical claims and local acceptance testing.

The clearance includes an authorized Predetermined Change Control Plan. A PCCP permits prespecified model modifications without a new 510(k) when the change remains within the authorized plan. Version logging and revalidation therefore must follow the software actually deployed, not only the original clearance date.

Mammography AI

Breast Cancer Detection: - First-generation CAD (1990s-2000s): Increased recalls without improving cancer detection Lehman et al., 2015 - Deep learning CAD (2015+): Improved performance, reduces false positives - Evidence: Retrospective external evaluations and reader studies established diagnostic-performance signals, followed by randomized and prospective screening-workflow studies. Lotter et al. was a retrospective and reader-study evaluation, not a randomized screening trial (Lotter et al., 2021) - FDA-cleared systems: Multiple (iCAD, Hologic, Lunit, Therapixel)

Breast Density Assessment: - Purpose: Standardized BI-RADS density classification - Benefit: Reduces inter-reader variability - FDA-cleared: Volpara, Densitas

Which screening workflows does the evidence support?

Applicability depends on the software version, population, imaging modality, and reading pathway. These studies evaluated different interventions. Cancer detection at screening, cancers diagnosed between screens (interval cancers), recall for further assessment, and reader workload are separate outcomes.

Table 8.1: Mammography screening evidence by product, population, and reading pathway. Results apply to the evaluated configurations.
Study, technology, and population Reading pathway and comparator Findings and limits of transfer
MASAI: Transpara 1.7.0; randomized screening trial at four Swedish sites, women aged 40–80 (Lång et al., 2023; Gommers et al., 2026). AI scores 1–9 received one human reader; score 10 received two. AI-supported, risk-adapted reading was compared with standard double reading. The interval-cancer endpoint met the prespecified noninferiority criterion; sensitivity improved with unchanged specificity. This supports the tested risk-adapted pathway. It does not establish interval-cancer superiority, autonomous screening, or a mortality benefit.
PRAIM: Vara, device versions 1.0.5–2.6.2 during the study; prospective observational digital mammography screening at 12 German sites, women aged 50–69; 461,818 examinations analyzed (Eisemann et al., 2025). Standard double reading was retained. Radiologists chose whether to use AI support, including normal tags and a safety net for suspicious findings. AI-supported readings were compared with readings without AI. Model-adjusted cancer detection was 6.7 versus 5.7 per 1,000; recall met noninferiority. Voluntary use and nonrandom allocation limit causal inference. Normal-tagged examinations took less reading time than examinations not tagged normal within the AI group; control reading time was not measured. This was not a trial of removing human readers.
AITIC: Transpara 1.7; 31,301 women aged 50–71 at one Spanish center, with digital mammography or digital breast tomosynthesis (Elías-Cabot et al., 2026). Paired comparison: every participant received standard double reading. The additional AI strategy classified scores 1–7 as normal and sent scores 8–10 to AI-supported double reading. There was no arbitration; any reader could trigger recall. The AI strategy reduced reading workload by 63.6% and increased cancer detection from 6.3 to 7.3 per 1,000. Recall was 14.8% higher and failed noninferiority. Results differed by imaging modality. Because standard reading also assessed everyone, this design cannot establish the interval-cancer rate of an independently deployed partially autonomous pathway.

United Kingdom technical-feasibility study: A multicenter study evaluated 115,973 mammograms retrospectively and then deployed the system prospectively in a noninterventional feasibility phase involving 9,266 cases at 12 sites. The prospective phase established technical feasibility and exposed distribution shift requiring recalibration; it did not compare an AI-directed screening pathway with usual care (Kelly et al., 2026).

US single-reading implementation: A real-world pre/post comparison evaluated Transpara v1.7.4-A during US single-reading digital breast tomosynthesis screening (24,520 examinations after implementation versus 21,630 before). After AI implementation, recall and cancer detection rates did not significantly change; AI risk strata correlated with recall; and the tool missed 15 of 171 cancers, including all cancers scored as low-risk (Ambinder et al., 2026). This US single-reading result is a counterpoint to European double-reading programs already on this page (MASAI, PRAIM, AITIC); it does not show that Transpara fails in every workflow. FDA cleared Transpara 2.1.0 through a Special 510(k) substantial-equivalence determination on August 14, 2026 (K260898, Class II, product code QDQ); that increment is not a new intended-use class, and Ambinder studied v1.7.4, so local evidence and labeling must be paired to the version actually deployed (FDA K260898 record).

Version-to-version variation within one product: Two versions of the same commercial model scored 117,709 BreastScreen Norway examinations (2009–2018) retrospectively. The highest malignancy score went to 87.1% (642/737) of screen-detected cancers with Transpara 1.7 and 93.5% (689/737) with version 2.1. The aggregate share of interval cancers at that score was unchanged (45.0% versus 44.5%), but 11.5% (23/200) of interval cancers rose to it and 12.0% (24/200) fell from it. The share of all examinations at the highest score moved from 10.1% to 10.7%, the share of the 3,071 false positive screening results at that score from 25.2% to 31.4%, and AUC with screen-detected and interval cancers as positive cases from 0.908 to 0.928 (Larsen et al., 2026). MASAI and AITIC evaluated version 1.7 (Lång et al., 2023; Elías-Cabot et al., 2026). A version update changes which examinations a risk-adapted pathway routes to double reading, so it needs its own local evaluation. The study used a single mammography equipment manufacturer at two centers and scored archived examinations against recorded program outcomes rather than reader decisions made with AI support. The authors state that the resulting recall and false positive screening rates, and with them the net benefit, remain unknown.

Limitations: - Performance varies by breast density, age, implants - Most training data from screening populations (may not generalize to diagnostic mammography) - Racial bias: many systems trained predominantly on white women

Device and threshold variation: A retrospective multi-site, multi-vendor analysis of 22,673 screening mammograms found that the same commercial AI’s scores and performance varied substantially across devices. Device-specific thresholds lowered consensus-conference workload by nearly one-third versus a global threshold, but performance on one device remained poor after recalibration (Blum et al., 2026). This finding supports local acceptance testing on each scanner and vendor; it does not mean mammography AI is unusable across devices.

Applying and reassessing the evidence: Before using a study to support a local decision, record the proposed population, screening or diagnostic use, imaging modality and equipment, software version and threshold, human-reader configuration, and recall or arbitration process. Identify each mismatch with the evaluated pathway. Reopen the assessment when these change, when interval-cancer follow-up or contradictory evidence becomes available, or when a source is corrected or retracted. Use the handbook’s evaluation framework to define local outcomes and acceptance criteria; a published detection gain alone does not establish local benefit.

Head CT AI

Intracranial Hemorrhage (ICH) Detection: - Commercial examples: Aidoc, Viz.ai, and RapidAI. Each product’s exact FDA record, finding, image type, user, and workflow must be verified separately. - Evidence: Worklist reprioritization cut median time to diagnosis from 512 to 19 minutes for outpatient head CT Arbabshirani et al., 2018 - Deployment: Widely adopted in EDs and trauma centers - Performance: Varies by hemorrhage type, size, acquisition, population, and operating threshold; no class-wide sensitivity should be assumed

Large Vessel Occlusion (LVO) Stroke: - Purpose: Identify strokes requiring thrombectomy - Systems: Viz.ai LVO, RapidAI, and Brainomix are examples, but evidence must remain tied to the exact product and workflow studied - Viz.ai observational evidence: In an 82-patient pre/post implementation study, overall door-to-groin time changed from 50 to 39 minutes without statistical significance (P=0.066). The prespecified off-hours subgroup improved from 157 to 95 minutes (P=0.009) (Figurelle et al., 2023) - Viz.ai randomized workflow evidence: A stepped-wedge cluster-randomized trial involving 243 patients treated with endovascular thrombectomy found an adjusted 11.2-minute reduction in door-to-groin time and a 9.8-minute reduction in CT-to-treatment time. It did not show a significant improvement in 90-day functional independence (Martinez-Gutierrez et al., 2023) - Workflow: Automated alerts to stroke team + neurointerventional

These Viz.ai studies do not establish equivalent effects for RapidAI, Brainomix, or every implementation of stroke-triage AI.

Current evidence boundary: A 2025 Radiology systematic review of 380 deep-learning studies in adult acute ischemic stroke found that 45.0% focused on lesion segmentation, 33.9% on classification or triage, 8.2% on outcome prediction, and 3.9% on generative AI or large language model applications (Jiang et al., 2025). For radiologists, the strongest evidence still supports narrow imaging and workflow triage tasks, not comprehensive stroke management or autonomous outcome prediction.

ASPECTS Scoring (ischemic stroke extent): - Purpose: Quantify early ischemic changes to guide treatment decisions - Benefit: Standardized scoring (high inter-rater variability in manual scoring) - Commercial example: RapidAI ASPECTS; verify the current authorized version and intended use in the FDA record

Limitations across head CT AI: - Artifact sensitivity (motion, beam hardening) - Subtle bleeds (small subarachnoid hemorrhage) may be missed - Overcalling: small calcifications, imaging artifacts flagged as hemorrhage

Chest CT AI

Pulmonary Embolism (PE) Detection: - Performance: High sensitivity for large, proximal PE - Challenge: Subsegmental PE (clinical significance debated, hard to detect) - Commercial examples: Aidoc and Avicenna.AI; product-specific labeling and evidence apply

Lung Nodule Detection and Characterization: - Screening context: Lung cancer screening CT - Performance: Comparable to radiologists for significant nodules - Systems: Many available, integration with lung-RADS reporting - Limitation: Sub-4mm nodules, part-solid nodules

CT-first lesion risk (2026): Peng et al. trained a cohort-aware model on 28,105 cases across LIDC-IDRI, Lung-PET-CT-Dx, NLST, and NSCLC-Radiomics, using CT as the default input and PET only when available (Peng et al., 2026). Internal malignancy-classification AUC was 0.955–0.978 (Brier 0.038–0.056). External Dice was 0.878; patient-level survival C-index was 0.709 in NLST and 0.857 on NSCLC-Radiomics. This is retrospective discrimination and spatial agreement on public cohorts, not evidence that deploying the model reduces unnecessary biopsy, interval cancer, or lung-cancer mortality. Prospective multicenter validation is absent.

Opportunistic Esophageal Malignancy Detection:

Zhou and colleagues report EAGLE, a two-stage model for esophageal malignancy on noncontrast chest CT, with external testing across eight centers in three countries (n = 11,466) at about 98.5% specificity and about 90% sensitivity for cancer (lower for HGIN), plus reader-assist gains, real-world false-positive reduction after calibration, and a prospective hospital run with PPV about 42% (Zhou et al., 2026). LDCT cohorts and an endoscopy-program simulation explore piggyback screening and triage, but do not establish population mortality benefit. Multicenter diagnostic performance and implementation studies: not an EC-mortality RCT; several industry-affiliated authors; closed code/data.

Incidental Findings: - Examples: Vertebral compression fractures, aortic aneurysms, coronary calcium - Benefit: Identifies findings on studies ordered for other indications - Risk: Overdiagnosis, patient anxiety, unnecessary workups

### Abdominal CT AI {#sec-abdominal-ct-ai}

Multi-Condition Triage:

Abdominal CT has historically required separate AI algorithms for each finding (one for appendicitis, another for bowel obstruction, a third for solid organ injury). Standard first-in, first-out reading queues mean urgent abdominal findings may wait hours during ED crowding.

  • Aidoc CARE Multi-triage CT Body: FDA cleared K252970 on January 7, 2026, for concurrent triage and notification of 11 specified findings on adult CT examinations (FDA K252970 record)
  • K252970 findings: Diverticulitis, abdominal-pelvic abscess, appendicitis, intestinal ischemia or pneumatosis, obstructive renal stone, small-bowel obstruction, large-bowel obstruction, spleen injury, liver injury, kidney injury, and pelvic fracture
  • Architecture: The clearance covers a multi-finding triage system. Other Aidoc findings have separate authorization records and cannot be counted as part of K252970 merely because they appear in the same commercial platform.

Regulatory evidence boundary:

  • The FDA labeling describes indication-specific performance testing. It does not support treating one mean sensitivity or specificity as the performance of every finding in every population.
  • The software runs in parallel with standard care, provides a preview and notification, is not intended for diagnosis, and must not be used to deprioritize studies.
  • Vendor comparisons with other products require independent, head-to-head evidence and should not be inferred from a clearance summary.

K252970 authorizes a defined triage function for 11 findings. It does not authorize autonomous abdominal CT interpretation.

Limitations:

  • Pivotal study was retrospective and blinded; prospective real-world performance data are pending
  • Performance for patients with multiple simultaneous acute conditions or atypical presentations not yet characterized
  • Integration with existing PACS workflows varies by institution

Liver malignancy CE-CT (additional reader): A 2026 Nature Medicine multicenter study plus single-arm trial developed LiON, a multiphase CE-CT additional reader for patient-level liver-malignancy diagnosis, distinct from K252970’s concurrent triage of acute liver injury (Zhang et al., 2026). In 10,333 routine-practice patients (NCT07153783), standalone malignancy AUC was 0.952 (95% CI 0.942–0.961), and AI-human review identified 51 previously overlooked lesions, including 15 malignancies. The trial is China-only, unmasked, and single-arm; it is not an RCT, not FDA-cleared, several authors are Alibaba employees, and it does not license skipping radiologist over-read.

Zhang and colleagues report RADAR, a vision-language model for broad contrast-enhanced abdominal CT interpretation trained on 424,911 examinations with anatomy-aware report supervision (more than 15 million anatomy-specific image-text pairs), reaching mean AUC about 0.913 across 146 findings, with external multi-center AUC about 0.895 and emergency-case AUC about 0.904 on more than 27,000 exams, plus reader-assist gains in sensitivity and reading time (Zhang et al., 2026). Research generalist VLM for abdominal CT (RADAR), not the Nat Med LiON liver model; not FDA-cleared autonomous abdominal CT interpretation; several Alibaba DAMO-affiliated authors; reader-study details synthesized from secondary coverage pending full-text lock.

Generalist Multimodal Radiology Foundation Models (Emerging)

A recent NEJM AI study introduced MedVersa, a generalist multimodal imaging foundation model trained on 29 million instances compiled from 91 public datasets. Unlike task-specific models built for a single finding, MedVersa was evaluated across classifications, segmentations, visual question answering, and report-generation workflows from one unified architecture (Zhou et al., NEJM AI, 2026).

Why this matters for radiologists:

  • Chest radiograph report equivalence: In blinded evaluation, board-certified radiologists rated developer-generated and human-written reports as clinically equivalent in 64% of cases overall (91% for normal chest radiographs).
  • Workflow signal: Developer-conducted user studies reported reduced report-writing time and fewer clinically relevant discrepancies versus standard workflows. Independent replication is pending.

Caution before clinical adoption:

  • Results are promising but primarily benchmark and study-environment based; independent prospective validation across heterogeneous health systems remains necessary.
  • Generalist performance can mask modality- or task-specific failure modes that become apparent only in local deployment.
  • This is not equivalent to broad regulatory authorization for autonomous radiology reporting.

CLEAR benchmark: A 2026 Nature Biomedical Engineering study introduced an auditable chest-radiograph foundation model trained on more than 870,000 image-report pairs from 239,391 patients and evaluated on four physician-annotated external datasets (Han et al., 2026). It extends the evidence base for interpretable imaging models, but it remains a benchmark study rather than a prospective clinical-deployment trial. External test performance is not a substitute for evidence that a model improves care when inserted into a live workflow.

A Science abdominal CT generalist, RADAR, extends the same report-supervised vision-language pattern beyond chest-centric foundation models (Zhang et al., 2026; see Abdominal CT AI).

Acar et al. evaluated GPT-5.2, Gemini 3 Pro, and Claude Opus 4.5 on 252 paywall-protected RSNA Case Collection intracranial and spinal cases, comparing open-ended generative diagnosis with multiple-choice recognition (Acar et al., 2026). Primary diagnostic accuracy was 51.6–62.7% open-ended (GPT-5.2 59.1%, Gemini 3 Pro 62.7%, Claude Opus 4.5 51.6%) and rose to 75.4–81.0% when answer options were provided (15.1–23.8 percentage-point gap). A neurosurgical panel graded clinically significant errors in 17.8–26.9% of all cases. Open-ended accuracy and the clinical-error burden do not support autonomous neuroradiology diagnosis; multiple-choice scores overstate generative performance and are not a deployment endpoint.

MRI AI

Prostate MRI: - PI-RADS Scoring: AI-assisted lesion detection and classification - Evidence: AI achieved AUROC 0.91 vs. 0.86 for radiologists in a 10,207-exam confirmatory study (Saha et al., Lancet Oncology, 2024) - Benefit: Reduces reading time, standardizes scoring - Limitation: Peripheral zone vs. transition zone performance differs

Brain MRI: - Multiple sclerosis lesion quantification: Automated lesion segmentation and volume - Tumor segmentation: Glioma characterization, treatment response assessment - Status: Research tools becoming clinically available

Cardiac MRI: - Automated chamber quantification: Ejection fraction, volumes, strain - Benefit: Faster, standardized measurements - FDA-cleared: Circle Cardiovascular Imaging, Arterys

Liver MRI (non-contrast FLL): A multicenter npj Digital Medicine study introduced PRISM, a three-stage deep-learning system for focal liver lesions on non-contrast MRI, developed and tested in 12,823 patients from 9 institutions (Dong et al., 2026). Mean accuracy was 0.981 for benign-versus-malignant classification and 0.859 for four-subtype classification across five test cohorts; a 13-radiologist reader study reported +4.4% (Protocol B vs A) and +12.1% (Protocol E vs C) accuracy with 30.5% (10.9 s, 95% CI 9.8–12.0) and 66.6% (61.8 s, 95% CI 58.9–64.7) shorter interpretation time. In 1,147 consecutive prospective patients, PRISM labeled 79.3% low-risk at NPV 99.1%. This is assisted NC-MRI diagnosis in Chinese centers, not autonomous liver MRI and not a claim that contrast MRI can be omitted.

MRI Reconstruction: - Purpose: Accelerated imaging (shorter scan times) - Method: Deep learning fills in undersampled k-space data - FDA-cleared: Several vendors integrating into scanners (e.g., Philips SmartSpeed Precise, 510(k) July 2025; up to 3x faster scans) - Benefit: Reduced scan times improve patient experience, throughput

Evidence gap in commercial MRI AI: A systematic review of 14 commercial MRI AI reconstruction products found that 29% had zero peer-reviewed validation studies, and only 19% of published articles prospectively demonstrated clinical impact (Fransen et al., European Radiology, 2025). No studies analyzed hallucinatory artifacts; no economic analyses existed. This mirrors the broader pattern seen with FDA clearances: market deployment is outpacing clinical evidence.

Limitations: - Longer acquisition times than CT make AI deployment more challenging - Sequence variability across institutions affects generalizability - Artifacts from implants, motion difficult for AI to handle

Other Modalities

Ultrasound: - Cardiac echo: Automated view identification, EF calculation, valve assessment - OB/GYN: Fetal biometry, anomaly screening - Status: Emerging, especially for point-of-care ultrasound guidance

Nuclear Medicine: - PET/CT: Automated lesion detection, SUV quantification - Bone scans: Metastasis detection, reporting automation

Interventional Radiology: - Real-time guidance: Needle/catheter tracking - Dose reduction: AI-enhanced low-dose fluoroscopy - Status: Early-stage development

The “Will AI Replace Radiologists?” Debate:

2016 Prediction: Geoffrey Hinton (deep learning pioneer): “It’s quite obvious that we should stop training radiologists”

2024 Reality: AI has not replaced radiologists. AI is augmenting radiologists.

Why Radiologists Remain Essential:

  1. Complex reasoning across modalities: AI excels at narrow tasks, struggles with pulling together information across studies, clinical context, prior imaging

  2. Incidental findings: AI trained for specific task (e.g., PE detection) may miss unrelated critical findings (e.g., aortic dissection, lung cancer)

  3. Clinical integration: Radiologists communicate with referring physicians, provide consultative input beyond reports

  4. Edge cases and artifacts: AI performs poorly on unusual presentations, technical limitations, patient factors

  5. Accountability: Responsibility must be defined through product labeling, institutional policy, professional duties, contracts, and applicable law. FDA authorization alone does not allocate civil liability.

  6. Quality assurance: Someone must validate AI outputs, identify failures, oversee system performance

What Has Changed:

  • Radiologists increasingly use AI as assistive tools (second reads, quantification, prioritization)
  • Workflow efficiency improvements (faster reads for straightforward cases)
  • New roles: radiologist-informaticists curating datasets, validating AI, implementing systems
  • Training emphasis: understanding AI capabilities/limitations, effective human-AI collaboration

Future Trajectory:

One plausible trajectory is human-AI partnership where: - AI handles routine detection, quantification, triage - Radiologists focus on complex cases, integration, communication, oversight - Radiologist supply challenges (shortages in many regions) partially addressed by AI efficiency gains

Quantitative Workforce Projections

A task-based scenario model posted as a preprint in December 2025 projected a 33% reduction in radiologist hours worked over five years under its central assumptions (range: 14–49%), with the largest modeled effects from AI-assisted report drafting and delegation of selected mammography and radiography studies (Langlotz, 2025, preprint). These are modeled scenarios, not observed workforce changes or a validated forecast.

Key drivers of projected productivity gains:

AI Application Modality Affected Projected Effect
Report drafting All modalities 20% reduction in interpretation time (range 10–30%)
Study delegation Mammography 50% of studies may not require human interpretation (range 30–60%)
Study delegation Radiography 40% of outpatient studies modeled as delegable (range 30–50%)
Automated protocoling CT, MR 60% reduction in protocoling time

Model conclusion, not an observed outcome: The preprint argues that productivity gains could be offset by continued growth in imaging volume and a relatively static workforce. The direction and magnitude of actual labor effects will depend on demand, payment, regulation, product performance, and whether institutions use saved time to increase volume, improve quality, or reduce staffing.

Important caveats:

  • Preprint status: This analysis has not been peer-reviewed. The model combines published evidence with author judgment where data gaps exist.
  • Author conflicts of interest: The author holds equity in multiple radiology AI companies (Bunkerhill Health, Sirona Medical, Whiterabbit.ai, and others), relevant context for interpreting projections about AI adoption benefits.
  • Assumptions may not hold: Implementation will occur in phases, not all institutions will adopt all tools, and some applications (e.g., delegating normal studies to AI-only interpretation) may face significant resistance.
  • U.S.-centric analysis: Other countries with different imaging volumes, practices (e.g., double-reading of mammograms), and workforce dynamics will see different effects.

Evidence supporting key assumptions:

The delegation and report drafting projections draw on prospective studies:

  • Mammography delegation: AI triage can identify 35-63% of screening mammograms as normal without requiring radiologist review, with equivalent or improved cancer detection (Hickman et al., 2023; Leibig et al., 2022)
  • Radiography delegation: In a retrospective rule-out study, AI classified 25–53% of unremarkable chest radiographs as normal at a prespecified sensitivity of at least 98% (Plesner et al., 2024). This does not establish the safety of autonomous delegation in routine care.
  • Report drafting productivity: Generative AI-assisted radiograph reporting shows 15% improvement in documentation efficiency without compromising accuracy (Huang et al., 2025)

FDA-Authorized Imaging AI Devices:

FDA’s periodically updated AI-enabled medical device list remains dominated by radiology devices and is not a comprehensive inventory of every authorized AI-enabled device (FDA AI-Enabled Medical Devices).

The table is an orientation map, not an authorization register. A platform name can cover multiple devices, indications, and versions, so the primary FDA record remains controlling.

Application Vendor Examples Regulatory note
Intracranial hemorrhage detection Aidoc, Viz.ai, RapidAI Verify each finding, modality, workflow, and software version
Large vessel occlusion stroke Viz.ai, RapidAI, Brainomix Evidence below is Viz.ai-specific unless another source is named
Pulmonary embolism Aidoc, Avicenna.AI Verify whether the function is detection, triage, or notification
Pneumothorax Oxipit ChestLink, Lunit INSIGHT CXR Autonomous and assistive claims require different records
Breast cancer detection iCAD, Hologic, Lunit Confirm screening versus diagnostic use and reader configuration
Diabetic retinopathy IDx-DR, EyeArt IDx-DR received De Novo authorization under DEN180001; do not call the pathway 510(k) clearance
Cardiac MRI quantification Arterys, Circle CVI Confirm the authorized measurement and acquisition context
Bone age assessment 16bit, Carestream Confirm intended users, age range, and reference standard

Clinical evidence reported in FDA documents

A systematic review identified 950 FDA-authorized AI or machine-learning devices through June 2024, including 723 radiology devices. Among 717 radiology authorization documents available for analysis, the authors reported the following evidence characteristics (Sivakumar et al., 2025):

  • Only 5% underwent prospective testing before clearance
  • Only 8% included a human-in-the-loop validation component
  • Only 29% incorporated any clinical testing at all

These document-level findings do not mean that every device lacking a described clinical study was never tested. They show that clinical, prospective, and human-operator evidence was uncommon in the publicly available authorization documents reviewed.

Post-market safety: Recalls and adverse events

A companion study of the same 950-device cohort found that 60 devices (6.3%) were involved in 182 recall events (Lee et al., 2025):

  • 43.4% of recalls occurred within the first 12 months of clearance (roughly double the rate for all 510(k) devices)
  • Devices whose public authorization documents did not describe clinical validation had a reported 2.8-fold higher recall risk, an association that does not by itself establish causation
  • Publicly traded companies accounted for 53% of devices but 91.8% of recall events
  • Primary failure modes: diagnostic/measurement errors (109 recalls), functionality delays (44 recalls)

These studies identify evidence-transparency and post-market safety concerns, but they do not establish that limited premarket clinical testing caused each recall. A separate 2026 cohort analysis identified recalls in 43 of 903 AI-enabled devices (4.8%) through its study cutoff, with a median 458 days from authorization to recall. Its estimate linking the 510(k) pathway to recall was imprecise (hazard ratio 1.39; 95% credible interval 0.84–3.52) (Ren et al., 2026).

FDA authorization establishes that a device met the applicable regulatory pathway. It does not establish transportability, local clinical utility, or freedom from post-market failure.

FDA Regulatory Pathways:

  • 510(k) clearance: Substantial equivalence to existing device (most radiology AI)
  • De novo: Novel device, no predicate (e.g., IDx-DR first autonomous diagnostic)
  • PMA: Highest scrutiny (rare for software)

FDA lifecycle policy: - Predetermined change control plans can authorize specified future modifications when the plan is included in the marketing submission and reviewed by FDA (FDA PCCP guidance) - Version-specific labeling and authorization records remain essential; a commercial platform name does not establish that every version or feature is authorized - Real-world performance monitoring should be risk-proportionate and linked to predefined action rules

Implementation Challenges:

Technical Integration: - PACS integration: AI must fit radiology workflow (not separate system) - HL7/FHIR standards: Data exchange between EHR, PACS, AI - DICOM compatibility: Handle different scanner manufacturers, protocols - Network bandwidth: Some AI requires cloud processing → upload/download delays

Workflow Disruption: - Alert fatigue: Too many AI flags → radiologists ignore - Optimal threshold: Balance sensitivity (catch everything) vs. specificity (minimize false alarms) - Worklist changes: AI triage changes reading order → radiologist adaptation needed

Validation and Monitoring: - Performance varies by context: Local validation essential before deployment - Continuous monitoring: Detect performance drift, scanner changes, patient population shifts - Failure mode identification: Document when/how AI fails at your institution

Radiologist Training: - Understanding AI: Capabilities, limitations, failure modes - Interpreting AI outputs: When to trust, when to override - Providing feedback: Improving AI through error identification

Economic and Reimbursement:

A rapid systematic scoping review screened 8,013 records and included 140 studies of diagnostic-radiology AI. It found extensive technical research but sparse evidence on implementation, staff and patient experience, quantitative workflow effects, and cost (Lawrence et al., 2025). A separate review screened 1,879 publications and included 21 economic evaluations; estimated value varied with task, prevalence, staffing, payment, licensing, and whether the analysis modeled rather than measured deployment (Molwitz et al., 2026).

A time saving is not a cost saving unless the surrounding workflow can use it.

Costs: - Licensing fees: Annual per-scanner or per-study fees - Infrastructure: Hardware (GPUs), networking, storage - Personnel: Informaticists, IT support, radiologist time for validation

Reimbursement: - Procedure-specific coding: Coding and coverage differ by application and change over time; verify current payer and CPT guidance - Separate payment is not guaranteed: Economic value may depend on measurable efficiency, quality, access, or outcome gains - Value-based care: AI may support quality metrics (reduced errors, faster turnarounds)

ROI Measurement Considerations: - Net change in interpretation and downstream clinical time - Recall, addenda, repeat imaging, escalation, and error-review burden - Infrastructure, integration, monitoring, retraining, and downtime costs - Access, turnaround, quality, safety, and patient-relevant outcomes

Reduced malpractice claims, referral growth, or cost savings should be treated as hypotheses until measured in the institution’s own workflow.

Medical Liability:

Potential Duties on Both Sides of Adoption:

Legal scholarship has discussed claims arising from both inappropriate reliance on AI and failure to use a tool that later becomes part of accepted practice (Mello & Guha, 2024). That possibility is not a universal legal rule, and no product becomes the standard of care merely because it is FDA-authorized or widely marketed.

Mammography is a useful test case because multiple comparative trials now exist, but legal duties still depend on jurisdiction, available evidence, professional guidance, local resources, and the facts of a particular case. Trial evidence cannot predict how a court will allocate responsibility.

Evidence on AI and Perceived Liability:

A 2025 NEJM AI experiment examined mock-juror judgments in hypothetical radiology cases. It measured perceptions, not actual litigation outcomes or governing law (Bernstein et al., 2025). A follow-up vignette experiment involving 282 participants reported that 74.7% found a duty-of-care breach after a single review of AI-flagged imaging versus 52.9% after a documented independent review followed by AI review (Bernstein et al., 2026). These results generate hypotheses about workflow and documentation; they do not prove that a specific sequence prevents liability. See Liability and Malpractice for full analysis.

Standard of Care Implications

Standard of care is fact-specific and jurisdiction-specific. FDA authorization does not determine civil liability, and adoption rates alone are not dispositive. Relevant considerations may include:

  • Adoption rates at comparable institutions (academic medical centers, community practices)
  • Professional society guidelines recommending specific AI applications
  • Malpractice insurers requiring or incentivizing AI use
  • Published evidence demonstrating improved outcomes (reduced miss rates, faster treatment)

Several mammography trials report noninferior or improved detection under defined AI-assisted workflows with lower reading workload (Lång et al., 2023). Whether those results influence a legal standard in a particular setting remains an open question.

Key Liability Questions:

  • Who had which duty if AI misses a finding? The answer can involve clinicians, institutions, developers, manufacturers, and contractual parties.
  • Does using AI change standard of care? (Emerging legal question, varies by jurisdiction)
  • Is a radiologist obligated to use an available tool? No universal duty follows from availability alone.
  • What documentation is required? Product labeling, institutional policy, applicable law, and clinical relevance determine the answer.
  • Can a claim allege failure to use AI? Such a claim is possible, but success cannot be predicted without the governing law and case facts.

Risk Mitigation:

  • Thorough local validation before clinical deployment
  • Clear protocols for radiologist review (AI is assistive, not autonomous for most applications)
  • Documentation proportionate to clinical relevance, product labeling, and institutional policy
  • Informed consent where appropriate
  • Verify malpractice insurance covers AI use, including both errors made while using AI and allegations of failing to use available AI
  • Request specific policy language confirming coverage for AI-assisted clinical decision making
  • Monitor professional society guidance, product labeling, local performance, and applicable law

Cross-reference: For comprehensive legal framework on AI liability, standard of care evolution, and documentation strategies, see Physician AI Liability and Regulatory Compliance.

Evidence-Based Assessment:

Systematic Reviews and Meta-Analyses:

  • Chest X-ray AI: Systematic review of 46 studies found standalone AI models performed comparably to or better than radiologists, with AUC values of 0.82–0.96 depending on pathology, though heterogeneity remains high (Ahmad et al., 2023)

  • Breast cancer detection: AI achieves AUC 0.883 (UK dataset), comparable to radiologists (McKinney et al., 2020, Nature)

  • Intracranial hemorrhage: Retrospective assessment of a commercial detection algorithm (200 ICH, 102 non-ICH patients) reported 93% sensitivity and 93% specificity for hemorrhage slices Rava et al., 2021

Randomized Controlled Trials (RCTs):

  • Mammography (MASAI trial): The largest completed RCT of AI-supported mammography screening randomized 105,934 Swedish women to AI-assisted vs. standard double-reading. Interim results showed non-inferior cancer detection with 44% reduction in screen-reading workload (Lång et al., 2023). Full two-year follow-up found AI-supported screening improved sensitivity (80.5% vs. 73.8%, P=0.031) at identical specificity (98.5%), with 12% fewer interval cancers and 27% fewer aggressive (non-luminal A) cancers in the AI group, though these reductions did not reach statistical significance in this non-inferiority design (P=0.41) (Gommers et al., 2026, The Lancet)

  • Mammography triage (AITIC): A prospective paired noninferiority study in Spain’s Córdoba screening program (NCT04949776) tested AI-based triage to classify examinations as low risk and omit them from radiologist reading. AI triage detected 7.3 versus 6.3 cancers per 1,000 and reduced reading workload by 63.6%, but recall was 14.8% higher and failed the prespecified noninferiority endpoint (Elías-Cabot et al., 2026).

  • CXR triage simulation: A UK retrospective simulation estimated that AI-based prioritization could reduce reporting delay for critical findings from 11 to 3 days. It did not prospectively deploy the pathway or measure patient outcomes (Annarumma et al., 2019).

  • CXR pathway trial: In LungIMPACT, a prospective multicenter randomized study of 93,326 chest radiographs, AI-supported management did not reduce time to chest CT or time to lung cancer diagnosis (Woznitza et al., 2026). A favorable prioritization simulation and a null randomized pathway trial answer different questions; the simulation should not be treated as proof of clinical benefit.

Prospective Validation Studies:

  • IDx-DR diabetic retinopathy: Prospective validation at 10 primary care sites showed 87.2% sensitivity, 90.7% specificity Abramoff et al., 2018 - led to FDA clearance

  • Viz.ai LVO stroke: Figurelle et al. reported an 82-patient pre/post implementation study. The overall door-to-groin comparison was not statistically significant, while the off-hours subgroup improved from 157 to 95 minutes (Figurelle et al., 2023). A separate stepped-wedge cluster-randomized trial found shorter process times but no significant improvement in 90-day functional independence (Martinez-Gutierrez et al., 2023). Both studies were Viz.ai-specific.

Common Failure Modes:

Artifact Sensitivity: - Motion artifacts, metal artifacts, beam hardening - Skin folds, ECG leads, monitoring equipment mimicking pathology - Positioning issues (e.g., rotated chest X-ray)

Edge Cases: - Rare diseases not well-represented in training data - Unusual presentations of common diseases - Pediatric patients if trained on adults - Post-surgical anatomy

Confounding by Context: - Detecting “portable” keyword not lung findings - Learning ICU location as proxy for disease severity - Scanner-specific image characteristics

Segmentation Errors: - Misidentifying anatomy (aorta vs. pulmonary artery) - Including/excluding wrong structures (lung nodule vs. vessel) - Partial volume effects

Overconfidence: - High confidence scores for incorrect predictions - No “uncertainty” measure for out-of-distribution cases

Best Practices for Radiology AI Deployment:

Implementation Checklist

Pre-Deployment: - Literature review: published validation for your use case? - Vendor vetting: FDA clearance? Peer-reviewed publications? Customer references? - Local pilot: test on retrospective cases from YOUR institution - Failure mode analysis: identify when/how system fails on your data - Workflow design: how will AI integrate into radiologist workflow?

Deployment: - Radiologist training: capabilities, limitations, workflow integration - Technical validation: PACS integration, network performance, uptime - Initial monitoring: high-frequency performance checks - Feedback mechanism: radiologists report errors, unexpected behaviors

Post-Deployment: - Continuous monitoring: performance drift, false positive/negative rates - Quarterly reviews: aggregate performance, user feedback, failure patterns - Version control: document AI updates, re-validate when model changes - Outcome tracking: impact on reporting times, error rates, clinical outcomes


Human Factors: Training the Next Generation with AI

As radiology AI enters training environments, a critical question emerges: what happens to radiologists who learn with AI assistance from the start? Direct longitudinal evidence in radiology remains limited. Human-factors experiments nevertheless show that presentation order, explanations, user expertise, and erroneous AI advice can change decisions, which makes curriculum design a patient-safety question rather than a software-training exercise.

Cognitive De-Skilling in AI-Assisted Training

The Evidence Boundary:

No cited longitudinal study establishes that radiology residents trained with CAD acquire permanently weaker independent interpretation skills. A 2025 observational colonoscopy study reported lower adenoma detection during non-AI examinations after clinicians had worked with AI assistance, a risk signal that does not prove permanent deskilling or transfer directly to radiology (Budzyń et al., 2025). A prospective single-center observational study found that non-AI detection performance was maintained after CADe implementation (Okumura et al., 2026). A prospective multicenter pragmatic trial involving 13 endoscopists and 5,013 colonoscopies found no significant overall upskilling or deskilling after CADe removal, although individual trajectories varied (Pedersen et al., 2026).

The defensible conclusion is that deskilling should be monitored, not that it has already been proved in radiology training.

Why De-Skilling Happens:

Automation bias: The tendency to over-rely on automated systems and under-weight contradictory information from other sources. Automation bias is stronger in less experienced clinicians, who lack the pattern recognition library to confidently override AI suggestions.

Reduced deliberate practice: Learning diagnostic radiology requires repeated exposure to cases with immediate feedback. When AI provides the answer before the trainee has fully reasoned through the case, the learning opportunity is lost.

Skill fade without reinforcement: Expertise requires continuous use. If AI handles routine detection tasks, radiologists lose practice opportunities for foundational skills.

Practical Implications for Training Programs:

Radiology residency programs must now balance AI proficiency with independent skill development. Suggested approaches:

Illustrative rotation structure to preserve independent skills:

The following structure is a curriculum design option, not an evidence-based mandate or a universal postgraduate-year rule:

  1. AI-free rotations (R1-R2 years): Early trainees develop pattern recognition without AI assistance. Build foundational skills on normal variants, common pathology, artifact recognition.

  2. AI-assisted rotations (R3-R4 years): After foundational skills solidify, introduce AI as decision support. Trainees learn to integrate AI outputs with independent assessment.

  3. AI-off practice sessions: Regular exercises where trainees interpret studies without AI, then compare to AI outputs. Identifies cases where human and AI disagree, forcing critical reasoning.

Workflow modifications:

  1. Interpret first, then view AI: Trainees form independent impression before reviewing AI flags. Prevents anchoring to AI assessment.

  2. Override documentation: When trainees override AI, require documentation of reasoning. Builds habit of critical evaluation rather than reflexive acceptance.

  3. AI failure mode review: Quarterly conference reviewing cases where AI failed. Builds pattern recognition for AI limitations.

Assessment considerations:

Board examinations and competency assessments currently test independent interpretation without AI. Trainees who learned exclusively with AI assistance may struggle on exams that remove the assistive technology they rely on in clinical practice.

The Paradox of AI-Enhanced Training:

Training programs face a fundamental tension. AI improves diagnostic accuracy when used appropriately, so withholding AI from trainees seems to deny them valuable decision support. Yet providing AI throughout training may prevent development of the independent pattern recognition skills needed to use AI effectively.

The solution requires staged competency development:

Phase 1 (Foundation): Build and document unassisted diagnostic competence. This provides a baseline for detecting when an AI output is implausible.

Phase 2 (Integration): Introduce AI as assistive technology after independent skills solidify. Trainees learn to integrate AI suggestions with their own assessments, developing critical evaluation skills.

Phase 3 (Mastery): Senior residents use AI as experienced clinicians do, as one input among many informing final interpretation.

Programs can adapt these phases to local competencies and workflows. The relevant safeguard is demonstrable independent competence plus the ability to recognize and escalate AI failure, not adherence to one fixed sequence.

AI Anchoring Bias in Clinical Practice

Beyond training concerns, AI introduces anchoring bias even in experienced radiologists. When AI flags a region as suspicious, radiologists anchor to that assessment rather than conducting fully independent evaluation.

The Mechanism:

CAD systems trained on biopsy-confirmed cancers learn to detect “lesions suspicious enough to biopsy” rather than true cancer. The training data contains selection bias: biopsied lesions are enriched for features that triggered clinical suspicion, not necessarily features that distinguish malignant from benign pathology.

When a radiologist reviews an AI-flagged region, cognitive anchoring occurs:

  1. AI presents confident assessment (e.g., “92% probability of malignancy”)
  2. Radiologist anchors to AI confidence level rather than independently evaluating imaging features
  3. Confirmation bias activates: Radiologist selectively notices features supporting AI assessment, discounts contradictory features
  4. Result: Overcalling borderline findings flagged by AI, undercalling regions AI missed

Evidence from Mammography CAD:

First-generation mammography CAD increased recall rates without improving cancer detection (Lehman et al., 2015). Radiologists followed CAD prompts for equivocal findings they would have dismissed without CAD, leading to unnecessary biopsies.

The problem was not CAD sensitivity (which was high) but specificity (which was poor). Radiologists anchored to CAD confidence scores and upgraded BI-RADS assessments for CAD-flagged findings that lacked independent suspicion.

Mitigation Strategies:

Workflow design matters: Two approaches to AI integration:

  1. Concurrent mode: AI displays results alongside images during interpretation
    • Risk: Radiologist sees AI assessment before forming independent opinion, anchoring occurs
    • Advantage: Efficient, fits existing workflow
  2. Second-reader mode: Radiologist interprets study independently, then reviews AI outputs
    • Advantage: Prevents anchoring, preserves independent reasoning
    • Disadvantage: Adds time, requires workflow adjustment

Human-factors studies suggest that workflow order can affect reliance, but the direction and magnitude depend on task and interface. In a study of 140 radiologists across 15 chest-radiograph tasks, incorrect AI advice reduced aggregate performance and harmed performance for half of the individual pathologies studied (Yu et al., 2024). A prospective randomized study of 220 physicians found that local feature-based explanations improved performance when advice was correct but also increased trust regardless of whether the advice was correct (Prinster et al., 2024).

Confidence calibration: Radiologists must learn to calibrate AI confidence scores to actual positive predictive value in their practice. A CAD system reporting “85% probability of malignancy” may have 40% PPV in a screening population, 70% PPV in a diagnostic population. Understanding this prevents over-reliance on AI confidence metrics.

The Role of Presentation Order:

Presentation order can change how AI information is used, but the chapter should not prescribe one sequence from an uncited mammography percentage. A randomized controlled vignette study involving 2,020 radiology assessments found that explanation format affected performance and that chain-of-thought explanations improved diagnostic accuracy by 12.2% in that controlled setting (Spitzer et al., 2026). The endpoint was diagnostic accuracy in study vignettes, not patient outcomes, trainee skill retention, or proof that one interface is universally preferable.

Fairness interventions can also fail to travel. Across medical-imaging datasets and tasks, locally optimized fairness did not reliably persist under distribution shift (Yang et al., 2024). Subgroup performance must be reassessed after deployment, not inferred from development data.

Override documentation: Institutions should track AI overrides, both false positive corrections (AI flagged, radiologist dismissed) and false negative corrections (AI missed, radiologist identified). This data informs:

  • Whether radiologists are appropriately skeptical of AI outputs
  • Patterns of AI failure in local practice
  • Training needs for radiologists over- or under-reliant on AI

Measuring Appropriate Skepticism:

Override rates can reveal how radiologists interact with AI, but no universal range defines appropriate skepticism. Institutions should stratify overrides by indication, confidence, patient subgroup, device version, scanner, and final adjudication. A low override rate can reflect either strong performance or over-reliance; a high rate can reflect poor performance, threshold mismatch, or an interface that presents too many low-value alerts.

Override metrics are signals for investigation, not stand-alone quality targets.

Cross-reference: For broader discussion of human-AI collaboration patterns and strategies to mitigate automation bias across specialties, see Integration into Clinical Workflow.

The Training Challenge Ahead:

Radiology faces a dilemma: AI tools improve efficiency and reduce errors when used appropriately, but may degrade the skills needed to recognize when AI is wrong. The solution requires deliberate curriculum design that builds independent expertise first, then layers AI proficiency on top of solid foundational skills.

Programs that integrate AI throughout training without preserving AI-free skill development risk producing radiologists who cannot function when AI fails, as it inevitably will in edge cases, technical failures, or deployment to new contexts where validation is incomplete.


Professional Society Guidelines on Radiology AI

Multi-Society Statement on AI in Radiology (2024)

In January 2024, five major radiology societies jointly published guidance on developing, purchasing, implementing, and monitoring radiology AI. The complete statement should be read with local policy because it addresses the lifecycle of the human-AI system, not a checklist for one product (Brady et al., 2024). The participating societies:

  • ACR (American College of Radiology)
  • CAR (Canadian Association of Radiologists)
  • ESR (European Society of Radiology)
  • RANZCR (Royal Australian and New Zealand College of Radiologists)
  • RSNA (Radiological Society of North America)

Core Principles:

  1. Patient well-being first: AI in radiology should increase patient well-being, minimize harm, respect human rights, and ensure benefits and harms are distributed equitably.

  2. Data ethics central: Ethical issues relating to acquisition, use, storage, and disposal of data are central to patient safety and appropriate AI use.

  3. Local evaluation required: Even FDA-authorized products require acceptance testing and performance monitoring in the intended environment.

  4. Workflow integration essential: AI systems must integrate with existing PACS and clinical workflows to achieve value.

  5. Continuous monitoring mandatory: Post-deployment surveillance for performance drift is not optional.

ACR AI Quality Programs

ARCH-AI Program (2024):

The ACR ARCH-AI program recognizes facilities that implement defined AI governance and quality-assurance practices. It is a recognition program, not imaging accreditation. The program establishes:

  • Expert consensus-based building blocks for AI infrastructure
  • Governance frameworks for AI implementation
  • Process standards for safe AI deployment
  • Quality metrics for ongoing monitoring

Assess-AI National Registry:

The ACR Assess-AI registry collects clinical AI results and contextual information to support:

  • Track accuracy over time
  • Identify performance shifts
  • Enable comparison across radiology departments
  • Provide baseline benchmarking data

ACR-SIIM Practice Parameter (2026):

The ACR-SIIM Practice Parameter for Imaging AI was approved on May 5, 2026, and is scheduled to take effect on October 1, 2026 (ACR approval and effective-date notice). It addresses selection, implementation, use, monitoring, privacy, and continuous quality improvement. The parameter is approved but is not yet effective as of August 14, 2026.

ESR Guidelines and Recommendations

EU AI Act Guidance (2025):

The ESR AI Working Group published recommendations for implementing the European AI Act in radiology (ESR, 2025), addressing:

  • Nine key articles particularly relevant to medical imaging
  • Post-market surveillance requirements
  • Data governance standards
  • Risk classification for imaging AI

ESR Essentials Series (2024-2025):

The ESR has published practice recommendations covering:

  • Health Technology Assessment for AI (European Radiology, December 2024): Framework for evaluating AI tools across their complete lifecycle
  • Common Performance Metrics (European Radiology, 2025): Guidance on selecting task-specific metrics and ensuring real-world performance
  • AI in Breast Imaging (European Radiology, 2025): Endorsed by ESR and EUSOBI for mammography AI deployment

RSNA Standards and Education

CLAIM Checklist (2024 Update):

The CLAIM 2024 update provides reporting standards for medical-imaging AI manuscripts, covering:

  • Classification algorithms
  • Image reconstruction
  • Text analysis
  • Workflow optimization

Multi-Society AI Education Syllabus:

AAPM, ACR, RSNA, and SIIM jointly developed a multisociety radiology AI syllabus with competency recommendations for four personas:

  1. Users of AI systems
  2. Purchasers of AI systems
  3. Clinical collaborators providing expertise during AI development
  4. Developers building AI systems

Pediatric Radiology AI Statement (2025)

A multisociety pediatric radiology statement, including ACR, ESPR, SPR, SLARP, AOSPR, and SPIN, addresses AI adoption across four pillars:

  1. Regulation and purchasing
  2. Implementation and integration
  3. Interpretation and post-market surveillance
  4. Education

Key recommendation: Pediatric-specific validation is essential, as AI trained on adult populations may perform differently in children.

Clinical Reporting Systems

AI does not replace the applicable reporting and management standard. Primary ACR resources include Lung-RADS, BI-RADS, and PI-RADS. Institutions should link the operative version from local protocols rather than restating a score from memory.


Future Directions:

Emerging Applications: - Multi-modal integration: Combining imaging with EHR data, genomics, pathology - 3D and 4D imaging: Better volumetric analysis, motion analysis - Synthetic data generation: Training AI without real patient data privacy concerns - Federated learning: Multi-institutional AI training without data sharing

Technical Advances: - Explainable AI: Better visualization of AI reasoning - Uncertainty quantification: AI signals when it’s unsure - Few-shot learning: AI that learns from small datasets (valuable for rare diseases) - Self-supervised learning: AI learns from unlabeled images

Regulatory Evolution: - Continuous learning systems: FDA frameworks for AI that improves post-deployment - Real-world evidence requirements: Post-market surveillance mandatory - International harmonization: Aligning FDA, CE Mark, other regulatory standards

The Clinical Bottom Line:

Key Takeaways for Radiologists and Referring Physicians
  1. Radiology AI is here and expanding: FDA’s periodically updated AI-enabled medical device list remains dominated by radiology devices, although FDA cautions that the list is not comprehensive (FDA AI-Enabled Medical Devices)

  2. AI augments, does not replace: Radiologists remain essential for complex reasoning, incidental findings, clinical integration, and oversight

  3. Performance varies dramatically by context: Local validation essential before deployment

  4. Evidence is maturing: RCTs and prospective studies increasingly available for major applications

  5. Workflow integration is challenging: Technical integration easier than changing radiologist practices

  6. Responsibility must be explicit: FDA authorization does not allocate civil liability; product labeling, workflow, institutional policy, professional duties, contracts, and jurisdiction all matter

  7. Alert fatigue is real: Balance sensitivity vs. specificity carefully

  8. Continuous monitoring essential: AI performance can drift with scanner changes, population shifts, software updates

  9. Subspecialty expertise still critical: AI handles routine cases, experts needed for complex cases

  10. Future is collaborative: Effective human-AI partnership, not replacement

Questions About Radiology AI

How accurate is AI in radiology?

There is no single accuracy figure for radiology AI. Performance depends on the task, threshold, prevalence, scanner, population, comparator, and workflow. Sensitivity and specificity from one study should not be transferred to another product, indication, or site.

Can AI replace radiologists?

Current evidence supports automation or assistance for selected imaging tasks, not replacement of the radiologist’s full interpretive, procedural, consultative, communication, and quality-assurance functions. Workforce effects should be treated as scenarios rather than measured inevitabilities.

What is CAD in radiology?

Computer-aided detection uses algorithms to mark or flag suspected findings for review. Its role differs from triage, which changes worklist priority, and from autonomous diagnosis. The authorized intended use and local workflow determine what the output can support.

How many radiology AI devices are FDA approved?

FDA updates its AI-enabled medical-device list periodically and states that the list is not comprehensive, so a fixed count becomes stale. Radiology is the largest specialty category, and the primary FDA record should be checked for each device’s pathway, intended use, date, and labeling.

Next Chapter: We’ll examine AI applications in Internal Medicine and Hospital Medicine, where data challenges differ substantially from imaging.


Hypothetical Decision Exercises

Fictional teaching cases

The three cases below are fictional decision exercises. Product behavior, patient outcomes, numerical performance changes, costs, policies, and legal arguments are illustrative unless a source is linked in the same sentence. They are not reported adverse events, predictions of a verdict, or representations of a named product’s performance.

Scenario 1: Missed Intracranial Hemorrhage After AI Triage

Assume a radiologist at a Level I trauma center uses an FDA-authorized generic AI triage system that flags suspected intracranial hemorrhage for expedited review.

Case: 68-year-old woman presents to ED after ground-level fall. Non-contrast head CT ordered. AI system does NOT flag the study as critical. You read the study 4 hours later during routine workflow (normal turnaround time for non-critical studies).

Your interpretation: Small (8mm) left frontoparietal subdural hematoma, no mass effect, no midline shift. You call the ED immediately.

ED physician: “We discharged her 2 hours ago. She seemed fine. GCS 15, no focal deficits. Why didn’t the AI flag this as critical?”

Patient outcome: Patient returns 6 hours later with worsening headache, confusion. Repeat CT shows subdural expansion to 15mm with early mass effect. She requires emergent craniotomy.

Answer 1: What went wrong?

AI false negative - The AI triage system failed to detect an 8mm subdural hematoma that met criteria for neurosurgical consultation.

Possible contributors to the miss: - Small hemorrhage size: Performance can vary with lesion size, but the training distribution is unknown in this fictional case - Attenuation and chronicity: Isodense or mixed-density subdural collections may be less conspicuous than acute hyperdense blood - Edge case: No sensitivity value should be assumed without the exact device labeling, indication, population, and operating point

Radiologist workflow failure: - You relied on AI triage system to identify critical studies - Without AI flag, study entered normal workflow queue (4-hour turnaround) - You did not have system in place to expedite all head trauma CTs regardless of AI flagging

Answer 2: Are you liable for malpractice?

No verdict can be predicted from this vignette. The relevant legal and safety questions include:

Standard of care: What is expected turnaround time for trauma head CTs? - What turnaround time is required by applicable institutional policy, accreditation requirements, clinical urgency, and local practice? - Did the queue design allow an absent AI flag to lower the priority of a study that should have remained urgent?

Reliance on AI: Did you inappropriately defer to AI for triage decisions? - AI is assistive tool, NOT substitute for radiologist judgment - You remain responsible for timely interpretation of all studies

Plaintiff argument: - Radiologist abdicated professional responsibility to AI - 4-hour delay in diagnosis led to delayed treatment, worse outcome, need for surgery - 8mm SDH with trauma mechanism should have been read urgently, AI flag or not

Defense argument: - The device was used within its labeling and did not replace the standard worklist - 8mm SDH without mass effect/midline shift is clinically stable in most cases - Patient was GCS 15 at discharge, clinical exam reassuring - Expansion of SDH is not predictable, could have occurred even with immediate diagnosis

Legal boundary: Breach, causation, and damages would require jurisdiction-specific analysis and case evidence. A chapter cannot forecast settlement or verdict.

Answer 3: How should this have been prevented?

System-level safeguards:

  1. AI triage is adjunct, not gatekeeper
    • AI flags accelerate critical studies
    • BUT: Absence of AI flag does NOT downgrade study priority
  2. Clear institutional policies
    • Define and audit an institution-specific urgent-turnaround policy independent of AI output
    • AI triage supplements, does not replace, standard workflows
  3. Continuous AI performance monitoring
    • Track false negative rate for AI triage
    • Measure false negatives using an appropriate reference standard and denominator
    • Review discordant cases at a frequency proportionate to risk and volume
  4. Radiologist training
    • AI is assistive, not definitive
    • Clinical urgency determined by clinical context + exam, not AI alone

Documentation: Follow institutional policy and document AI output when it is clinically material, required by labeling, or relevant to a quality investigation. Routine templated documentation of every concordant or discordant output can create noise without improving safety.

Lesson: AI triage improves workflow for true positives but CANNOT be trusted to catch 100% of critical findings. Standard turnaround time policies for trauma/ED studies must remain in place regardless of AI flagging.

Scenario 2: CAD False Positive Leading to Unnecessary Biopsy

Assume a breast imaging radiologist uses an AI-based computer-aided detection system for screening mammography.

Case: 52-year-old woman, average-risk screening mammogram. You interpret study as BI-RADS 1 (negative). CAD system flags 6mm focal asymmetry in right upper outer quadrant as “suspicious mass, recommend additional imaging.”

Your initial assessment: This area appears benign: normal glandular tissue, no associated calcifications, present on prior mammogram 1 year ago (stable). You are confident this is false positive.

Decision point: Do you: - A) Override CAD, final report BI-RADS 1 (negative) - B) Follow CAD suggestion, upgrade to BI-RADS 0 (needs additional imaging)

You choose: Option B (upgrade to BI-RADS 0) - “Out of abundance of caution and because CAD flagged it.”

Outcome: - Patient recalled for diagnostic mammogram + ultrasound - Diagnostic mammo: Focal asymmetry unchanged from prior (stable) - Ultrasound: No correlate - You recommend 6-month follow-up (BI-RADS 3) - Patient anxious, insists on biopsy “to be sure” - Stereotactic biopsy performed: Benign fibroglandular tissue - Patient develops hematoma requiring drainage, files complaint about “unnecessary procedure”

Answer 1: What was the error?

Over-reliance on CAD - You upgraded a finding you clinically assessed as benign ONLY because CAD flagged it, despite: - Your expert interpretation: benign - Stability on prior imaging (1 year) - No suspicious features (no calcifications, no mass characteristics)

CAD false positive: False-positive burden depends on the exact product, threshold, population, and workflow. The system’s access to priors, clinical context, and risk information must be verified rather than assumed.

Failure to apply clinical judgment - You allowed CAD to override your expert assessment.

Answer 2: Was this the right decision medicolegally?

The decision cannot be declared legally wrong from the vignette alone. The quality issues include:

Unsupported deference: “Out of abundance of caution because CAD flagged it” does not explain why the imaging finding justified escalation - CAD is decision support tool, not clinical decision-maker - Your role: Expert interpretation integrating CAD output with clinical judgment, priors, patient factors

Overdiagnosis harm - Unnecessary recall, anxiety, biopsy, hematoma - The fictional patient experienced harms after a cascade that began with an inadequately justified recall - Cascade of interventions triggered by inappropriate deference to AI

Liability risk paradox: - Following an AI output does not create a legal safe harbor - Overcalling and undercalling can both create harm; legal consequences remain fact-specific

Answer 3: What is the appropriate use of CAD?

CAD as “second reader”: 1. You interpret first - Form your own impression BEFORE looking at CAD marks 2. Review CAD marks - Consider areas CAD flagged 3. Apply clinical judgment - Decide if CAD finding warrants further action

When to follow CAD: - CAD flags area you initially overlooked → Re-review carefully, may represent true finding - CAD flags area you were uncertain about → May support upgrading to BI-RADS 0

When to override CAD: - CAD flags area you assessed as clearly benign AND stable on priors → Override, document rationale - CAD flags known benign finding (e.g., lymph node, surgical scar) → Override

Documentation: - If you override CAD, document: “CAD system flagged [location]. Reviewed; consistent with benign [finding]. Stable on [date] prior. Assessed BI-RADS 1.” - If you follow CAD, document: “CAD system identified [finding] not initially appreciated. Recommend additional imaging for further characterization.”

Lesson: CAD is adjunct to expert interpretation, not substitute. Radiologist clinical judgment integrating CAD output with priors, risk factors, and clinical context determines final assessment. Blindly following CAD recommendations leads to over-diagnosis, patient harm, and liability.

Scenario 3: AI Algorithm Degradation After Scanner Upgrade

Assume a radiology director discovers that the institution replaced conventional CT scanners with a materially different photon-counting platform without completing the AI change-control process.

AI tools affected: The fictional institution uses generic vendor-neutral algorithms for: - Pulmonary embolism detection - Lung nodule detection - Intracranial hemorrhage detection

All AI tools validated on prior scanner model (Somatom). No re-validation performed after NAEOTOM upgrade.

Week 1 post-upgrade: PE detection AI generates 15 false positive alerts (vs. typical 2-3/week). Radiologists dismiss as “AI acting up.”

Week 3 post-upgrade: Quality assurance review reveals: - PE AI sensitivity drop: 95% → 78% (17 percentage point degradation) - PE AI false positive rate increase: 5% → 22% - Root cause: New scanner’s photon-counting technology produces different image noise characteristics, contrast-to-noise ratios than conventional CT - AI algorithm trained on conventional CT data does not generalize to photon-counting CT

Clinical impact: - 8 PE cases missed by AI over 3-week period (all eventually detected by radiologist) - 1 subsegmental PE missed by both AI and radiologist (patient decompensated 24 hours later, escalated to ICU)

Answer 1: What was the failure?

Scanner upgrade without AI re-validation - Critical error - AI algorithms are scanner-specific, trained on specific image characteristics - New scanner technology (photon-counting CT) produces different imaging data - No validation performed before clinical deployment on new scanner

Lack of continuous performance monitoring - AI performance drift not detected for 3 weeks - No system to track AI sensitivity/false positive rates in real-time - Radiologists noticed “AI acting up” but did not escalate to formal investigation

Communication breakdown - IT upgraded scanners without radiology notification - AI deployment requires cross-departmental coordination (IT, radiology, vendors)

Answer 2: Who is liable for the missed PE?

Responsibility cannot be allocated from the vignette alone. Potential systems issues include:

Plaintiff argument: - Hospital deployed AI on a new scanner without determining whether acceptance testing or revalidation was required by labeling, quality policy, or the change-control plan - Hospital failed to monitor AI performance - Scanner upgrade performed without radiology notification, impact assessment - Radiologist relied on AI for sensitivity (standard practice with validated AI)

Radiology department argument: - IT department made scanner upgrade unilaterally - Radiology not informed of upgrade - No opportunity to re-validate AI or pause AI deployment during transition

Radiologist individual argument: - Subsegmental PE is difficult to detect (small, distal vessels) - Relied on AI as validated tool (95% sensitivity on prior scanner) - Individual radiologist cannot be expected to detect systemic AI failure

Legal boundary: Institutional and individual duties would depend on notice, policy, product labeling, contracts, standard of care, and causation. The facts support a quality and governance investigation, not a predicted liability allocation.

Answer 3: How should scanner upgrades be managed?

Pre-upgrade protocol:

  1. Radiology notification required - Use a locally defined advance-notice period sufficient for impact assessment and testing
  2. AI impact assessment - Identify all AI tools potentially affected
  3. Vendor consultation - Contact AI vendors to determine if re-validation needed
  4. Re-validation plan:
    • Pause AI deployment during scanner transition
    • Select a validation sample and reference standard based on intended use, prevalence, uncertainty, and the consequence of error
    • Compare AI performance on new scanner vs. prior scanner
    • Predefine risk-based acceptance and stop rules rather than applying a universal 5% threshold

Post-upgrade protocol:

  1. Continuous performance monitoring:
    • Track clinically relevant performance at a frequency proportionate to risk, volume, and the speed with which harm could accumulate
    • Automated dashboards with alert thresholds
  2. Quality assurance:
    • Review random sample of AI-negative studies (detect false negatives)
    • Review AI-positive studies (measure false positive rate)
    • Present AI performance metrics at radiology QA conferences monthly

Governance: - Radiology AI Committee with authority to approve/pause AI tools - IT-Radiology protocol: No scanner/PACS/software changes without radiology sign-off if AI tools deployed - Vendor agreements: Require vendors to notify customers when imaging equipment changes may affect AI performance

Documentation: - Maintain AI validation log documenting scanner model, software version, validation dates - When scanner upgraded, document: “AI tool [name] paused pending re-validation on new scanner”

Lesson: AI algorithms are tightly coupled to imaging equipment characteristics. Scanner upgrades, even within same vendor, can degrade AI performance. Continuous monitoring and re-validation after equipment changes are essential.