Diagnostic Imaging and Radiology

Radiology has the largest regulatory footprint for medical AI, but regulatory volume is not evidence of clinical benefit. A cross-sectional review of FDA authorizations supports that characterization (Sivakumar et al., 2025). The strongest current studies show a mixed pattern: AI can improve detection, reading efficiency, and selected workflows, while other randomized and prospective studies show no time benefit, higher recall, or clinically important misses. The clinically relevant question is therefore not whether an algorithm is accurate, but whether a defined human-AI workflow improves care in the population where it will be used.

Learning Objectives

After reading this chapter, you will be able to:

  • Distinguish diagnostic accuracy from workflow and patient outcomes
  • Interpret randomized, prospective, silent, retrospective, reader, regulatory, and vendor evidence
  • Evaluate imaging AI across radiography, mammography, CT, MRI, ultrasound, and report generation
  • Recognize automation bias, inappropriate reliance, distribution shift, and de-skilling risk
  • Apply current ACR, RSNA, multisociety, ESR, EUSOBI, and pediatric radiology guidance
  • Design local acceptance testing and postdeployment monitoring for imaging AI
  • Explain what FDA clearance establishes and what it does not establish

What the evidence supports

  • Randomized trials show that AI can increase actionable lung nodule detection on chest radiographs and improve mammography screening sensitivity in defined workflows.
  • Prospective implementation studies show that imaging AI can increase cancer detection and reduce documentation time, but results depend on workflow, thresholds, population, and endpoint.
  • Negative and mixed studies matter. Lung cancer worklist prioritization did not shorten time to CT or diagnosis in a large randomized trial, and an AI mammography triage strategy failed its recall-rate noninferiority endpoint.

What remains unproven

  • FDA clearance does not establish local clinical utility, improved patient outcomes, economic value, or suitability for autonomous use.
  • Most cleared radiology AI products have not been evaluated prospectively with clinicians in the loop.
  • No evidence supports replacing the radiologist’s full interpretive, consultative, procedural, and communication functions with current AI systems.

Implementation rule

Define the intended use, validate locally, train users on failure modes, preserve a safe non-AI workflow, monitor performance by site and subgroup, and establish stop rules before deployment.

Introduction

Imaging AI now spans detection, worklist triage, segmentation, quantitative measurement, reconstruction, report drafting, protocol selection, and quality assurance. These functions have different evidentiary requirements. A model may classify images accurately yet fail to improve reporting time, treatment, outcomes, or cost when inserted into a real clinical system.

The evidence base should be read in descending order of clinical proximity:

  1. Randomized clinical utility studies: Compare complete care or screening workflows and measure downstream outcomes.
  2. Prospective implementation studies: Evaluate a deployed workflow without randomization.
  3. Prospective silent studies: Run the model on live cases without exposing outputs to clinicians.
  4. Retrospective external testing: Test stored data from institutions outside model development.
  5. Reader studies: Measure performance under controlled interpretation conditions.
  6. Regulatory evidence: Establishes that the device met requirements for its pathway and intended use.
  7. Vendor or benchmark evidence: Generates hypotheses but does not establish clinical utility.

AUC, sensitivity, and specificity do not establish that a deployed workflow benefits patients. The endpoint must match the claim. A triage product should be evaluated for time to interpretation or action, a screening product for detection, recall, interval disease, and overdiagnosis, and a report generator for final-report accuracy, editing burden, and consequential errors.

Where AI Operates in the Imaging Workflow

Detection and triage

Computer-aided detection marks suspected findings for review. Triage software changes worklist priority or notifies clinicians about suspected time-sensitive findings. These are not interchangeable: detection may change what is seen, while triage may change when a study is read.

Quantification and reconstruction

AI can segment anatomy, measure lesions, estimate volumes, quantify calcium or emphysema, reduce noise, and reconstruct images from accelerated or lower-dose acquisitions. Technical image quality and diagnostic equivalence should be assessed separately. A visually plausible reconstruction can still alter or suppress pathology.

Classification and decision support

Classification systems estimate whether an image or region is normal, abnormal, benign, malignant, or consistent with a specific diagnosis. The output must be interpreted at the local disease prevalence. Positive and negative predictive values can differ sharply from the development study even when sensitivity and specificity remain stable.

Report generation and communication

Generative systems can draft findings, compare prior studies, structure reports, identify possible communication failures, or translate reports for patients. The final report remains a clinical document. Drafting efficiency must be balanced against omission, insertion, laterality, comparison, and recommendation errors.

In a prospective cohort across a 12-hospital academic system, documentation time was 15.5% lower for 11,980 model-assisted radiograph interpretations than for 11,980 matched baseline interpretations. A blinded peer-review sample of 800 reports found no detected difference in clinical accuracy or textual quality (Huang et al., 2025). This was observational evidence from one system, not a randomized demonstration of better patient outcomes.

Clinical Evidence by Modality

Chest radiography

Chest radiography illustrates why endpoint selection matters.

  • A pragmatic, open-label, single-center randomized trial of 10,476 health-screening participants found that AI assistance increased detection of actionable lung nodules from 0.25% to 0.59% without a statistically significant increase in positive reports or false referrals (Nam et al., 2023). The trial measured detection, not mortality or long-term outcomes.
  • A prospective multicenter silent study across five National Health Service sites analyzed 63,083 adult chest radiographs. The model classified 20% as normal, achieved 97% sensitivity and 35% specificity for abnormal studies, and had 31 clinically significant misses after discrepant-case review (Storey et al., 2026). A silent trial cannot establish the safety of autonomous exclusion from radiologist review.
  • The LungIMPACT randomized trial analyzed 93,326 chest radiographs and found that AI worklist prioritization did not shorten median time to CT or lung cancer diagnosis (Woznitza et al., 2026). AI prioritization did not improve the pathway merely because the worklist changed.
  • In a large retrospective multicenter study, a pneumothorax model showed high diagnostic performance; a separate real-world study found faster oxygen-therapy initiation when CAD was paired with electronic clinical alerts (Hillis et al., 2022; Oh et al., 2025). The treatment finding applies to that specific alert pathway.

COVID-19 chest radiograph models also exposed shortcut learning. Models that appeared accurate often relied on site-specific markers and acquisition artifacts rather than disease pathology (DeGrave et al., 2021). This failure mode remains relevant whenever a model is moved to a new hospital, scanner fleet, or protocol.

Mammography

Mammography has the strongest program-level imaging AI evidence, but results differ by implementation strategy.

  • The MASAI randomized trial assigned 105,934 women to AI-supported screening or standard double reading. Full follow-up found higher sensitivity with AI-supported screening (80.5% vs 73.8%) at the same specificity (98.5%); interval cancer and aggressive cancer differences favored AI but were not statistically significant (Lång et al., 2023; Gommers et al., 2026). MASAI supports one defined screening workflow, not mammography AI as a class.
  • PRAIM was a prospective observational implementation study at 12 German screening sites, not a randomized trial. Among 463,094 women, radiologists voluntarily using AI had a higher cancer detection rate and a noninferior recall rate compared with standard double reading (Eisemann et al., 2025). Voluntary use and nonrandom allocation limit causal interpretation.
  • A multicenter United Kingdom study evaluated 115,973 mammograms retrospectively and then deployed the system prospectively in a noninterventional feasibility phase involving 9,266 cases at 12 sites. It demonstrated technical feasibility and strong retrospective diagnostic performance, but it did not test an AI-directed screening pathway against usual care (Kelly et al., 2026).
  • The AITIC prospective paired noninferiority study included 31,301 women. Its partially autonomous strategy reduced radiologist workload by 63.6% and increased cancer detection, but recall was 14.8% higher and failed the prespecified noninferiority endpoint (Elías-Cabot et al., 2026; NCT04949776). The recall result prevents describing AITIC as an uncomplicated success.

First-generation mammography CAD remains an essential counterexample. A large observational study found increased recall without improved cancer detection (Lehman et al., 2015). New deep-learning systems require evaluation on their own evidence, but novelty does not erase the need to measure false-positive downstream effects.

The EUSOBI practice recommendations for AI in breast imaging regard current systems as aids to the reporting radiologist and emphasize postmarket surveillance and the absence of demonstrated long-term outcome benefit.

Head CT and stroke imaging

Intracranial hemorrhage and large-vessel occlusion systems can detect or prioritize suspected emergencies. A prospective stepped-wedge cluster-randomized trial at four comprehensive stroke centers included 243 patients who underwent endovascular thrombectomy. AI activation reduced adjusted door-to-groin time by 11.2 minutes and CT-to-EVT time by 9.8 minutes, but did not significantly improve 90-day functional independence (odds ratio 1.3, 95% CI 0.42–4.0) (Martinez-Gutierrez et al., 2023). A retrospective implementation study found that worklist reprioritization shortened time to diagnosis for outpatient head CT with intracranial hemorrhage (Arbabshirani et al., 2018). An observational stroke-network study associated automated notification with shorter off-hours door-to-groin time (Figurelle et al., 2023). Randomized and observational evidence supports shorter selected workflow intervals, but patient-outcome benefit remains unproven for isolated imaging-alert systems.

A triage flag must never remove an unflagged examination from the standard queue. False negatives remain possible, and the clinical impact of reprioritization depends on staffing, scanner-to-PACS latency, notification routing, stroke-team response, and existing process times.

Chest, abdominal, and pelvic CT

Body CT AI now includes pulmonary embolism, aortic disease, bowel obstruction, inflammatory disease, trauma, and opportunistic measurements. The intended use is often triage rather than diagnosis.

One 2025 Nature Medicine study developed and evaluated an AI warning system for acute aortic syndrome on noncontrast CT across a staged program that included multicenter retrospective testing, reader evaluation, large-scale real-world testing, and prospective workflow assessment in China (Hu et al., 2025). It is a substantial prospective translational study, but the diagnostic pathway and initial use of noncontrast CT reflect a specific health-system context.

The FDA cleared BriefCase-Triage: CARE Multi-triage CT Body through the 510(k) pathway on January 7, 2026 (K252970, Class II). Its labeling covers notification for 11 specified findings on adult chest, abdominal, or pelvic CT. The device operates in parallel with standard interpretation, does not remove or deprioritize cases, and its preview images are not intended for diagnostic use (FDA K252970 record and summary). The FDA labeling is narrower than claims of autonomous multi-disease diagnosis.

MRI

MRI AI supports reconstruction, acquisition acceleration, segmentation, lesion classification, and quantitative analysis. In a 10,207-examination confirmatory study, an AI system for prostate MRI showed higher standalone discrimination than participating radiologists for clinically significant cancer, but the intended workflow still required radiologist interpretation (Saha et al., 2024).

Commercial reconstruction evidence remains uneven. A systematic review of 14 products found that 29% had no peer-reviewed validation studies and that prospective evidence of clinical impact was uncommon (Fransen et al., 2025). Faster acquisition is not sufficient if pathology preservation has not been tested.

Ultrasound, interventional radiology, and nuclear medicine

Ultrasound AI is used for view acquisition, measurements, lesion characterization, and point-of-care guidance. Interventional applications include segmentation, navigation support, and procedure planning. Nuclear medicine applications include PET and SPECT reconstruction, attenuation correction, lesion segmentation, dosimetry, and theranostic response assessment.

Evidence is heterogeneous and usually task-specific. Products should be evaluated against the acquisition device, operator experience, tracer or contrast protocol, intended population, and downstream decision. No broad inference should be made from performance in one organ, modality, or device platform to another.

Generalist Models and Generative AI

Foundation models seek to support multiple imaging tasks through one pretrained architecture. MedVersa, for example, was trained on 29 million instances from 91 public datasets and evaluated across classification, segmentation, visual question answering, and report generation (Zhou et al., 2026). Such studies establish breadth under experimental conditions, not permission for unrestricted clinical use.

Generalist models create additional risks:

  • The valid task and population may be unclear to the user.
  • Performance can vary substantially across modalities and conditions.
  • A single update can affect multiple downstream functions.
  • Generated text can introduce findings that are absent or omit clinically important abnormalities.
  • Benchmark breadth can obscure the absence of prospective clinical evaluation.

A generalist architecture does not create a general clinical indication. Each deployed function still requires a defined intended use, evidence, version control, local testing, and monitoring.

Human Factors and Radiologist Performance

AI performance and human-AI team performance are separate quantities. A model can improve average reader performance while harming some readers or some pathologies.

In a study of 140 radiologists across 15 chest radiograph tasks, incorrect AI advice reduced aggregate performance and harmed performance for half of the individual pathologies studied. Years of experience, thoracic subspecialization, and familiarity with AI did not reliably predict who would benefit (Yu et al., 2024). Incorrect AI can lower radiologist performance, including among experienced users.

A prospective randomized study of 220 physicians found that local feature-based explanations improved performance when advice was correct, but also increased simple trust regardless of whether the advice was correct (Prinster et al., 2024). An explanation can increase reliance without making the underlying advice correct.

Fairness interventions also may not travel. Across medical imaging datasets and tasks, locally optimized fairness did not reliably persist under distribution shift (Yang et al., 2024). Subgroup performance must be reassessed after deployment, not inferred from development data.

Training and de-skilling

Direct evidence that radiology trainees become de-skilled from routine AI exposure remains limited. The stronger current evidence is broader: incorrect AI can degrade radiologist performance, explanations can increase inappropriate reliance, and a multicenter observational colonoscopy study found that adenoma detection during non-AI procedures declined after clinicians were exposed to AI-assisted colonoscopy (Budzyń et al., 2025). The endoscopy finding should not be represented as direct evidence about radiology residents.

Programs should teach independent image review, AI failure modes, calibration, subgroup performance, and safe override. The AAPM, ACR, RSNA, and SIIM multisociety syllabus provides a current curricular foundation. Periodic assessment without visible AI output is a reasonable safeguard, but specific rotation structures and thresholds have not been validated.

FDA Clearance and the Evidence Gap

The FDA’s AI-enabled medical device list is an important discovery tool, but FDA states that it is not comprehensive. The device decision, summary, intended use, user population, contraindications, and pathway should be checked in the primary FDA record.

FDA clearance establishes a regulatory decision, not local clinical utility. Most imaging AI devices enter through 510(k), which is based on substantial equivalence to a predicate rather than a requirement to prove better patient outcomes.

A cross-sectional review identified 950 FDA-authorized AI or machine-learning devices through June 2024, including 723 in radiology. Among 717 radiology devices with available submission documentation, 5% had prospective testing, 8% included a human operator, and 29% included clinical testing (Sivakumar et al., 2025). These percentages describe the reviewed documentation, not proof that all other devices were clinically ineffective.

A separate analysis linked 60 of 950 devices to 182 recall events. Of the recall events, 43.4% occurred within 12 months of clearance, and absence of reported clinical validation was associated with greater odds of recall (adjusted OR 2.8, 95% CI 1.6–4.7) (Lee et al., 2025). A newer cohort study found that 43 of 903 devices were recalled (4.8%), with a median 458 days from authorization to recall. Missing information about supporting clinical studies had an estimated hazard ratio of 1.39, but its 95% credible interval was wide and included no association (0.84–3.52) (Ren et al., 2026). These public-data associations do not prove that missing validation caused a recall.

Economic and Workforce Evidence

A rapid systematic scoping review screened 8,013 records and included 140 studies of diagnostic radiology AI. It found a large technical literature but sparse evidence on implementation, staff and patient experience, quantitative workflow effects, and cost (Lawrence et al., 2025).

A systematic review of economic evidence screened 1,879 publications and included 21. Economic value varied with task, disease prevalence, staffing, payment model, and licensing structure; many evaluations were modeled rather than based on mature deployed systems (Molwitz et al., 2026). A time saving is not a cost saving unless the surrounding workflow can use it.

Current evidence does not support a precise forecast of radiologist replacement. Radiologists perform protocol selection, interpretation across modalities and priors, procedures, consultation, communication, quality oversight, and management of uncertainty. AI is more likely to redistribute tasks than eliminate the specialty. Workforce projections should be labeled as scenarios, not demonstrated effects.

Implementation Standard

The 2024 multisociety statement from ACR, CAR, ESR, RANZCR, and RSNA covers development, purchasing, implementation, monitoring, and long-term safety (Brady et al., 2024). The ACR and SIIM practice parameter was approved on May 5, 2026, and is scheduled to take effect on October 1, 2026. It formalizes expectations for governance, inventory, acceptance testing, privacy, monitoring, and continuous quality improvement (ACR, 2026; practice parameter).

Minimum Local Deployment Package

Before clinical display of AI output, the institution should have:

  1. A named clinical owner and multidisciplinary governance group
  2. The FDA or other regulatory record and the exact intended use
  3. Local acceptance testing using representative scanners, protocols, populations, and prevalence
  4. A documented human-AI workflow, escalation pathway, downtime plan, and stop rules
  5. User training on known false positives, false negatives, and prohibited uses
  6. Baseline metrics for diagnostic, workflow, subgroup, and safety outcomes
  7. Version control and a plan to revalidate after material software, scanner, or protocol changes
  8. Postdeployment review of discordant cases, drift, complaints, and safety events

Local acceptance testing is required even when a device is cleared and externally tested. The local question is whether the exact version performs acceptably in the intended workflow, not whether it worked somewhere else.

What to measure

Measure only outcomes relevant to the intended use:

Function Minimum measures
Detection Sensitivity, specificity, PPV, NPV, false-negative review, subgroup performance
Triage Time to read, report, notification, action, and treatment; queue displacement effects
Screening Detection, recall, interval disease, biopsy yield, stage, workload, overdiagnosis indicators
Quantification Agreement, repeatability, failure rate, edit burden, downstream decision concordance
Reconstruction Diagnostic equivalence, pathology preservation, artifacts, repeat scans, acquisition time
Report drafting Final-report accuracy, omissions, insertions, laterality, editing time, addenda, communication failures

The ACR Assess-AI registry supports real-world monitoring and benchmarking. The ACR ARCH-AI program recognizes facilities that implement defined AI governance and quality-assurance practices. ARCH-AI is a recognition and quality program, not imaging accreditation.

The European Society of Medical Imaging Informatics recommends local validation using clinically relevant metrics and institutional prevalence (Klontzas et al., 2025). Postdeployment monitoring is part of the intervention, not an optional audit.

Professional Society and Reporting Guidance

Society documents should be read in full because their scope differs.

Practice, governance, and regulation

Quality programs and research reporting

Clinical reporting systems

AI does not replace the applicable reporting and management standard. Relevant primary ACR resources include Lung-RADS, BI-RADS, and PI-RADS.

Education and population-specific guidance

Adult performance cannot be assumed to transfer to pediatric imaging. Differences in anatomy, disease prevalence, acquisition, radiation considerations, and available training data require pediatric-specific evaluation.

Liability and Documentation

Liability remains jurisdiction-dependent and fact-specific. Mock-juror studies describe perceptions in hypothetical cases, not adjudicated legal rules. The operational response is to define responsibility, document material use or override when required by institutional policy, preserve version and output logs, and investigate discordant cases.

The full evidence and legal analysis belongs in Liability and Malpractice. This chapter’s narrower point is practical: clearance, use, nonuse, and override can each become relevant after an adverse event, but none creates a universal rule that the radiologist must accept the algorithm.

Hypothetical Decision Vignettes

Negative triage output

An urgent examination receives no AI flag. The absence of a flag must not lower its institutional priority or replace the standard interpretation pathway. The event should be captured if subsequent review identifies a clinically significant miss.

Positive mammography prompt

AI marks a subtle region after the radiologist has formed an independent assessment. The radiologist reviews the prompt against all views, priors, and the applicable BI-RADS standard. A positive prompt alone does not determine the assessment category or management recommendation.

Scanner or software change

A scanner upgrade, reconstruction change, model update, or new patient population alters the deployed environment. The institution assesses whether the change falls within the tested configuration, conducts targeted revalidation, and intensifies monitoring before returning to routine oversight.

Key Takeaways

  1. Radiology leads medical AI regulation, but clearance volume is not clinical utility.
  2. Evidence should be matched to the claim: accuracy, workflow, program, patient, and economic outcomes are different endpoints.
  3. Positive randomized evidence exists for defined chest radiograph and mammography workflows.
  4. Negative and mixed evidence must remain visible, including LungIMPACT’s null time result and AITIC’s recall result.
  5. Incorrect AI and persuasive explanations can worsen human performance.
  6. Local acceptance testing, subgroup analysis, monitoring, version control, and stop rules are required for responsible deployment.
  7. Society guidance should be linked and used directly, including the ACR-SIIM practice parameter, ARCH-AI, Assess-AI, CLAIM, and relevant RADS standards.
  8. Current AI augments selected radiology tasks. It does not replace the specialty’s complete clinical function.