Appendix C — Quick Reference: All Chapter Summaries (TL;DRs)

Using This Appendix

This appendix compiles selected chapter TL;DRs (Too Long; Did not Read summaries) for rapid reference. It is designed for:

  • Quick review before implementing AI tools
  • Refreshing key concepts
  • Finding specific information across chapters
  • Sharing with colleagues who need executive summaries

Each TL;DR includes: - The clinical context - Key evidence and applications - What works vs. what does not - Critical takeaways

For full details, citations, limitations, and implementation guidance, use each linked chapter as the controlling source.


Part I: Foundations

Chapter 1: History of AI in Medicine

Key Lesson: Controlled technical performance does not establish clinical adoption or patient benefit.

Major Failures: - MYCIN (1970s): Experts rated 65% of its proposed regimens acceptable in a ten-case blinded evaluation, but it never entered routine clinical care (Yu et al., 1979) - IBM Watson for Oncology: Published evaluations mainly measured concordance, not patient outcomes, and results varied across cancers and settings - Google Flu Trends: Ran high for 100 of 108 weeks from August 2011 through September 2013 and exposed the risks of opaque, unstable data-generating processes (Lazer et al., 2014)

What Works: Narrow, well-defined tasks with an explicit population, reference standard, workflow, and evidence boundary. IDx-DR is one example, authorized through De Novo classification for a specified diabetic-retinopathy screening use after a prospective single-arm pivotal study (Abràmoff et al., 2018).


Chapter 2: AI Fundamentals for Clinicians

Core Concept: AI learns patterns from data, does not follow explicit rules

Critical Metrics: - PPV depends on prevalence and the evaluated population; no single metric is sufficient for every clinical use. - AUC summarizes discrimination across thresholds but does not establish calibration, utility, safety, or patient benefit - Sensitivity vs. Specificity trade-offs

Key Limitations: - Black-box problem (cannot explain reasoning) - Distribution shift (works at Hospital A, fails at Hospital B) - Bias amplification (training data biases → algorithmic biases)


Chapter 3: Clinical Data Challenge

Reality: Clinical data is messy - missingness, heterogeneity, temporal complexity, bias

Critical Issues: - Missing data is NOT random (sicker patients have more data) - EHR data quality variable (copy-paste errors, billing optimization) - External validation addresses transportability within tested conditions; prospective workflow and comparative outcome studies answer different questions

Demand: Evidence that matches the intended population, setting, task, comparator, reference standard, workflow, and clinical consequence.


Part II: Clinical Specialties

Chapter 4: Radiology

Maturity: Radiology has the largest representation on FDA’s periodically updated public list of AI-enabled medical devices, but authorization volume is not a measure of clinical benefit.

Strong Evidence: - Diabetic retinopathy screening (IDx-DR): De Novo-authorized for a narrow autonomous use after a prospective single-arm study, not an RCT - ICH triage: selected products have diagnostic or workflow evidence; verify the exact product and endpoint - LVO stroke: a Viz.ai cluster-randomized trial shortened selected treatment intervals without significantly improving 90-day functional independence (Martinez-Gutierrez et al., 2023) - Mammography: the MASAI randomized trial provides screening-specific evidence; results should not be transferred to every mammography product (Lång et al., 2023)

Workforce boundary: Current evidence supports task redistribution and augmentation scenarios, not a settled prediction of wholesale radiologist replacement.


Surgery, Anesthesiology, and Perioperative Care

Current State: Surgical AI is strongest before and after the operation: risk prediction, planning, video review, quality measurement, and postoperative monitoring. Intraoperative AI remains high risk because surgical decisions are immediate and often irreversible.

Robotic Surgery: Clinical robotic surgery is teleoperated. AI is emerging mainly in surgical video analytics, privacy-preserving multicenter model training, skill assessment, case review, OR scheduling, and research-stage supervised autonomy (Saldanha et al., 2026).

Evaluation Standard: - Separate platform safety, surgeon learning curve, procedure outcomes, AI perception, human factors, and physical autonomy - Verify FDA indication and document number before procurement or patient-facing claims - Demand procedure-specific evidence, not generic robotic-surgery claims - Run silent-mode validation before displaying intraoperative AI outputs

Bottom Line: Robotic surgery evidence is procedure-specific. Preoperative risk prediction has stronger evidence than real-time intraoperative AI. No surgical AI should guide irreversible intraoperative action without independent surgeon verification.


Primary Care, Family Medicine, and Preventive Medicine

Best Applications: - Diabetic retinopathy screening (IDx-DR): narrow authorized use with prospective pivotal evidence - Ambient documentation: emerging randomized and observational evidence, with product, workflow, correction burden, and endpoint boundaries - Chronic disease monitoring (BP, diabetes)

Weak Evidence: - General diagnostic AI (too complex for current systems) - General symptom checkers, where performance varies by task, benchmark, population, and evaluation method

Critical: Workflow integration essential (no time for separate systems in 15-min visits)


Emergency Medicine

Better-supported applications: - LVO stroke detection: Viz.ai has randomized and observational evidence for shorter selected intervals; RapidAI and Brainomix require separate product-specific evidence - ICH and PE triage: evidence must stay tied to the authorized finding, modality, software version, site, and measured workflow endpoint

Controversial: - One external validation of Epic Sepsis Model version 1 reported 33% sensitivity and 12% PPV at the evaluated threshold; it did not test causal patient benefit. (Wong et al., 2021) - Deterioration prediction - Variable results, implementation-dependent

Challenge: Alert fatigue in high-volume EDs


Part III: Implementation

Chapter 23: Evaluating AI Clinical Decision Support Systems

Evaluation logic: 1. Match the claim to the study design and endpoint 2. Verify the exact model, version, population, comparator, and reference standard 3. Distinguish internal validity, transportability, workflow effect, clinical utility, and patient outcomes 4. Treat regulatory authorization, diagnostic accuracy, usability, and outcome benefit as separate questions 5. Plan independent local monitoring before deployment

20 Essential Questions: See full chapter for complete vendor evaluation checklist

Red Flags: - No peer-reviewed publications - No external validation - Vendor refuses to share performance data - Claims 99%+ accuracy


Part IV: Practical Tools

Chapter 29: AI Tools Every Physician Should Know

Evidence-aware examples: - IDx-DR: De Novo-authorized autonomous diabetic-retinopathy screening for a defined population, supported by a prospective single-arm pivotal study - Viz.ai: randomized and observational stroke-workflow evidence, without established functional-outcome benefit from the alert alone - Ambient documentation systems: evaluate documentation quality, correction burden, time, clinician experience, privacy, and downstream errors separately - Aidoc, Paige, and Circle CVI product families: verify the exact authorized device, version, intended use, and product-specific evidence before drawing conclusions

Avoid: - Unvalidated symptom checkers - Tools without peer-reviewed publications - “AI diagnoses everything” systems


Chapter 30: Large Language Models in Clinical Practice

What LLMs Can Do: - Literature synthesis - Documentation drafts (WITH REVIEW) - Patient education materials (WITH VERIFICATION) - Differential diagnosis brainstorming

Critical Limitation: Hallucinations (confident but false information)

Safety boundaries: - Do not enter protected or sensitive patient data into an unapproved service. - Verify clinically consequential output against authoritative sources and the patient record - Do not rely on an unvalidated model as the sole basis for urgent or irreversible decisions - Do not treat a model response as a substitute for indicated specialist consultation

Rule: Oversight and accountability must be defined for the exact task. Some systems perform narrow autonomous functions within authorized labeling, while clinicians and institutions retain separate duties.


Clinical Trials and AI

Evidence boundary: Trial retrieval, criterion review, consent, enrollment, retention, and patient outcomes are different endpoints. TrialGPT demonstrated promising matching performance on synthetic patient summaries and shorter screening time in a small user study, but it did not establish increased enrollment or clinical benefit.

Outcome prediction: CTO is a large outcome-label benchmark, not a predictive model. HINT, CTO, TOP, and TrialBench support retrospective research. Prospective timestamped forecasting and decision-impact studies are still required.

Investigator standard: Validate protocol logic, preserve model and data versions, review false negatives and disparities, retain human authority for safety and endpoint decisions, and use the current core trial guidance with SPIRIT-AI, CONSORT-AI, or DECIDE-AI as applicable.


Quick Decision Trees

“Should I Deploy This AI Tool?”

START: Does the intended function fall within FDA device oversight, another regulatory regime, or non-device clinical practice? - Verify the applicable pathway and exact record - Do not infer clinical utility from authorization alone

Does the evidence match the intended population, workflow, version, and endpoint? - NO → Do not make unsupported performance or benefit claims; define the evidence gap - YES → Continue

Has use in practice been evaluated with an appropriate comparator? - NO → Consider silent evaluation or a governed pilot proportionate to risk - YES → Determine whether the measured endpoint was diagnostic, workflow, behavioral, or patient relevant

Does it match the local population and operating conditions? - NO → Do not assume transportability; evaluate locally before consequential use - YES → Continue

Can I afford false positives? - Calculate expected false alerts - Assess alert fatigue risk - Pilot before full deployment

Decision: Deploy only when the evidence, regulation, workflow, governance, fallback, and monitoring plan jointly support the intended use.


“Should I Trust This LLM Output?”

Is it medical fact (drug dose, diagnosis, treatment)? - YES → VERIFY against authoritative source - NO → Continue

Could an error cause clinical harm? - YES → Use independent verification proportionate to the consequence - NO → Continue with ordinary quality controls

Can I cite the source? - LLM provides citation → CHECK IT (often fabricated) - No citation → Treat as unverified

Decision: Use as draft/idea, never final answer for medical decisions


The Ultimate Clinical Bottom Lines

For All Physicians:

  1. Demand claim-matched evidence: Design, population, comparator, and endpoint determine what can be concluded.
  2. Use multiple metrics: PPV at the relevant prevalence matters, alongside sensitivity, specificity, calibration, uncertainty, workload, and consequences
  3. Evaluate transportability: Test performance and workflow locally when risk and evidence gaps justify it
  4. Continuous monitoring: Performance drifts over time
  5. Define responsibility: Clinical, institutional, manufacturer, and legal duties are task- and jurisdiction-specific.
  6. Start small: Pilot, learn, expand cautiously
  7. Alert fatigue is real: Optimize thresholds carefully
  8. Privacy first: Use patient data only within approved technical, contractual, and legal controls
  9. Transparency matters: Patient communication and consent should reflect the use, risk, setting, and applicable requirements
  10. Stay informed: Field evolving rapidly

Red Lines (Do NOT Cross):

  • DO NOT Deploy consequential AI without evidence and governance proportionate to risk
  • DO NOT Trust vendor claims without verification
  • DO NOT Ignore high false positive rates
  • DO NOT Skip local pilot testing
  • DO NOT Enter protected or sensitive patient data into an unapproved LLM service
  • DO NOT Rely on AI for urgent life-threatening decisions without verification
  • DO NOT Assume AI works equally for all patient populations

Next Steps

To implement AI safely: 1. Read relevant specialty chapter (Part II) 2. Review evaluation framework (Evaluating AI Clinical Decision Support Systems) 3. Check physician toolkit for specific tools (AI Tools Every Physician Should Know) 4. Assess LLM use cases if applicable (Large Language Models in Clinical Practice) 5. Plan local pilot with monitoring 6. Document everything 7. Iterate based on real-world performance

Remember: AI is a set of task-specific tools, not a general substitute for clinical judgment. Use it within a defined evidence boundary, monitor performance, and prioritize patient safety.


For complete details, evidence, and implementation guidance, read the full chapters.


Questions About This Quick Reference

What is the most important AI metric for clinicians?

No single metric is most important in every use case. Positive predictive value estimates how often a positive result is correct at the evaluated prevalence. Clinical evaluation should also address sensitivity, specificity, calibration, uncertainty, workflow consequences, and patient-relevant outcomes.

Will AI replace physicians?

Current evidence does not establish wholesale replacement of physicians. Some authorized systems perform narrow autonomous functions, while most clinical systems support bounded tasks that still require defined human, institutional, and manufacturer responsibilities.

What are the biggest red flags when evaluating medical AI?

Important warning signs include evidence that does not match the intended population or workflow, undisclosed denominators, unclear reference standards, unsupported performance claims, inaccessible regulatory records, absent subgroup analysis, and no plan for postdeployment monitoring.

Can clinicians use ChatGPT for patient care?

Use depends on the task, product, contract, configuration, institutional approval, and applicable privacy rules. Do not enter protected or sensitive patient information into an unapproved service. Clinically consequential output requires verification against authoritative sources and the patient record.

Which AI tools have the strongest evidence for clinical use?

Evidence is product-, version-, task-, population-, comparator-, and endpoint-specific. IDx-DR had a prospective single-arm pivotal study, not a randomized trial. Viz.ai has randomized workflow evidence for selected stroke intervals without established functional-outcome benefit. Evidence from one product must not be transferred to RapidAI, HeartFlow, or another tool.