Using This Appendix
This appendix compiles selected chapter TL;DRs (Too Long; Did not Read summaries) for rapid reference. It is designed for:
- Quick review before implementing AI tools
- Refreshing key concepts
- Finding specific information across chapters
- Sharing with colleagues who need executive summaries
Each TL;DR includes: - The clinical context - Key evidence and applications - What works vs. what does not - Critical takeaways
For full details, citations, limitations, and implementation guidance, use each linked chapter as the controlling source.
Part I: Foundations
Chapter 1: History of AI in Medicine
Key Lesson: Controlled technical performance does not establish clinical adoption or patient benefit.
Major Failures: - MYCIN (1970s): Experts rated 65% of its proposed regimens acceptable in a ten-case blinded evaluation, but it never entered routine clinical care (Yu et al., 1979) - IBM Watson for Oncology: Published evaluations mainly measured concordance, not patient outcomes, and results varied across cancers and settings - Google Flu Trends: Ran high for 100 of 108 weeks from August 2011 through September 2013 and exposed the risks of opaque, unstable data-generating processes (Lazer et al., 2014)
What Works: Narrow, well-defined tasks with an explicit population, reference standard, workflow, and evidence boundary. IDx-DR is one example, authorized through De Novo classification for a specified diabetic-retinopathy screening use after a prospective single-arm pivotal study (Abràmoff et al., 2018).
Chapter 2: AI Fundamentals for Clinicians
Core Concept: AI learns patterns from data, does not follow explicit rules
Critical Metrics: - PPV depends on prevalence and the evaluated population; no single metric is sufficient for every clinical use. - AUC summarizes discrimination across thresholds but does not establish calibration, utility, safety, or patient benefit - Sensitivity vs. Specificity trade-offs
Key Limitations: - Black-box problem (cannot explain reasoning) - Distribution shift (works at Hospital A, fails at Hospital B) - Bias amplification (training data biases → algorithmic biases)
Chapter 3: Clinical Data Challenge
Reality: Clinical data is messy - missingness, heterogeneity, temporal complexity, bias
Critical Issues: - Missing data is NOT random (sicker patients have more data) - EHR data quality variable (copy-paste errors, billing optimization) - External validation addresses transportability within tested conditions; prospective workflow and comparative outcome studies answer different questions
Demand: Evidence that matches the intended population, setting, task, comparator, reference standard, workflow, and clinical consequence.
Part II: Clinical Specialties
Chapter 4: Radiology
Maturity: Radiology has the largest representation on FDA’s periodically updated public list of AI-enabled medical devices, but authorization volume is not a measure of clinical benefit.
Strong Evidence: - Diabetic retinopathy screening (IDx-DR): De Novo-authorized for a narrow autonomous use after a prospective single-arm study, not an RCT - ICH triage: selected products have diagnostic or workflow evidence; verify the exact product and endpoint - LVO stroke: a Viz.ai cluster-randomized trial shortened selected treatment intervals without significantly improving 90-day functional independence (Martinez-Gutierrez et al., 2023) - Mammography: the MASAI randomized trial provides screening-specific evidence; results should not be transferred to every mammography product (Lång et al., 2023)
Workforce boundary: Current evidence supports task redistribution and augmentation scenarios, not a settled prediction of wholesale radiologist replacement.
Surgery, Anesthesiology, and Perioperative Care
Current State: Surgical AI is strongest before and after the operation: risk prediction, planning, video review, quality measurement, and postoperative monitoring. Intraoperative AI remains high risk because surgical decisions are immediate and often irreversible.
Robotic Surgery: Clinical robotic surgery is teleoperated. AI is emerging mainly in surgical video analytics, privacy-preserving multicenter model training, skill assessment, case review, OR scheduling, and research-stage supervised autonomy (Saldanha et al., 2026).
Evaluation Standard: - Separate platform safety, surgeon learning curve, procedure outcomes, AI perception, human factors, and physical autonomy - Verify FDA indication and document number before procurement or patient-facing claims - Demand procedure-specific evidence, not generic robotic-surgery claims - Run silent-mode validation before displaying intraoperative AI outputs
Bottom Line: Robotic surgery evidence is procedure-specific. Preoperative risk prediction has stronger evidence than real-time intraoperative AI. No surgical AI should guide irreversible intraoperative action without independent surgeon verification.
Primary Care, Family Medicine, and Preventive Medicine
Best Applications: - Diabetic retinopathy screening (IDx-DR): narrow authorized use with prospective pivotal evidence - Ambient documentation: emerging randomized and observational evidence, with product, workflow, correction burden, and endpoint boundaries - Chronic disease monitoring (BP, diabetes)
Weak Evidence: - General diagnostic AI (too complex for current systems) - General symptom checkers, where performance varies by task, benchmark, population, and evaluation method
Critical: Workflow integration essential (no time for separate systems in 15-min visits)
Emergency Medicine
Better-supported applications: - LVO stroke detection: Viz.ai has randomized and observational evidence for shorter selected intervals; RapidAI and Brainomix require separate product-specific evidence - ICH and PE triage: evidence must stay tied to the authorized finding, modality, software version, site, and measured workflow endpoint
Controversial: - One external validation of Epic Sepsis Model version 1 reported 33% sensitivity and 12% PPV at the evaluated threshold; it did not test causal patient benefit. (Wong et al., 2021) - Deterioration prediction - Variable results, implementation-dependent
Challenge: Alert fatigue in high-volume EDs
Part III: Implementation
Chapter 23: Evaluating AI Clinical Decision Support Systems
Evaluation logic: 1. Match the claim to the study design and endpoint 2. Verify the exact model, version, population, comparator, and reference standard 3. Distinguish internal validity, transportability, workflow effect, clinical utility, and patient outcomes 4. Treat regulatory authorization, diagnostic accuracy, usability, and outcome benefit as separate questions 5. Plan independent local monitoring before deployment
20 Essential Questions: See full chapter for complete vendor evaluation checklist
Red Flags: - No peer-reviewed publications - No external validation - Vendor refuses to share performance data - Claims 99%+ accuracy
Quick Decision Trees
“Should I Trust This LLM Output?”
Is it medical fact (drug dose, diagnosis, treatment)? - YES → VERIFY against authoritative source - NO → Continue
Could an error cause clinical harm? - YES → Use independent verification proportionate to the consequence - NO → Continue with ordinary quality controls
Can I cite the source? - LLM provides citation → CHECK IT (often fabricated) - No citation → Treat as unverified
Decision: Use as draft/idea, never final answer for medical decisions
The Ultimate Clinical Bottom Lines
For All Physicians:
- Demand claim-matched evidence: Design, population, comparator, and endpoint determine what can be concluded.
- Use multiple metrics: PPV at the relevant prevalence matters, alongside sensitivity, specificity, calibration, uncertainty, workload, and consequences
- Evaluate transportability: Test performance and workflow locally when risk and evidence gaps justify it
- Continuous monitoring: Performance drifts over time
- Define responsibility: Clinical, institutional, manufacturer, and legal duties are task- and jurisdiction-specific.
- Start small: Pilot, learn, expand cautiously
- Alert fatigue is real: Optimize thresholds carefully
- Privacy first: Use patient data only within approved technical, contractual, and legal controls
- Transparency matters: Patient communication and consent should reflect the use, risk, setting, and applicable requirements
- Stay informed: Field evolving rapidly
Red Lines (Do NOT Cross):
- DO NOT Deploy consequential AI without evidence and governance proportionate to risk
- DO NOT Trust vendor claims without verification
- DO NOT Ignore high false positive rates
- DO NOT Skip local pilot testing
- DO NOT Enter protected or sensitive patient data into an unapproved LLM service
- DO NOT Rely on AI for urgent life-threatening decisions without verification
- DO NOT Assume AI works equally for all patient populations
Next Steps
To implement AI safely: 1. Read relevant specialty chapter (Part II) 2. Review evaluation framework (Evaluating AI Clinical Decision Support Systems) 3. Check physician toolkit for specific tools (AI Tools Every Physician Should Know) 4. Assess LLM use cases if applicable (Large Language Models in Clinical Practice) 5. Plan local pilot with monitoring 6. Document everything 7. Iterate based on real-world performance
Remember: AI is a set of task-specific tools, not a general substitute for clinical judgment. Use it within a defined evidence boundary, monitor performance, and prioritize patient safety.
For complete details, evidence, and implementation guidance, read the full chapters.
Questions About This Quick Reference
What is the most important AI metric for clinicians?
No single metric is most important in every use case. Positive predictive value estimates how often a positive result is correct at the evaluated prevalence. Clinical evaluation should also address sensitivity, specificity, calibration, uncertainty, workflow consequences, and patient-relevant outcomes.
Will AI replace physicians?
Current evidence does not establish wholesale replacement of physicians. Some authorized systems perform narrow autonomous functions, while most clinical systems support bounded tasks that still require defined human, institutional, and manufacturer responsibilities.
What are the biggest red flags when evaluating medical AI?
Important warning signs include evidence that does not match the intended population or workflow, undisclosed denominators, unclear reference standards, unsupported performance claims, inaccessible regulatory records, absent subgroup analysis, and no plan for postdeployment monitoring.
Can clinicians use ChatGPT for patient care?
Use depends on the task, product, contract, configuration, institutional approval, and applicable privacy rules. Do not enter protected or sensitive patient information into an unapproved service. Clinically consequential output requires verification against authoritative sources and the patient record.