Primary Care, Family Medicine, and Preventive Medicine
Primary care combines undifferentiated symptoms, prevention, chronic disease, multimorbidity, longitudinal records, and care coordination under substantial time pressure. AI applications range from ambient documentation and autonomous retinal screening to general-purpose copilots. Evidence must remain tied to the exact task, product, population, workflow, and endpoint. A narrow diagnostic result or simulated-vignette score does not establish broad primary care utility.
After reading this chapter, readers will be able to:
- Identify AI tools that integrate into primary care workflows
- Evaluate evidence for AI-assisted screening and prevention
- Understand chronic disease management AI applications
- Assess clinical decision support systems for primary care
- Navigate the unique challenges of AI in outpatient settings
- Recognize appropriate use cases vs. overhyped applications
- Address patient communication about AI-assisted care
Professional Society Guidelines on Primary Care AI
American Academy of Family Physicians (AAFP)
The AAFP’s official Ethical Application of Artificial Intelligence in Family Medicine policy addresses the patient-physician relationship, primary-care functions, transparency, bias, training-data diversity, privacy, systems design, accountability, and trustworthiness. The current official page is dated August 22, 2023.
AAFP Core Principles for Ethical Application of AI in Family Medicine:
Preserve and Enhance the Patient-Physician Relationship: As the patient-physician dyad expands to a triad with AI, the relationship must at minimum be preserved, and ideally enhanced.
Support the 4 C’s of Primary Care: AI/ML must enhance first Contact, Comprehensiveness, Continuity, and Coordination of care.
Expand Capacity and Capability: AI should help achieve the Quintuple Aim in family medicine practice.
Transparency Requirements: Companies must provide:
- Clear information on training data used
- Understandable descriptions of how AI makes predictions
- Documentation addressing implicit bias in design
Rigorous Evaluation: AI/ML should be evaluated with the same rigor as any other healthcare tool.
AAFP Survey Initiative (2024–2025):
In 2024, the AAFP and Rock Health surveyed primary-care physicians and clinicians about digital-health and AI adoption, needs, and concerns. The survey is descriptive member-research, not a clinical practice guideline or effectiveness study (AAFP survey report, 2025).
February 2025 survey reporting emphasized:
- Strong interest in AI among family physicians, tempered by cautious optimism
- Primary concern: AI could reduce administrative burdens, allowing more focus on patient care
- Key barriers: Implementation concerns, training needs, and data privacy
Health IT End-Users Alliance Consensus Statement (April 2025):
The AAFP participated in the Health IT End-Users Alliance consensus statement calling for:
- Common principles balancing AI innovation with appropriate guardrails
- Regulatory oversight as AI adoption accelerates
- Responsible and secure AI development, training, implementation, and monitoring
The Alliance later described the statement in an official 2026 response to the U.S. health-sector AI request for information (Health IT End-Users Alliance, 2026). It is policy guidance, not evidence that a particular product improves outcomes.
For current AAFP guidance: aafp.org/artificial-intelligence
AiM-PC Curriculum (AAFP/STFM/ABFM)
The Artificial Intelligence and Machine Learning for Primary Care curriculum was developed with support from the Society of Teachers of Family Medicine and American Board of Family Medicine Foundation. The current page lists AAFP credit approval from December 13, 2025, through December 12, 2026. Accreditation is not a clinical endorsement of any product.
Curriculum Goals:
- Equip learners with skills to be engaged AI stakeholders
- Guide appropriate use of AI/ML in practice
- Ensure responsible and ethical AI/ML application
For the AiM-PC curriculum: stfm.org/aim-pc
AI Applications in Primary Care
### 1. Clinical Decision Support (CDS)
Evidence-based guideline prompts: - Preventive care reminders (screening mammography, colonoscopy, vaccinations) - Chronic disease management (diabetes targets, hypertension goals) - Drug-drug interaction checking - Reality check: Most CDS predates modern AI, suffers from alert fatigue - Evidence: Mixed. Can improve adherence to guidelines but often ignored (Bates et al., 2003)
Diagnostic support: - Differential diagnosis generation (Isabel, DXplain) - Symptom checking (limited evidence for accuracy) - Limitation: Broad differential does not narrow possibilities without clinical judgment - Use case: Educational tool, rare disease consideration, confirmation of thinking
Autonomous diagnosis in primary care: General-purpose systems have not established safe autonomous use across undifferentiated presentations, comorbidity, longitudinal context, and care transitions.
Conversational diagnostic AI: AMIE, a diagnostic-dialogue system optimized through simulated patient conversations, was evaluated in a randomized, double-blind crossover study using 159 clinical scenarios. It outperformed primary care physicians on most prespecified axes of history taking, diagnostic accuracy, management, communication, and empathy in a text-based standardized setting (Tu et al., 2025). The result is a serious signal for future primary care copilots, but it is not real-world deployment evidence. The study used controlled scenarios, simulated interactions, and no live patient outcomes.
A 2026 JAMA Network Open cross-sectional vignette evaluation of 21 off-the-shelf LLMs found that final-diagnosis failure rates were below 0.40, while differential-diagnosis failure rates exceeded 0.80 in every model (Rao et al., 2026). Strong final-diagnosis accuracy on complete vignettes can hide weak differentials. This is a vignette benchmark, not bedside utility, and it does not authorize unsupervised deployment.
LLM-based clinical copilots:
A pragmatic cluster-randomized trial evaluated AI Consult 2.0 across 16 primary-care facilities in Kenya. The trial included 9,691 patients and 103 clinical officers. Expert-adjudicated treatment failure within 14 days occurred in 102 of 4,693 patients (2.2%) receiving AI-supported care and 94 of 4,654 (2.0%) receiving usual care (adjusted odds ratio 0.77, 95% CI 0.55–1.08; P = .13). The primary clinical outcome did not differ significantly, and no serious adverse event was judged related to the intervention (Agweyu et al., 2026). The strongest patient-level trial was null on its primary clinical endpoint.
How AI Consult works:
- LLM runs in background during patient visits, integrated into EHR
- Analyzes documentation at key points; flags potential errors
- Three-tier alert system: green (no concern), yellow (minor concern), and red (critical concern intended to focus clinician attention)
- Clinician retains full control; AI identifies errors for verification, does not take autonomous actions
Key implementation factors:
- Clinically aligned design: The model was integrated into clinical documentation and adapted to local epidemiology and guidelines
- Version-specific evidence: The effect estimate applies to AI Consult 2.0, the trial’s implementation program, facility network, and follow-up period
- Endpoint discipline: Process or documentation improvement should not be reported as improved patient health when the prespecified patient outcome was null
Separate safety evaluation:
A retrospective safety review evaluated AI Consult V1 in the same 16-clinic network and examined 1,469 records (Agweyu et al., 2026). Clinical management guidance aligned with local guidelines in 99% of encounters, and hallucinations occurred in 3.4%. Reviewers identified potentially harmful recommendations in 115 encounters (7.8%); clinicians fully or partially adopted 67 of those recommendations. No LLM-induced change was identified in 917 encounters. The safety review evaluated V1, while the randomized trial evaluated AI Consult 2.0; results should not be transferred across versions without qualification.
Why this matters for primary care:
The randomized effectiveness result and the separate safety review illustrate a recurring pattern: process measures can improve while the primary clinical outcome remains unchanged, and potentially harmful outputs can still reach care when clinicians defer to generated content.
Specialist-transition communication: A randomized trial involving 111 specialists across 24 disciplines at two centers and 2,069 patients found that an AI-supported transition workflow shortened specialist consultation time and improved customized communication while adding preconsultation work. It did not evaluate clinical outcomes (Tao et al., 2026).
Evaluation framework: PRIMARY-AI is a proposed framework for evaluating AI in primary care, not a professional-society consensus standard or proof of effectiveness (Zeng et al., 2026).
2. Diabetic Retinopathy Screening
IDx-DR (De Novo DEN180001, 2018): - FDA authorized a bounded autonomous diabetic-retinopathy screening use through De Novo classification (FDA DEN180001) - Evidence: 87.2% sensitivity and 90.7% specificity in a prospective pivotal study, not a randomized trial (Abràmoff et al., 2018) - Use case: Primary care offices lacking ophthalmology access - Workflow: Non-mydriatic retinal camera + AI interpretation - Boundary: Authorization and pivotal diagnostic performance do not by themselves establish improved screening completion, referral completion, or visual outcomes
EyeArt (Eyenuk, FDA 510(k) K200667, August 2020): - The authorized indications, output categories, compatible cameras, and workflow are described in the FDA 510(k) summary - Evidence and labeling for EyeArt should not be transferred to IDx-DR or another retinal system
Clinical pathway: - A positive output requires the product-specific referral pathway - A screening system does not manage confirmed retinopathy or evaluate every ocular disease - Screening completion, imageability, referral completion, treatment, and visual outcomes are separate endpoints
Limitation: - Image quality requirements (dilated pupils often needed for optimal performance) - Cannot assess other eye pathology (glaucoma, macular degeneration) - Requires retinal camera acquisition and infrastructure
3. Cardiovascular Risk Prediction
Traditional risk calculators enhanced with AI: - ASCVD risk calculator - Framingham risk score - AI additions: Incorporating more variables (retinal imaging, ECG patterns, genetic data)
Retinal imaging for CV risk: - AI analyzes retinal fundus photos to predict CV events - Evidence: Correlates with cardiovascular disease risk Poplin et al., 2018 - Status: Research stage, not yet clinical standard - Promise: Non-invasive risk assessment during diabetic eye screening
ECG-based risk prediction: - AI analysis of 12-lead ECG predicts atrial fibrillation, heart failure, mortality - Regulatory status and intended use are product-specific; research prediction studies should not be presented as authorized clinical interventions - Clinical utility: Identifying high-risk patients for preventive interventions
Caveat: Adding complexity to risk assessment requires evidence that it improves outcomes, not just correlates with risk
4. Hypertension Management
Remote blood pressure monitoring and decision support: - Smartphone-connected BP cuffs - Digital workflows can flag concerning trends or support medication adjustment; remote monitoring is not automatically AI - Evidence: An expert position paper reviews telemedicine and remote blood-pressure management, not a single machine-learning intervention (Omboni et al., 2020) - Limitation: Requires patient adherence to monitoring, connectivity
Atrial fibrillation detection: - Smartwatch-based screening; the evidence below is specific to Apple Watch - Evidence: Apple Heart Study found 0.52% of participants received irregular pulse notifications. Among those notified, PPV was 84% when compared to simultaneous ECG recording. On subsequent ECG patch monitoring (applied ~13 days later), 34% showed AFib (Perez et al., 2019) - Clinical challenge: Low notification rate but high patient anxiety. Many who receive alerts will have normal subsequent monitoring (AFib is intermittent) - Best use: High-risk populations, not universal screening
5. Documentation and Administrative AI
Ambient clinical documentation: - Systems: Nuance DAX, Abridge, Suki, DeepScribe - Function: Listen to patient encounter, auto-generate clinical note - Evidence: Randomized trials show product-specific, not class-wide, documentation effects. A 2025 three-arm parallel trial of 238 outpatient physicians found that Nabla reduced time in notes by 9.5%, while the DAX comparison did not significantly change that endpoint (Lukac et al., 2025). A stepped-wedge trial of 66 practitioners found that Abridge reduced time on notes by 0.36 hours per day and reduced work exhaustion (Afshar et al., 2025). - Physician satisfaction: High. More face time with patients, less screen time; results vary by specific tool (DAX showed non-significant change in the UCLA RCT)
How it works: 1. AI records and transcribes conversation 2. Natural language processing extracts key information 3. Auto-generates SOAP note draft 4. Physician reviews, edits, signs
Limitations: - Requires review (AI makes errors, misses nuance) - Privacy concerns (recording conversations) - May miss non-verbal cues - Accuracy varies with accents, background noise, complex cases
Prior authorization automation: - AI extracts required information from EHR - Auto-fills prior auth forms - Potential benefit: Reduces manual data extraction and form completion - Limitation: Payer requirements remain variable, and generated submissions require review - Market context: A March 2026 trade-press report described venture funding and health-system reach for one vendor (MobiHealthNews, 2026). Funding and customer counts are not clinical-effectiveness evidence. Medication-approval outputs require human review before action.
Inbox management: - AI triages patient messages (urgent vs. routine) - Auto-generates response drafts for common questions - Status: Emerging, variable quality
6. Preventive Care and Screening
Identifying patients due for screenings: - AI scans EHR to identify missed screenings (mammography, colonoscopy, cervical cancer, lung cancer) - Outreach campaigns targeting overdue patients - Evidence: Improves screening rates when paired with outreach (Sequist et al., 2009)
Electronic flu-vaccine nudges (older adults): Yang et al. randomized 2,167 adults aged 60 or older from 80 family-doctor teams in Lanzhou New Area to seven electronic nudge messages or control (Yang et al., 2026). Overall influenza-vaccination willingness was 19.20%. In previously vaccinated participants (n = 458), most nudges were associated with lower willingness versus control (loss-framing OR 0.40, 95% CI 0.21–0.77). The authors report that a video nudge increased willingness among unvaccinated participants, particularly those worried about side effects. This is a willingness endpoint from one district, not evidence that blasting EHR or SMS flu reminders raises documented uptake or reduces influenza.
Lung cancer screening eligibility: - AI identifies patients meeting USPSTF criteria (age, smoking history) - Auto-generates orders or alerts - Challenge: Requires accurate smoking history documentation
Social determinants of health (SDOH) and care management: - AI identifies at-risk patients from EHR data (housing instability, food insecurity) - Evidence: Can predict risk, but interventions for SDOH remain challenging - Emerging: A mixed-methods study of 3,175 Medicaid beneficiaries compared an offline reinforcement-learning recommender with experience-based care management. Counterfactual analyses estimated a 12-percentage-point reduction in acute-care events (95% CI 2.2–21.8), but the study did not prospectively deploy the recommendations or measure an observed treatment effect (Basu et al., 2025).
7. Chronic Disease Management
Diabetes: - Continuous glucose monitor (CGM) data interpretation: - Pattern recognition (nocturnal hypoglycemia, post-prandial spikes) - Insulin dosing recommendations (emerging) - Limitation: Not yet integrated into most primary care EHRs
Medication adherence prediction:
- AI predicts which patients likely to be non-adherent
- Targeted interventions
Evidence boundary: Predictive discrimination does not establish that targeting patients improves adherence or outcomes
Risk stratification:
- Predicting progression to complications (retinopathy, nephropathy, neuropathy)
- Identifying high-risk patients for intensive management
Hypertension: - Home BP monitoring with AI alerts - Medication optimization algorithms - Evidence: Mixed. Some studies show BP improvement, others no benefit over usual care
Asthma/COPD: - Spirometry interpretation - Exacerbation prediction from symptom tracking - Status: Limited deployment in primary care
Depression: - Screening: AI-enhanced PHQ-9 interpretation - Monitoring: Smartphone apps with AI-based mood tracking - Chatbots: AI-driven cognitive behavioral therapy (Woebot, Wysa) - Evidence: Some chatbots show modest benefit for mild-moderate depression Fitzpatrick et al., 2017 - Limitation: The cited trial does not establish safety or efficacy for severe depression, crisis care, or suicidality
8. Patient Triage and Scheduling
Symptom checkers: - Examples: Ada, Babylon, K Health, Buoy Health - Function: Patient inputs symptoms → AI suggests possible diagnoses, urgency level - Evidence: Accuracy variable (30-60% for correct diagnosis in top 3) (Semigran et al., 2015) - Limitation: Patients often do not know what information is relevant, AI cannot examine - Boundary: Diagnostic-vignette accuracy does not validate a real-world triage or disposition pathway
AI-assisted telehealth triage: A Cedars-Sinai study of 461 virtual urgent care visits found AI recommendations matched or exceeded final physician recommendations in guideline adherence, with AI flagging antibiotic-resistant UTIs more consistently than clinicians (Zeltzer et al., 2025). However, AI was vendor-developed (K Health), not independently validated, and concordance between initial AI and final physician recommendations was only 56.8%.
Digital front-door systems are not always AI, but their adoption patterns define the workload and safety environment for AI triage. In a BMJ Digital Health & AI observational study of eConsult across Great Britain, Kerr et al. analyzed 43,657,891 online consultations from 2019 to 2023; 89.0% were submitted to the patient’s practice and 7.7% were directed to urgent or emergency care, with use varying by country, region, and deprivation (Kerr et al., 2025). For family physicians, portal triage should be evaluated as a workload, safety, and equity intervention, not only as an access convenience.
Osteoporotic vertebral-fracture video triage (OVCF): A multicenter 8 August 2026 npj Digital Medicine study developed a two-stage screening framework that combines prompt-guided multimodal large-language-model feature extraction from standardized posture images and functional videos with machine-learning classification, trained on 204 participants (102 OVCF, 102 controls). On geographically independent external validation (n = 56), a Gradient Boosting model reached AUC 0.838 (95% CI 0.712–0.940), sensitivity 89.3%, and specificity 71.4%, with ΔAUC < 0.001 from internal to external testing and a Brier score of 0.168 (0.151 after post-hoc isotonic recalibration) (Zhang et al., 2026). The authors frame this as community and primary-care triage to prioritize confirmatory imaging, not as a replacement for radiography; external n = 56 is small, and a related method patent is disclosed.
Appointment scheduling optimization: - AI predicts no-show risk - Optimizes schedule templates - Impact: Reduces gaps in schedule, improves access
9. Patient Education and Communication
Large language models (LLMs) for patient questions: - Examples: ChatGPT, Google Med-PaLM 2 - Use case: After-visit question answering, general health information - Evidence: Can provide accurate information for common questions (Singhal et al., 2023) - Critical limitations: - Hallucinations (confidently incorrect information) - Consumer systems may lack verified access to complete, current patient-specific data - Responsibility and escalation pathways may be unclear - Medical-question benchmark performance does not establish safe patient-specific advice
Personalized patient education materials: - AI generates health literacy-appropriate explanations - Tailored to specific diagnosis, reading level - Status: Emerging
10. Population Health Management
Risk stratification: - Predicting which patients will have ER visits, hospitalizations - Use case: Targeting high-risk patients for care management programs - Evidence: A systematic review found substantial heterogeneity and generally limited discrimination among readmission-risk models; it does not establish accurate identification across primary-care populations (Kansagara et al., 2011) - Challenge: Effective interventions for identified high-risk patients remain elusive
Care gap identification: - AI identifies patients with unmet preventive care, chronic disease management needs - Prioritizes outreach - Impact: Supports value-based care, quality metrics
What Does NOT Work Well in Primary Care:
General diagnostic AI: - No cited study establishes safe autonomous diagnosis across the full breadth of primary care - Context, patient history, social factors critical, hard to capture - No validated AI for broad primary care diagnostic reasoning
Replacing physician clinical judgment: - Complexity, uncertainty, patient preferences require human judgment - Longitudinal relationships and trust central to primary care
Autonomous treatment decisions: - Medication management requires considering allergies, interactions, patient preferences, cost - AI recommendations often lack context
Complex visit summarization: - AI struggles with nuanced discussions (goals of care, family dynamics, complex psychosocial issues)
Workflow Integration Challenges:
EHR Integration: - Most AI tools require separate logins, interfaces → workflow disruption - Poor integration → physician resistance - Integration should minimize duplicate entry while preserving review, provenance, and recovery pathways
Time Constraints: - 15-20 minute visits leave little time for AI interaction - AI must be faster than physician’s current workflow or provide substantial value
Heterogeneity: - Primary care sees all ages, all conditions - AI trained on specific populations may not generalize
Data Quality: - Outpatient data less structured than inpatient - Medication lists often inaccurate - Social history, family history poorly documented
Evidence-Based Assessment:
What Has Strong Evidence:
Diabetic retinopathy screening: Prospective pivotal diagnostic study plus bounded De Novo authorization, not a randomized trial (Abràmoff et al., 2018; FDA DEN180001)
Clinical decision support for preventive care: Multiple RCTs show improved screening rates (though mixed quality)
Remote monitoring for chronic diseases: Digital-care evidence exists for selected workflows, but an intervention should not be labeled AI unless the evaluated system used AI
Ambient documentation: High user satisfaction, time savings (long-term outcomes pending)
What Needs More Evidence:
Symptom checkers: Accuracy variable, clinical impact uncertain
AI-enhanced risk prediction: Correlates with outcomes but unclear if changes management improves outcomes
Mental health chatbots: Modest evidence for mild symptoms, not suitable for moderate-severe
Most population health AI: Can identify high-risk patients, but interventions often ineffective
What Lacks Evidence:
General diagnostic AI for primary care: No evidence cited here establishes autonomous clinical utility across undifferentiated primary-care presentations
Autonomous treatment recommendations: Not ready for unsupervised deployment
Most patient-facing AI: Limited validation for accuracy, clinical impact
Practical Implementation Guidance:
Clinical Value: - Does this solve a defined problem in the intended practice? - Will it improve patient outcomes, efficiency, or satisfaction? - What is the evidence from primary-care settings rather than a transferred specialty task?
Workflow Integration: - Does it integrate with the local EHR? - Will it add clicks/time or save time? - Can medical assistants/nurses operate it?
Patient Acceptability: - Will patients accept AI-assisted care? - How do I explain AI use to patients? - What if patient refuses AI?
Financial: - What is the total cost, including licensing, hardware, integration, personnel, monitoring, and incident response? - Is there reimbursement? - Which measured benefits and costs define the local business case?
Validation: - Has it been tested in primary care (not just specialty/hospital)? - Does it work for the intended population, including age, language, comorbidity, disability, and relevant subgroups? - What’s the false positive rate (will it create more work)?
Liability: - How do applicable law, institutional policy, contracts, insurance, and workflow allocate responsibilities? - Has the organization reviewed coverage for the intended AI-assisted use? - What documentation is required?
Patient Communication About AI:
What to Tell Patients:
“We use AI as an assistive tool to help with [specific task: screening, documentation, etc.]”
“I review all AI recommendations before making decisions”
“AI helps me focus more time on you rather than the computer”
“The evidence for this specific tool, task, and population is [describe accurately]”
What to Avoid:
“The AI makes the diagnosis” when the system is not authorized and used for an autonomous diagnostic function
“The AI is always right” (it makes errors)
“We’re using you to test AI” (only use validated tools clinically)
Permission, disclosure, and consent: - For ambient documentation, explain recording, data use, retention, review, and available alternatives under applicable policy and law - For autonomous diagnostics, explain the system’s role, possible outputs, follow-up pathway, and limitations - Research and quality-improvement activities require the review and permissions applicable to their design and jurisdiction
Addressing Patient Concerns:
“I don’t want AI involved in my care” - Respect preference - Explain AI role (assistive, not autonomous) - Offer alternative (traditional care)
“Will AI replace my doctor?” - No. AI is tool, physician judgment remains central - Emphasize relationship, trust, personalized care
“Is my data being used to train AI?” - Explain your practice’s data use agreements - HIPAA protections - Patient rights to opt out if available
Future Directions in Primary Care AI:
Demonstrated or deployed: - Product-specific ambient documentation can change time-in-note and work-exhaustion endpoints - Bounded autonomous diabetic-retinopathy screening has prospective pivotal evidence and FDA authorization - Preventive reminders and digital monitoring can improve selected process measures when linked to a response pathway
Under evaluation: - General-purpose copilots for longitudinal clinical decision support - AI-supported specialist transitions, inbox management, and prior authorization - Population-health recommendation systems and patient-facing education
Beyond current evidence: - Autonomous replacement of comprehensive primary-care diagnosis and treatment - Continuously learning local systems without validated change control and monitoring - Assured time savings, equity, or patient-outcome benefit from deployment alone
Regulatory and Reimbursement:
Current State: - Coverage and payment are code-, payer-, service-, setting-, and product-specific - FDA authorization does not establish reimbursement, cost-effectiveness, or a positive return on investment - Value-based arrangements can alter incentives, but financial value must be measured locally
Advocacy Needed: - CPT codes for AI-assisted services - Quality measures that AI demonstrably improves - Liability clarity - Interoperability standards
The Clinical Bottom Line:
Start with a bounded use case: Autonomous diabetic-retinopathy screening has a specific authorized workflow, while ambient-documentation evidence is product- and endpoint-specific.
Evaluate workflow and clinical endpoints separately: Adoption and time savings do not establish diagnostic accuracy or patient benefit.
Do not claim replacement: Current evidence does not establish safe replacement of primary-care physicians across undifferentiated and longitudinal care.
Treat ambient evidence as product-specific: Some randomized trials report documentation-time or exhaustion benefits, while effects differ by product and outcome.
Separate simulations from care: Diagnostic-vignette performance is not evidence of safe real-world triage, treatment, or follow-up.
Make communication truthful: Explain the system’s actual role, data flow, review process, uncertainty, and route to human help.
Require setting-specific evidence: Specialty, hospital, or benchmark validation does not guarantee primary-care performance.
Measure net workload: Include review, correction, preconsultation work, alert handling, and downstream follow-up rather than measuring generation time alone.
Evaluate the response pathway: Risk identification matters only when a feasible intervention reaches the intended patient and its outcome is measured.
Keep the claim no broader than the design and endpoint: A workflow benefit should not be reported as a diagnostic or patient-outcome benefit.
Specialty-Specific Considerations:
Family Medicine: Broad age range, multimorbidity, prevention, acute symptoms, and continuity make transportability and workflow evaluation especially important.
General Internal Medicine: Adult focus, chronic diseases. AI for disease management most applicable
Pediatrics: Age-specific development, disease patterns, caregiver involvement, consent, and pediatric validation require separate evaluation.
Geriatrics: Multimorbidity, polypharmacy, cognition, function, caregivers, and goals of care can make a single-disease recommendation unsafe or incomplete.
Next Chapter: Pathology and Laboratory Medicine covers digital pathology, diagnostic AI, validation, and deployment boundaries.
Check Your Understanding
The following cases are explicitly hypothetical teaching exercises. Institutions, products, patient details, model outputs, costs, outcomes, legal arguments, fault allocations, verdicts, and settlements are illustrative unless a source is linked in the same sentence. They are not reports of actual events or predictions of legal outcomes.
Scenario 1: Hypothetical Symptom-Checker Under-Triage
A large multispecialty group has integrated a fictional symptom checker into its patient portal to reduce avoidable emergency visits and improve triage efficiency. The case is not about Babylon Health or any other named product.
The patient is someone you know well: 68-year-old retired teacher, type 2 diabetes for 15 years, chronic kidney disease (Stage 3b, baseline creatinine 1.4). She’s usually conscientious about her care.
Sunday evening, she wakes up with a fever. 101.2°F on her home thermometer. Her usual insomnia tonight comes with chills, shaking under two blankets. And that gnawing discomfort when she urinates.
Instead of calling the after-hours triage line (which she’s always found slow and frustrating), she pulls up the patient portal on her phone. The health system just rolled out this symptom checker, and the hospital newsletter made it sound convenient.
She types in: “fever,” “painful urination,” “chills.”
The AI analyzes for maybe three seconds, then displays its conclusion:
“Likely urinary tract infection. Schedule an appointment with your primary care doctor within 2-3 days. Stay hydrated. Monitor symptoms.”
Urgency Level: Low. Primary care visit recommended.
She feels reassured. Low urgency. Just a UTI. She remembers she has leftover ciprofloxacin from last year’s urinary infection, takes one, drinks some cranberry juice, and figures she’ll call your office Monday morning for an appointment.
Monday, 3 PM. She arrives at your clinic for a walk-in slot, and the moment your MA flags you down in the hallway, you know something’s wrong.
Vitals: Temp 103.1°F. Heart rate 118. Blood pressure 88/52. Respiratory rate 24. O2 sat 92% on room air.
She’s lethargic, barely tracking your questions. Altered mental status. Dry mucous membranes. You order STAT labs: WBC 18,500, creatinine 3.2 (baseline 1.4), lactate 4.2.
The diagnosis hits you immediately: sepsis secondary to pyelonephritis, with acute kidney injury layered on her chronic kidney disease.
You start IV fluids, IV ceftriaxone, stabilize her enough to transfer to the ED. She ends up in the ICU for five days, develops acute tubular necrosis requiring temporary dialysis. She survives, but her kidney function never recovers. Now CKD Stage 4, one step from dialysis.
Six weeks later, the family files a malpractice claim alleging delayed diagnosis and inappropriate triage by the symptom checker.
Question 1: What went wrong with the symptom checker triage?
The exercise is designed around a context failure that a structured clinical triage process should test explicitly:
The symptom checker could not see context. Age 68, diabetes, chronic kidney disease. This is not a healthy 25-year-old with simple cystitis. This is a patient at high risk for serious infection. Fever plus chills in an elderly woman with diabetes and impaired kidney function should trigger immediate alarm bells. The appropriate response would have been: “Seek emergency care immediately” or “Call your doctor’s after-hours line NOW,” not “schedule an appointment in 2-3 days.”
What the comparative evidence actually measured: In a 2015 audit of 23 symptom checkers using 45 standardized patient vignettes, the correct diagnosis appeared first in 34% of evaluations and within the top 20 in 58%. Triage advice was appropriate in 57% overall, with better performance for emergency-care vignettes than for self-care vignettes. The systems overtriaged more often than they undertriaged. These results describe the tested products and vignettes, not every contemporary symptom checker, and they do not prove that a particular product would produce the fictional recommendation above (Semigran et al., 2015). A benchmark can reveal important failure modes without establishing the error rate of every product, population, or deployment.
3. Cannot replace clinical assessment
What a clinician-led triage pathway should consider:
- Fever + dysuria + diabetes + CKD = high-risk UTI
- Chills = concern for pyelonephritis or bacteremia
- Current vital signs, mental status, oral intake, symptom trajectory, comorbidity, and access to timely reassessment
- Whether the available facts warrant same-day assessment, emergency evaluation, or another disposition under the local protocol
The symptom checker cannot examine the patient, cannot assess: - Vital signs (fever degree, heart rate, blood pressure) - Appearance (toxicity, altered mental status) - Costovertebral angle tenderness - Hydration status
4. The communication design did not make escalation conditions salient
The patient trusted the AI recommendation despite worsening symptoms: - “Low urgency” reassured her - Delayed seeking care - Self-medicated with leftover antibiotics (inadequate for pyelonephritis)
The fictional deployment did not clearly explain that the output was a limited triage recommendation rather than a diagnosis, that worsening or severe symptoms should override it, and that some patients require clinician review because the model may not adequately represent their risk. A useful interface should present the uncertainty, the escalation conditions, and a direct route to human help without burying them in a disclaimer.
Question 2: Are you liable for malpractice?
Plaintiff’s argument:
1. Vicarious liability for AI system deployed by your health system
“The health system implemented this symptom checker as part of patient care delivery. The physician practice is responsible for the AI system’s recommendations, just as they would be for a triage nurse’s advice.”
The plaintiff would argue that portal integration, institutional branding, workflow design, and the absence of an effective escalation route made the tool part of the health system’s care pathway. Whether that creates a particular duty or vicarious-liability theory depends on jurisdiction and facts. The vignette is not evidence of an established legal precedent.
2. Failure to warn patients about AI limitations
“Patients were not given clear, usable information about the symptom checker’s limitations or the conditions that should trigger immediate clinical contact. The system could have included warnings such as:
- ‘This tool is NOT appropriate for patients with diabetes, kidney disease, or other chronic conditions’
- ‘If you have fever with chills, seek immediate medical attention regardless of this tool’s recommendation’
- ‘This result does not account for every risk factor and can be wrong. Use the call-now option if the recommendation does not fit how ill you feel.’”
The plaintiff could frame the communication failure as inadequate disclosure, negligent design, or failure to warn. The legally controlling theory and required disclosure would depend on applicable law and the product’s role.
3. Alleged deviation from a reasonable triage process
“A triage nurse following standard protocols would have recognized this as high-risk UTI requiring same-day assessment or ED referral. The AI fell below the standard of care, and the health system deployed it anyway.”
The defensible comparison is the disposition that a reasonable clinician-led protocol would have produced from the information actually available. AAFP’s general AI policies do not create a disease-specific legal rule for this fictional case.
4. Foreseeable harm
“The health system knew or should have known that symptom checkers have poor accuracy for emergent conditions. Deploying this tool for all patients, including high-risk elderly with comorbidities, was foreseeable to cause harm.”
Published comparative studies had documented that symptom-checker performance varied across products and vignette types. That evidence would support predeployment testing and safety controls, but it would not by itself prove negligence or causation in this case.
Defense’s argument:
1. Patient did not follow up appropriately
“The symptom checker recommended seeing a primary care doctor within 2-3 days. The patient waited 38+ hours and did not call when symptoms worsened. A reasonable patient would have sought care sooner given fever and chills.”
Contributory negligence: Patient’s delay in seeking care contributed to poor outcome.
2. Symptom checker is patient tool, not physician tool
“This was a patient-initiated triage resource, not a physician’s clinical assessment. Patients use many online resources (WebMD, Google) without physician liability. The symptom checker is analogous.”
Counterargument: Unlike generic websites, this tool was integrated into the patient portal and endorsed by the health system, creating higher duty of care.
3. Patient self-medicated inappropriately
“The patient took leftover ciprofloxacin without physician advice, delaying appropriate care. The symptom checker did not recommend self-treatment.”
Contributory negligence: Patient’s self-medication was independent decision.
Answer 2: The facts create material exposure, but the outcome cannot be predicted from the vignette
Liability would be determined under the law of the relevant jurisdiction and a developed factual record. The relevant questions would include who designed and controlled the tool, how the portal represented it, whether it operated within its intended use, what validation and monitoring the institution performed, what information the patient supplied, and what escalation options were available.
1. Relationship between the tool and the care pathway
Portal integration could support an argument that the institution adopted the system as part of care delivery. It does not automatically establish liability or imply that every output is an institutional medical judgment. Labels, instructions, contracts, clinical governance, and actual workflow would matter.
2. Validation and risk controls
An investigation should ask:
- Was the exact product evaluated in the population and workflow in which it was used?
- Did testing include older adults and people with important comorbidities?
- Were emergency escalation rules tested for sensitivity, usability, and accessibility?
- Was there a direct human-triage option and a procedure for monitoring under-triage?
- Did the deployment remain within the product’s current regulatory status and intended use?
An unfavorable answer could support a breach argument, but no single answer proves negligence by itself.
3. Foreseeability and causation
Semigran and colleagues documented variable diagnostic and triage performance in standardized vignettes. That makes error a foreseeable design concern. It does not establish that this fictional outcome was inevitable, that every competent human triager would have chosen the same disposition, or that earlier evaluation would certainly have prevented the renal outcome.
The plaintiff would still need to connect the output to the delay and the delay to the injury using patient testimony, system logs, expert testimony, and medical evidence. The defense could examine the patient’s self-medication, symptom progression, access to care, and whether the interface instructed the patient to seek help if symptoms worsened. Regulatory status, product performance, institutional negligence, clinical causation, and damages are separate questions.
4. Outcome and allocation of responsibility
No responsible analysis can assign a verdict, settlement amount, or percentage of fault from these facts alone. Those outcomes vary with jurisdiction, evidence, parties, insurance, and procedural posture. The educational value lies in identifying preventable system weaknesses before an injury occurs.
Lessons for Primary Care Teams:
1. Validate AI tools before deployment
Before implementing a patient-facing triage system:
- Verify the exact product, version, intended use, and current regulatory status
- Review product-specific evidence rather than transferring performance from another symptom checker
- Test representative local cases, including common comorbidities, language needs, disability access, and emergency presentations
- Define under-triage and over-triage measures before launch
- Use a supervised pilot and investigate discordant recommendations
2. Warn patients about AI limitations
Communication should describe what the tool does, what information it cannot assess, when its recommendation should be overridden, and how to reach a clinician. A borrowed benchmark error rate should not be presented as the local product’s error rate.
3. Treat portal integration as a clinical-governance decision
External links are not inherently safer than integrated tools. The stronger design is one in which ownership, intended users, escalation routes, monitoring, and incident response are explicit. A disclaimer cannot compensate for an unsafe triage pathway.
4. Ensure triage nurse backup
If after-hours automated triage is offered, preserve a usable route to human assessment and define when the system transfers the patient rather than continuing automated questioning.
5. Document AI limitations in informed consent
When the tool materially participates in care, determine with legal, compliance, clinical, and patient-representation input what should be disclosed and documented:
- What the AI does
- The evidence and limitations applicable to the actual product and version
- When to override AI and seek immediate care
- Whether and when a clinician reviews the interaction
6. Advocate for AI regulation
Regulatory classification is product-specific. A symptom checker can fall inside or outside medical-device oversight depending on its claims, intended use, functions, and jurisdiction. Teams should verify the current record for the exact product rather than assuming that all symptom checkers are either regulated devices or unregulated wellness software.
Scenario 2: Hypothetical Ambient Note Omits a Medication Instruction
A busy primary-care clinic has adopted a fictional ambient documentation system to reduce documentation burden. The error and outcome below are illustrative and should not be attributed to Nuance DAX, Abridge, or another product.
System workflow: 1. AI listens to patient encounter via smartphone app 2. Transcribes conversation 3. Auto-generates SOAP note draft 4. You review, edit, sign
Patient: 74-year-old man with heart failure (HFrEF, EF 30%), atrial fibrillation, hypertension, type 2 diabetes - Current medications: Carvedilol 25 mg BID, lisinopril 20 mg daily, furosemide 40 mg daily, metformin 1000 mg BID, apixaban 5 mg BID
Office visit (15 minutes): - Chief complaint: “Feeling more short of breath, ankles swelling” - Exam: JVP elevated 10 cm, 2+ pitting edema bilaterally, crackles at lung bases - Assessment: Heart failure exacerbation (volume overload)
Your management discussion (captured by AI):
You: “Your heart failure is acting up. You’re retaining fluid. I’m going to increase your water pill from 40 to 80 milligrams daily. Take two of the 40 milligram tablets each morning.”
Patient: “Should I keep taking the blood thinner?”
You: “Yes, definitely keep taking the apixaban. That’s critical for your atrial fibrillation to prevent stroke. Don’t stop that.”
Patient: “What about the other medications?”
You: “Keep everything else the same. The baby aspirin we discussed last time, I don’t think you need that anymore since you’re on apixaban. So STOP the aspirin, but keep the apixaban.”
AI-generated SOAP note (draft):
Assessment/Plan: 1. Heart failure exacerbation (volume overload) - Increase furosemide to 80 mg PO daily - Daily weights - Restrict sodium - Follow up 1 week
- Atrial fibrillation
- Continue current management
- Stop aspirin (redundant with anticoagulation)
Medications reviewed and updated.
Your review: You quickly scan the note, looks reasonable, sign and close encounter.
What the AI MISSED: The note does NOT explicitly list all medications, does NOT confirm apixaban continuation, does NOT clarify which medications to continue vs. stop.
3 weeks later: Patient admitted to hospital with ischemic stroke (right MCA territory) - On admission: Patient stopped taking apixaban - Patient’s explanation: “The doctor said to stop the blood thinner and keep the aspirin. I stopped the apixaban and kept taking my baby aspirin.”
Chart review: The AI note says “stop aspirin” but does NOT say “continue apixaban.”
Your recollection: You clearly told the patient to continue apixaban and stop aspirin, but the AI note is ambiguous.
Outcome: Patient has permanent left-sided weakness, requires rehabilitation
Malpractice claim: Failure to ensure clear medication instructions, inadequate documentation
Question 1: What went wrong with the ambient AI documentation?
Critical failures in AI-generated documentation:
1. AI misinterpreted “stop the aspirin” conversation
The AI heard: - “Stop the aspirin” - “Keep taking the apixaban”
But the AI’s natural language processing (NLP) failed to capture the critical distinction: - Continue apixaban (anticoagulant) - Stop aspirin (antiplatelet)
The AI note documented: - “Stop aspirin” (correct) - MISSING: “Continue apixaban” (critical omission)
Why this can happen in the exercise: Negation, medication names, speaker attribution, and implied continuation are plausible failure points for any documentation pipeline. The fictional omission does not establish a measured error pattern or rate for a named ambient product.
2. Physician failed to catch the omission
The clinician remains responsible for determining whether a note is accurate enough to sign under applicable professional, institutional, and payer requirements. The product’s review instructions and the organization’s policy should make that responsibility operational rather than ceremonial.
You signed the note without verifying: - All medication changes explicitly documented - Critical medications (anticoagulation) confirmed - Unambiguous instructions
Cognitive error: Automation bias, trusting the AI output without critical review.
Time pressure: 15-minute visit, back-to-back patients, quick note review → inadequate verification.
3. No safety checks for high-risk medications
Potential safety controls that were absent from the fictional workflow: - Explicitly document anticoagulation continuation in assessment/plan - Include updated medication list showing apixaban unchanged - Print after-visit summary for patient showing all medications
For high-consequence medication instructions, explicit start, stop, continue, dose, and follow-up language can reduce ambiguity. The exact documentation requirement remains setting-specific.
4. Patient misunderstood verbal instructions
The patient conflated “blood thinner” with “aspirin”: - Heard: “Stop the blood thinner” - Interpreted: “Stop apixaban” (the anticoagulant I take for blood thinning) - Did NOT realize: “Blood thinner” meant aspirin in this context
Communication failure: Ambiguous terminology (“blood thinner” applies to both aspirin and apixaban).
Better communication: - Use medication names, not categories - “Continue apixaban. That’s the one in the orange bottle you take twice a day” - “Stop the baby aspirin. The small white pill” - Provide written after-visit summary with medication list
Question 2: Are you liable for malpractice?
Plaintiff’s argument:
1. Failure to document critical medication instructions
“The medical record contains NO documentation that the patient should continue apixaban. The note says ‘stop aspirin’ but does NOT say ‘continue apixaban.’ This ambiguity led the patient to stop his anticoagulant.”
The plaintiff would argue that a reasonable medication-safety workflow required explicit, unambiguous documentation of the anticoagulation plan in these circumstances.
Experts will testify: A reasonable physician would document: - “Continue apixaban 5 mg BID (no change)” - Or include updated medication list showing apixaban continued - Or provide written after-visit summary
2. Inadequate review of AI-generated note
“The physician relied on AI without adequate verification. The AI-generated note was incomplete and ambiguous, yet the physician signed it without editing.”
Negligence: Delegating documentation to AI does NOT relieve physician of responsibility to ensure accuracy.
Analogy: If a scribe writes an incomplete note, the physician is still liable for signing it.
3. Foreseeable risk from incomplete documentation
“Draft documentation can omit or distort clinically important information. The clinic should have tested the exact product and workflow, monitored medication discrepancies, and made the signing review commensurate with clinical risk.”
The randomized ambient-scribe studies summarized earlier evaluated product-specific workflow and documentation outcomes. They do not establish a universal accuracy rate or prove that medication instructions are the most frequent error type. The foreseeable risk is broader: a draft can be incomplete, and a signed record can propagate that incompleteness.
4. Failure to provide written medication instructions
“The patient was managing several medications and the plan changed two agents commonly described as blood thinners. A written medication list and teach-back could have exposed the misunderstanding before the patient left.”
Kripalani and colleagues reviewed communication and information-transfer problems at hospital discharge, not the exact outpatient ambient-scribe workflow in this vignette. The review supports the general importance of reliable transitions, but it does not establish a universal one-hour recall percentage or a categorical legal requirement for every office visit (Kripalani et al., 2007).
Defense’s argument:
1. Patient misunderstood clear verbal instructions
“The physician clearly stated ‘keep taking the apixaban’ and ‘stop the aspirin.’ The patient’s misunderstanding is not the physician’s fault.”
Counterargument: The adequacy of the communication process, including teach-back and written instructions, should be evaluated from the patient’s needs and the clinical risk rather than the spoken sentence alone.
2. AI note accurately reflected assessment/plan
“The note correctly documented the heart failure management and aspirin discontinuation. The AI is not required to list every medication unchanged.”
Counterargument: Expert review could conclude that explicit continuation language was expected for a high-consequence anticoagulant in this particular encounter.
3. Patient had access to medication list in patient portal
“The patient could have checked the medication list in the patient portal, which showed apixaban as an active medication.”
Counterargument: Elderly patients often do not use patient portals. Physician cannot rely on patient portal access to ensure medication safety.
Answer 2: Liability cannot be resolved from the vignette
The facts support a serious safety review and plausible claims, but they do not establish a verdict. An expert and court would evaluate the applicable standard of care, the product’s intended workflow, the clinician’s actual review, the after-visit instructions, the patient’s understanding, the medication record, and the medical evidence connecting interruption of apixaban to the stroke.
1. Duty and workflow ownership: The clinician-patient relationship creates ordinary clinical duties, but the respective duties of the clinician, institution, and vendor depend on who configured, controlled, represented, and monitored the documentation system.
2. Breach: The signed note’s omission would be relevant. It does not automatically establish breach. The inquiry would compare the complete encounter and communication process with the practice expected in that setting, including whether an accurate medication list and after-visit summary were available.
3. Causation: The patient would need to show that the communication or record defect caused the interruption and that the interruption caused the stroke. The note, audio if lawfully retained, portal messages, dispensing history, patient testimony, and clinical record could support or weaken that chain. An adverse event after an AI-assisted encounter does not itself prove that the AI, the clinician, or the institution caused it.
4. Damages and allocation: Permanent disability could create substantial damages, but no evidence presented here supports a percentage allocation or settlement range. Comparative-fault rules also vary by jurisdiction.
Lessons for Using Ambient Documentation AI:
1. Review the draft according to clinical risk
The signing workflow should not assume that the draft captured every important detail. A locally approved review checklist can include:
- All medication changes explicitly documented?
- High-risk medications (anticoagulants, insulin, immunosuppressants) confirmed?
- Clear instructions (start/stop/continue) for each medication?
- Diagnostic plan matches your intent?
- Follow-up clearly stated?
No universal number of minutes establishes an adequate review. Complexity, risk, product behavior, and the extent of human editing determine the work required.
2. Implement safety checks for high-risk medications
Illustrative high-risk medication template (adapt to local workflow):
Anticoagulation: - [ ] Continue [medication] [dose] [frequency] (no change) - [ ] OR: Change from [old] to [new] - [ ] OR: STOP [medication], start [new medication]
Insulin: - [ ] Current regimen confirmed - [ ] OR: Changes explicitly documented
3. Provide written medication instructions
An after-visit summary or equivalent written plan can distinguish:
- Updated medication list
- Medications to START (highlighted)
- Medications to STOP (highlighted)
- Medications to CONTINUE unchanged
The delivery format should reflect patient preference and access. Printing, portal delivery, interpreter-supported review, caregiver participation, and follow-up calls are options, not interchangeable mandates.
4. Use medication names, not categories
Avoid: “Stop the blood thinner.”
Prefer: “Stop the baby aspirin. Keep taking the apixaban.”
Avoid: “Increase the water pill.”
Prefer: “Increase the furosemide from 40 to 80 milligrams.”
5. Teach-back method for high-risk changes
One possible prompt is: “Tell me which medications you’re changing.”
Patient should state: - “I’m stopping the baby aspirin” - “I’m continuing the apixaban twice a day” - “I’m increasing the furosemide to 80 milligrams”
If the patient cannot teach back the plan, clarify it, provide the agreed written format, and consider follow-up appropriate to the clinical risk.
6. Follow the applicable transparency and documentation policy
Whether the chart, encounter notice, consent process, or another record should identify ambient technology depends on law, institutional policy, product configuration, and how recordings or transcripts are handled. A generic attestation does not prove that review was adequate. If an attestation is used, it should be truthful and product-neutral unless the product identity is operationally necessary.
7. Quality assurance audits
An audit plan can sample ordinary encounters and high-risk events, compare source conversation or clinician intent with the signed record when lawful and feasible, track medication discrepancies, and route recurring defects to clinical governance and the vendor. The sampling frequency, sample size, and action threshold should be prespecified from local volume, baseline error, and severity. A fixed ten-note sample or 5% threshold is not universally validated.
8. Design patient communication with legal and community input
A notice can explain what technology is used, whether audio is retained, who can access the data, how the note is reviewed, and how to ask questions or opt out when an alternative is available. Consent and notification requirements are jurisdiction-specific. Transparency language must describe the real data flow and review process, not merely reassure the patient that AI is present.
Scenario 3: Hypothetical Autonomous Retinal-Screening False Negative
A federally qualified health center serving a predominantly Latino, low-income population has implemented a fictional autonomous diabetic-retinopathy screening system called RetinaAccess. The patient event and retrospective image review are illustrative. The real IDx-DR pivotal evidence is discussed separately for comparison.
Background: - Your patient population has high diabetes prevalence (25%) - Ophthalmology access limited (6-month wait for appointments) - Many patients lack transportation to ophthalmology clinics - Goal: Screen diabetic patients for retinopathy in primary care setting
Equipment: A fictional non-mydriatic retinal camera and RetinaAccess software
Workflow: 1. Medical assistant obtains retinal photos (both eyes) 2. AI analyzes images 3. AI provides autonomous interpretation (no physician review of images) 4. Results: “Referable diabetic retinopathy detected” → Refer to ophthalmology OR “Negative for referable diabetic retinopathy” → Rescreen in 1 year
Patient: 52-year-old woman with type 2 diabetes × 12 years - HbA1c: 9.2% (poorly controlled) - Last eye exam: 3 years ago (“told everything was fine”) - No visual symptoms
RetinaAccess screening (performed by medical assistant): - Retinal photos obtained both eyes - Image quality: “Adequate” per AI assessment - Result: “Negative for referable diabetic retinopathy” - Recommendation: “Rescreen in 12 months”
Your assessment: Review AI result in chart, no physician review of actual images
Management: “Your diabetic eye screening looks good. We’ll recheck next year. Let’s focus on getting your blood sugar under better control.”
18 months later: Patient presents with vision changes - “Floaters in right eye, vision blurry” - Exam: Decreased visual acuity right eye (20/80)
Ophthalmology referral (expedited): - Diagnosis: Proliferative diabetic retinopathy (PDR), vitreous hemorrhage right eye - Findings: Neovascularization, dot-blot hemorrhages, hard exudates both eyes - Retrospective review of IDx-DR images from 18 months ago: 3/3 retinal specialists identify early retinopathy changes (microaneurysms, hard exudates) visible on the original images
Ophthalmologist’s note: “Findings suggest retinopathy was present 18 months ago and progressed due to lack of follow-up and poor glycemic control.”
Treatment: Panretinal photocoagulation (PRP), anti-VEGF injections - Outcome: Vision partially recovered (20/50 right eye) but permanent peripheral vision loss
Malpractice claim: Failure to diagnose diabetic retinopathy, reliance on AI false negative
Question 1: What went wrong with the AI screening?
Critical failures in AI implementation:
1. AI false negative
Real comparator evidence, not evidence about RetinaAccess: In the IDx-DR pivotal prospective study (Abràmoff et al., 2018): - Sensitivity: 87.2% (detects 87.2% of referable retinopathy) - Specificity: 90.7% - Among reference-positive participants represented in the sensitivity calculation, the complementary false-negative fraction was approximately 12.8%
That complement is not a universal statement that every deployment will miss one in eight patients. It depends on the study population, reference standard, image-acquisition success, operating threshold, and intended-use conditions. The pivotal study does not establish the performance of the fictional RetinaAccess system or prove why the fictional result was negative. Product-specific sensitivity must not be transferred to another model, version, camera, population, or workflow.
Why false negatives occur: - Image quality issues: Shadows, reflections, small pupils reduce AI accuracy - Early/subtle findings: Microaneurysms, small hard exudates harder for AI to detect - Population and site transportability: performance can differ when patient characteristics, cameras, operators, disease prevalence, or image-acquisition conditions differ from the evaluation setting - Poorly controlled diabetes: Rapid progression between screenings
2. “Autonomous” system with no physician oversight
IDx-DR was authorized through De Novo classification in 2018 for a bounded autonomous screening workflow described in DEN180001. “Autonomous” describes the labeled function. It does not mean that the software makes every eye-care decision or that its evidence transfers to RetinaAccess.
Problem: Your workflow followed the autonomous model: - Medical assistant obtains images - AI provides result - Physician accepts AI result without viewing images
No physician ever looked at the retinal photos.
Governance question: Did the clinic follow the exact intended-use workflow, ensure that image-acquisition operators were trained, manage ungradable results correctly, and define when symptoms or other clinical findings would prompt eye-care referral independent of the screening output? The pivotal workflow did not require a primary-care physician to reinterpret every image, so the absence of physician image review is not itself proof of misuse.
3. No risk stratification
Patient-specific context that should remain visible to the clinician: - Diabetes duration 12 years (longer duration = higher retinopathy risk) - HbA1c 9.2% (poor control is clinically relevant to retinopathy risk and progression) - No eye exam × 3 years (lack of baseline for comparison)
The exercise should ask whether symptoms, prior retinopathy, pregnancy, image quality, an overdue comprehensive examination, or another clinical factor requires a different pathway under current guidelines and the device label. It should not invent a universal rule that HbA1c or diabetes duration alone requires physician reinterpretation of every negative autonomous screen.
4. Over-reliance on AI, inadequate patient counseling
What you told the patient: “Your diabetic eye screening looks good.”
A more accurate explanation would distinguish screening from certainty: “The screening system did not detect referable diabetic retinopathy today. No screening test detects every case. Keep the recommended eye-care follow-up, and report new floaters, blurred vision, dark areas, or other vision changes promptly.” Whether additional disclosure or comprehensive examination is indicated depends on current guidance, the label, symptoms, and patient circumstances. The pivotal study’s complementary false-negative fraction should not be recited as though it were the local patient’s individualized probability.
Question 2: Are you liable for malpractice?
Plaintiff’s argument:
1. The clinic used a screening result outside a safe care pathway
The plaintiff would argue that the clinic treated a negative output as a complete eye-care conclusion, failed to reconcile it with the patient’s overdue eye care and clinical context, and communicated more certainty than a screening result justified. The retrospective specialist reviews would be important evidence, although hindsight review can introduce its own bias.
2. Product and workflow evidence were not established
RetinaAccess is fictional. The clinic would need product-specific evidence showing the model version, camera, operators, patient population, image-quality process, and threshold used. IDx-DR’s study and De Novo record cannot establish RetinaAccess performance. If RetinaAccess lacked appropriate authorization or was used outside its intended use, that would be a separate regulatory and governance problem.
3. Communication and follow-up were inadequate
The plaintiff could argue that “looks good” was materially misleading because it collapsed a probabilistic screening result into certainty. The patient should have received the actual result, its role in the care pathway, the planned follow-up, and symptom-triggered escalation instructions. Whether a particular consent disclosure was legally required would remain jurisdiction-specific.
4. Earlier detection could have changed treatment
The plaintiff would need expert evidence that the abnormalities were present and detectable under the contemporaneous standard, that the appropriate response would have produced earlier specialist evaluation, and that earlier treatment would probably have changed the outcome. The fictional retrospective review supports investigation of this causal chain, not automatic proof.
Defense’s argument:
1. The clinic followed an authorized autonomous workflow
If the actual product were an authorized autonomous system and the clinic followed its label, the defense could argue that routine primary-care reinterpretation was not part of the intended workflow. That fact would be relevant but not conclusive. FDA authorization establishes a lawful marketing pathway for a defined intended use; it does not decide civil liability or guarantee a clinical outcome.
2. A false negative does not necessarily mean malfunction or negligence
No screening system has perfect sensitivity. A case can be missed even when a model functions as designed. The defense would examine image quality, disease definition, contemporaneous findings, progression after the test, and whether the fictional system’s performance remained within its validated range.
3. Disease progression had multiple contributors
The course of diabetic retinopathy depends on disease severity, glycemic control, blood pressure, kidney disease, follow-up, and other factors. Poor control should not be converted into moral blame or an invented percentage of patient fault. It is medically relevant to progression and causation.
4. The counterfactual outcome is uncertain
Earlier evaluation can permit earlier treatment, but the record would need to establish what treatment was indicated at the earlier time and the probability that it would have prevented this degree of loss.
Answer 3: The vignette identifies serious questions, not a predetermined legal result
The strongest concern is not simply that a machine produced a false negative. It is that the fictional clinic may have lacked a product-specific validation record, precise patient communication, a pathway for symptoms and exceptions, postdeployment monitoring, and an incident-review process. Those facts would be examined separately from the device authorization and separately from clinical causation.
No percentage of fault, verdict, or settlement range can be derived responsibly from this hypothetical. The relevant legal rules vary by jurisdiction, and the evidentiary record would include the product label, training records, system logs, images, configuration history, clinical documentation, patient communications, and expert testimony.
Lessons for Autonomous Diabetic-Retinopathy Screening:
1. Preserve the bounded meaning of “autonomous”
Autonomous screening can produce a clinical output without specialist interpretation of every image when used within its authorized workflow. It does not replace comprehensive eye care, evaluate unrelated eye disease, override symptoms, or eliminate the responsibilities assigned to the ordering and implementing organization.
2. Verify the exact product and intended-use conditions
The implementation record should identify:
- Authorization number and current labeling
- Software version and compatible camera
- Eligible and excluded patient groups
- Operator training and image-acquisition requirements
- Handling of insufficient-quality or indeterminate outputs
- Referral and follow-up instructions attached to each result
3. Separate clinical risk from invented thresholds
HbA1c, diabetes duration, kidney disease, pregnancy, prior retinopathy, symptoms, and time since eye care can affect management. The clinic should map these factors to current professional guidance and the device label. The restored chapter’s fixed HbA1c, duration, six-month, and physician-review rules were illustrative rather than validated universal thresholds, so they should not be presented as standards.
4. Communicate the result without false reassurance
An accurate script can say:
“The screening system did not detect referable diabetic retinopathy today. This is a screening result, not a guarantee that every eye condition is absent. Follow the recommended eye-care schedule and report new floaters, blurred vision, dark areas, pain, or other vision changes promptly.”
The clinic should use the actual product’s patient-facing language where required and should not translate study sensitivity into a personalized certainty statement.
5. Document the real decision pathway
The record can include the named product and result, image-quality or gradability status, follow-up recommendation, symptom instructions, referral decision, and any reason the ordinary pathway was changed. Documentation should describe what occurred, not serve as a defensive attestation.
6. Monitor local performance and access outcomes
A postdeployment program should track acquisition failures, indeterminate results, referral completion, time to eye care, discordant findings, complaints, and detected safety events. Sampling and action thresholds should be based on local volume and risk rather than an unsupported rule such as reviewing 10 to 20 negatives each year or requiring sensitivity above 95%.
7. Investigate subgroup performance without inventing disparities
The clinic should examine whether performance or access outcomes differ across relevant patient groups. The absence of adequate subgroup evidence is a reason to measure and monitor, not a basis for asserting that the training population was predominantly White or assigning a disparity magnitude without a source.
8. Treat access and safety as joint endpoints
Autonomous screening may expand access where specialist capacity is limited. A successful program must still connect positive and ungradable results to care, prevent negative results from becoming false reassurance, and show that the pathway improves completed screening and appropriate follow-up. Technical accuracy alone does not establish that a screening program improves patient outcomes.