Large Language Models in Clinical Practice

Keywords

clinical LLM, clinical LLMs, evidence LLM, OpenEvidence evaluation, large language models in medicine, ER-Reason benchmark, GPT-4 medicine, Claude healthcare, medical AI chatbot, LLM hallucinations, HIPAA compliant AI, LLM reproducibility liability, UpToDate Expert AI, clinical decision support LLM

Large language models can draft, summarize, retrieve, and reorganize clinical text. They can also omit key facts, invent support, vary across repeated runs, and produce unsafe recommendations with confident language. Benchmark performance, simulated-case performance, clinician performance with AI, and patient outcomes are different endpoints and must not be treated as interchangeable.

Learning Objectives

After completing this chapter, clinicians should be able to:

  • Understand how LLMs work and their fundamental capabilities and limitations in medical contexts
  • Identify appropriate vs. inappropriate clinical use cases based on risk-benefit assessment
  • Recognize and mitigate hallucinations, citation fabrication, and knowledge cutoff problems
  • Navigate privacy (HIPAA), liability, and ethical considerations specific to LLM use in medicine
  • Evaluate medical-specific LLMs (Med-PaLM, GPT-4 medical applications) vs. general-purpose models
  • Implement LLMs safely in clinical workflows with proper oversight and verification protocols
  • Communicate transparently with patients about LLM-assisted care
  • Apply vendor evaluation frameworks before adopting LLM tools for clinical practice

The Clinical Context:

Large Language Models (ChatGPT, GPT-4, Med-PaLM, Claude) have exploded into medical practice since ChatGPT’s public release in November 2022. Unlike narrow diagnostic AI trained for single tasks, LLMs are general-purpose systems that can write clinical notes, answer medical questions, summarize literature, draft patient education materials, generate differential diagnoses, and assist with complex clinical reasoning.

They differ fundamentally from narrow diagnostic AI: LLMs communicate in natural language, appear to “understand” medical concepts, and can perform diverse tasks without task-specific training. This versatility makes them extraordinarily useful and extraordinarily dangerous if used incorrectly.

The fundamental challenge: LLMs are statistical language models trained to predict plausible next words, not to retrieve medical truth. They can generate confident, coherent, authoritative-sounding but completely false medical information (“hallucinations”). A physician who trusts LLM output without verification risks patient harm.

Key Applications:

  • Ambient clinical documentation: Nuance DAX and Abridge convert conversations into draft notes. Peer-reviewed evaluations report workflow benefits, but effects vary by product, setting, endpoint, and implementation (Rotenstein et al., 2026; Lukac et al., 2025)
  • Literature synthesis and summarization: Summarize guidelines, compare treatment options (with citation verification)
  • Patient education materials: Generate health literacy-appropriate explanations (with physician review)
  • Differential diagnosis brainstorming: Suggest possibilities for complex cases (treat as idea generation, not diagnosis)
  • Medical coding assistance: Suggest ICD-10/CPT codes from clinical narratives (with compliance review)
  • Clinical decision support: Glass Health, other LLM-based systems provide treatment suggestions (requires rigorous verification)
  • Medical education: Explaining concepts, generating practice questions (risk: teaching hallucinated “facts”)
  • Autonomous patient advice: Patients asking LLMs medical questions without physician oversight (dangerous false reassurance)
  • Medication dosing without verification: LLMs fabricate plausible but incorrect dosages
  • Citation generation: LLMs can fabricate references or attach real references to unsupported claims, so every citation requires direct verification

What Actually Works:

  1. Ambient documentation: Randomized and observational studies support reduced documentation burden for some implementations, not a universal percentage or guaranteed return on investment (Rotenstein et al., 2026; Lukac et al., 2025)
  2. Literature synthesis with source verification: LLMs can accelerate screening and drafting, but the synthesis is not reliable until each material claim and citation is checked against the underlying source
  3. Patient education draft generation: Language-level adaptation is useful when the clinician verifies medical content and individualizes the draft before distribution
  4. Care-transition communication: A randomized trial of the PreA chatbot measured consultation workflow in one setting. Its results should not be generalized to autonomous clinical decision-making (Tao et al., 2025)
  5. Clinician support in simulated care: A randomized trial involving 92 physicians found a 6.5 percentage-point increase in management-reasoning scores across five simulated vignettes. It did not measure patient outcomes (Goh et al., 2025)

What Does Not Work:

  1. Citation reliability: Error rates vary by model, prompt, domain, and task. A citation that exists may still fail to support the sentence, so no fixed cross-task fabrication rate is defensible
  2. Medication dosing: Unverified model output can use the wrong equation, extract the wrong patient parameter, or apply the wrong guideline
  3. Medical calculations: NIH research found GPT-4 achieves only 50.9% accuracy on clinical calculator tasks (CHADS-VASc, GFR, risk scores), with three failure modes: wrong equations, parameter extraction errors, arithmetic mistakes
  4. Autonomous diagnosis: LLMs lack patient-specific data, physical exam findings, cannot replace clinical judgment
  5. Real-time medical knowledge: All LLMs have training data cutoffs, meaning they may be unaware of newer drugs, guidelines, or treatments published after training
  6. AI manuscript associations: A cross-sectional submission analysis found associations between disclosed AI use and editorial outcomes. The observational design does not establish that AI use caused rejection or withdrawal (Perlis et al., 2026)

Critical Insights:

Unsupported output remains a system-level risk: Model training, prompting, retrieval, tool use, and workflow design can change the rate and type of error, but none removes the need for verification

HIPAA status is product- and workflow-specific: The organization must verify the exact service, account, configuration, retention and training terms, safeguards, and contractual relationship before protected health information is entered

Liability is fact- and jurisdiction-specific: A model output does not displace the duties of the clinician, institution, or manufacturer. FDA status, a BAA, or a vendor disclaimer does not by itself decide civil liability

Exam performance ≠ clinical utility: GPT-4 scores 86% on USMLE but multiple choice questions do not test clinical judgment, patient communication, or risk management. When answer patterns are disrupted (NOTA test), LLM accuracy drops 9-38% depending on the model, suggesting pattern matching over genuine reasoning (Bedi et al., 2025)

Collaboration depends on task and interface: One randomized vignette study found no significant improvement in diagnostic-reasoning score with GPT-4 access, while a separate randomized simulation found improved management reasoning. Neither trial established patient-outcome benefit (Goh et al., 2024; Goh et al., 2025)

Clinical LLM evidence remains thin relative to publication volume: A 2026 LLM-assisted review identified 4,609 peer-reviewed clinical LLM studies from January 2022 through September 2025, but only 19 prospective randomized trials. Most studies tested simulated scenarios or exam-style tasks rather than patient-centered outcomes (Chen et al., 2026).

Ambient documentation has the most mature workflow evidence: Benefits vary by product and workflow, and time saved is not the same as cash savings, clinical quality, or patient benefit.

Clinician-driven EHR monitoring is not a benchmark problem: ChatEHR’s Nature Medicine report states that benchmark-based evaluations are insufficient for monitoring clinician-driven EHR interactions (Shah et al., 2026).

Prompting quality matters enormously: Specific, detailed prompts with requests for sourcing and uncertainty yield better outputs than vague questions

Clinical Bottom Line:

LLMs are powerful assistants for documentation, education, and brainstorming, but dangerous if used autonomously for diagnosis, treatment, or urgent decisions.

Safe use requires: - Institutionally approved systems only after product-, account-, and workflow-specific privacy review - Always verify medical facts against authoritative sources - Treat LLM output as drafts requiring physician review, never final decisions - Document verification steps - Transparent communication with patients about LLM assistance

Demand evidence: - Ask vendors for prospective validation studies (not just retrospective accuracy) - Request HIPAA compliance documentation and Business Associate Agreement (BAA) - Validate locally before widespread deployment - Monitor continuously for errors, near-misses, and hallucinations

The evidence supports bounded assistance, not a universal clinical copilot. Match the claim to the study design and endpoint, verify the output, and monitor the deployed workflow.

Medico-Legal Considerations:

  • Liability allocation is not fixed: Duties and defenses depend on the jurisdiction, intended use, professional standard, workflow, warnings, contracts, and facts of the event
  • Standard of care is context-specific: Adoption alone does not establish a legal duty to use a model or excuse unsafe use
  • Reproducibility creates unique liability: LLMs produce different outputs for identical prompts, complicating documentation, peer review, and quality assurance
  • Documentation requirements: Note LLM assistance where material to decisions, document verification steps, record LLM version and timestamp
  • Testing before deployment: Repeat representative tasks enough to characterize consequential output variation, using a prespecified local test plan rather than an arbitrary run count
  • Informed consent emerging: Some institutions now inform patients when LLMs assist documentation or clinical reasoning
  • HIPAA consequences depend on the violation and enforcement framework: Verify current HHS rules rather than copying a fixed penalty range into deployment policy
  • Insurance terms vary: Obtain the carrier’s current written coverage position for the intended workflow
  • Fabricated citations compromise research integrity: Authors remain responsible for verifying references and complying with journal and institutional policies

Essential Reading:

  • Omiye JA et al. (2024). “Large Language Models in Medicine: The Potentials and Pitfalls: A Narrative Review.” Annals of Internal Medicine 177:210-220. (doi:10.7326/M23-2772) [Stanford thorough review covering LLM capabilities, limitations, bias, privacy concerns, and practical clinical applications]

  • Singhal K et al. (2023). “Large language models encode clinical knowledge.” Nature 620:172-180. [Med-PaLM 2 validation, 86.5% MedQA performance]

  • Thirunavukarasu AJ et al. (2023). “Large language models in medicine.” Nature Medicine 29:1930-1940. [Review of medical LLM capabilities and limitations]

  • Nori H et al. (2023). “Capabilities of GPT-4 on Medical Challenge Problems.” Microsoft Research. [GPT-4 USMLE performance: 86%+]

  • Ayers JW et al. (2023). “Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum.” JAMA Internal Medicine 183:589-596. [LLM vs. physician responses quality comparison]

  • Lee P et al. (2023). “Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine.” New England Journal of Medicine 388:1233-1239. [Clinical use cases and risk assessment]

For annual landscape overview:

See Also:


Introduction: How LLMs Differ from Narrow Medical AI

Every previous chapter in this handbook examines narrow AI: algorithms trained for single, specific tasks.

  • Radiology AI detects pneumonia on chest X-rays (and nothing else)
  • Pathology AI grades prostate cancer histology (and nothing else)
  • Cardiology AI interprets ECGs for arrhythmias (and nothing else)

Large language models are general-purpose systems that can perform diverse language tasks through a natural-language interface.

The same model can produce a guideline summary, draft a patient handout, or generate a differential diagnosis without task-specific retraining. The apparent versatility does not establish factual accuracy, clinical validity, or fitness for a particular workflow.

This versatility is unprecedented in medical AI. It’s also what makes LLMs uniquely dangerous.

A narrow diagnostic AI fails in predictable ways: - Pneumonia detection AI applied to chest X-ray might miss a pneumonia (false negative) or flag normal lungs as abnormal (false positive) - Failure modes are bounded by the task

LLMs fail in unbounded ways: - Fabricate drug dosages that look correct but cause overdoses - Invent medical “facts” that sound authoritative but are false - Generate fake citations to real journals (paper does not exist) - Provide confident answers to questions where uncertainty is appropriate - Contradict themselves across responses - Recommend treatments that were standard of care in training data but have been superseded

A safer clinical analogy: An LLM can produce fluent language from broad training data, but it does not have a clinician’s longitudinal knowledge of the patient, physical examination, professional duties, or accountability. Model access to patient-specific data depends on the deployed system, and access does not guarantee correct interpretation.


Part 1: How LLMs Work

The Technical Basics (Simplified)

Training: 1. Ingest massive text corpora (internet, books, journals, Wikipedia, Reddit, medical textbooks, PubMed abstracts) 2. Learn statistical patterns: “Given these words, what word typically comes next?” 3. Scale to billions of parameters (weights connecting neural network nodes) 4. Fine-tune with human feedback (reinforcement learning from human preferences)

Inference (when you use it): 1. You provide a prompt (“Generate a differential diagnosis for acute chest pain in a 45-year-old man”) 2. LLM predicts most likely next word based on learned patterns 3. Continues predicting words one-by-one until stopping criterion met 4. Returns generated text

Crucially: - A base model does not automatically verify generated claims against an authoritative database - Models can execute learned and tool-assisted reasoning procedures, but fluent intermediate steps are not proof that the conclusion is sound - They generate token sequences from statistical patterns conditioned on the prompt and system context - Truth and plausibility are not the same thing

Why Hallucinations Happen

Definition: LLM generates confident, coherent, plausible but factually incorrect text.

Mechanism: The training objective is “predict next plausible word,” not “retrieve correct fact.” As WHO’s 2025 guidance notes, LLMs have no conception of what they produce, only statistical patterns from training data (WHO, 2025). When uncertain, LLMs default to generating text that sounds correct rather than admitting uncertainty or refusing to answer.

Medical examples documented in literature:

  1. Fabricated drug dosages:
    • Prompt: “What is the pediatric dosing for amoxicillin?”
    • GPT-3.5 response: “20-40 mg/kg/day divided every 8 hours” (incorrect for many indications; standard is 25-50 mg/kg/day, some indications 80-90 mg/kg/day)
  2. Invented medical facts:
    • Prompt: “What are the contraindications to beta-blockers in heart failure?”
    • LLM includes “NYHA Class II heart failure” (false; beta-blockers are indicated, not contraindicated, in Class II HF)
  3. Fake citations:
    • Prompt: “Cite studies showing benefit of IV acetaminophen for postoperative pain”
    • GPT-4 generates: “Smith et al. (2019) in JAMA Surgery found 40% reduction in opioid use” (paper does not exist; authors, journal, year all fabricated but plausible)
  4. Outdated recommendations:
    • All LLMs have training data cutoffs (check the specific model’s documentation)
    • May recommend drugs withdrawn from market after training
    • Unaware of updated guidelines published post-training

Why this matters clinically: A physician who trusts LLM output without verification risks: - Incorrect medication dosing → patient harm - Reliance on outdated treatment → suboptimal care - Academic dishonesty from fabricated citations → career consequences

Mitigation strategies: - Always verify drug information against pharmacy databases (Lexicomp, Micromedex, UpToDate) - Cross-check medical facts with authoritative sources (guidelines, textbooks, PubMed) - Never trust LLM citations without looking up the actual papers - Use LLMs for drafts and idea generation, never final medical decisions - Higher stakes = more verification required

Retrieval-Augmented Generation (RAG) for Healthcare

A first comprehensive review of RAG for healthcare applications (Ng et al., NEJM AI, 2025) examines how retrieval-augmented generation addresses three core LLM limitations:

The Three Problems RAG Addresses:

Problem How RAG Helps
Outdated information Retrieves from current knowledge bases, bypassing training cutoff
Hallucinations Grounds responses in retrieved documents, enabling source verification
Reliance on public data Can query institutional guidelines, formularies, and proprietary sources

How RAG Works:

  1. Retrieval: When a query arrives, the system searches a curated knowledge base (guidelines, textbooks, institutional protocols)
  2. Augmentation: Retrieved passages are provided to the LLM as context
  3. Generation: LLM generates response grounded in retrieved documents, with source citations

Clinical Applications:

  • Guideline-grounded responses: RAG can pull from current treatment guidelines, reducing outdated recommendations
  • Institutional integration: Hospital-specific formularies and protocols as knowledge sources
  • Citation verification: Responses include source documents that users can verify
  • Pharmaceutical industry: Drug information queries grounded in package inserts and regulatory documents

Limitations:

  • RAG reduces but does not eliminate hallucinations. LLMs can still misinterpret or misstate retrieved content
  • Retrieval quality depends on knowledge base curation and query matching
  • Adds latency and infrastructure complexity compared to base LLMs
  • Retrieved context may be outdated if knowledge base is not maintained

Clinical Implication: When evaluating clinical LLM tools, ask whether they use RAG or similar grounding approaches. Systems that cite sources enable verification; those that do not require more skepticism.

State of Clinical AI 2026: LLM Diagnostic Performance

The inaugural State of Clinical AI Report from the Stanford-Harvard ARISE network provides a nuanced assessment of LLM diagnostic capabilities.

Impressive benchmark results:

Brodeur et al. add a peer-reviewed positive signal for text-based LLM clinical reasoning. In the Science version, an OpenAI o1-series model was evaluated across five clinical reasoning experiments plus a real emergency department second-opinion study. On 143 NEJM Clinicopathologic Conference cases, o1-preview included the correct diagnosis in its differential in 78.3% of cases and selected the correct next diagnostic test in 87.5%. In 79 emergency department cases, o1 identified the exact or very close diagnosis in 65.8% of initial triage cases, 69.6% after the emergency physician encounter, and 79.7% at admission or ICU transfer, exceeding two attending physician second-opinion baselines at each touchpoint (Brodeur et al., 2026).

ER-Reason adds a more workflow-aware benchmark for this question. The preprint uses de-identified emergency department data from 3,437 patients across 3,984 ER encounters, 25,174 longitudinal clinical notes, and 72 physician-authored rationales to test LLM performance across triage acuity, EHR review, treatment planning, final diagnosis, and disposition. Its central contribution is not a deployment claim. It is a more realistic test bed for whether LLMs can reproduce clinician reasoning across linked ED decisions, rather than answer isolated licensing-exam questions (Mehandru et al., 2025, preprint).

The reality check:

Performance depends heavily on how narrowly the problem is framed:

  • The Brodeur emergency department arm was a blinded second-opinion differential diagnosis task, not real-time triage, disposition, or treatment management; subgroup error patterns and patient outcome effects were not established (Brodeur et al., 2026; Hopkins & Cornelisse, 2026)
  • ER-Reason reported that stronger reasoning models still showed systematic workflow gaps: o3-mini compressed acuity predictions toward “Urgent,” underestimated discharges, overestimated admissions, and reached only 34.40% exact ICD-10 match accuracy for final diagnosis despite substantially higher category-level agreement (Mehandru et al., 2025, preprint)
  • When models had to ask follow-up questions, manage incomplete information, or revise decisions as new details emerged, performance dropped (Johri et al., Nature Medicine, 2025)
  • On tests measuring reasoning under uncertainty, AI systems performed closer to medical students than experienced physicians (McCoy et al., NEJM AI, 2025)
  • Models tended to commit strongly to answers even when clinical ambiguity was high
  • Accuracy dropped 9-38% when familiar answer patterns were disrupted, with reasoning models the most resilient (Bedi et al., JAMA Network Open, 2025)

Why this matters for clinical practice:

In everyday medicine, uncertainty is common. The gap between performance on fixed exam questions and performance in ambiguous real-world scenarios is substantial. The report concludes that much of what looks impressive in headline-grabbing studies may not hold up in clinical practice.

Clinical implication: Use LLMs for brainstorming and drafts where you can verify output, not for situations requiring judgment under uncertainty.

When structured data outperforms LLMs: For clinical prediction tasks on structured EHR data, LLMs may underperform traditional approaches. In antimicrobial resistance prediction for sepsis, LLMs analyzing clinical notes achieved AUROC 0.74 compared to 0.85 for deep learning on structured EHR data, with combined approaches offering no improvement (Hixon et al., 2026, conference abstract). LLMs excel at language tasks, not structured clinical prediction.

A Clinical LLM Evidence Ladder

The current evidence base spans different questions that should not be collapsed into one claim about whether an LLM “works.”

Clinician reasoning in simulation: In a randomized study involving 92 physicians and five simulated vignettes, access to an LLM increased the mean management-reasoning score by 6.5 percentage points. The study tested simulated decisions, not care delivery or patient outcomes (Goh et al., 2025). A positive simulation endpoint is evidence for that endpoint, not proof of clinical benefit.

Longitudinal simulated care: A virtual objective structured clinical examination compared an AI system and 21 primary care physicians across 100 simulated patients and three visits. The longitudinal design is more realistic than a single-answer benchmark, but the patients, encounters, and outcomes remained simulated (Liévin et al., 2026).

Pragmatic deployment with a null primary endpoint: A cluster-randomized primary-care trial assigned 103 clinician clusters and included 9,702 encounters. Fourteen-day treatment failure was 2.0% with AI consultation and 2.2% with usual care (adjusted odds ratio 0.77, 95% CI 0.55–1.08, P=.13). The primary endpoint was not statistically significant, so the trial does not establish reduced treatment failure. (Agweyu et al., 2026).

Patient-facing education: A randomized trial of 2,113 participants compared e-learning plus a chatbot with the chatbot alone. Because every participant had chatbot access, the result estimates the incremental effect of e-learning in that chatbot-mediated setting, not chatbot effectiveness versus usual care (Li et al., 2025).

Human interaction failure: In a randomized study of 1,298 participants, models tested alone identified relevant conditions in 94.9% of scenarios, but participants using the models identified relevant conditions in fewer than 34.5% and did not outperform the control condition on disposition. Standalone model accuracy did not predict the value of the human-model system. (Bean et al., 2026).

Triage stress testing: A simulated evaluation used 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions, producing 960 responses. Thirty-three of 64 emergency responses were undertriaged. This is a safety stress test, not an observed adverse-event rate among patients (Ramaswamy et al., 2026).

The defensible question is not whether an LLM is medically intelligent. It is whether a specified human-model workflow improves a prespecified endpoint in the intended population without unacceptable harm.


Part 2: Failure Modes and Hallucination Exercises

Case 1: A Hypothetical Composite Oncology Protocol

Scenario (hypothetical composite, not a reported patient event): A clinician asks a general-purpose LLM to reproduce a pediatric acute lymphoblastic leukemia consolidation protocol.

LLM response: Generated detailed protocol with drug names, dosages, timing that looked professionally formatted and authoritative.

The problem: - The model assembles plausible dose, timing, and duration elements from incompatible regimens - It omits protocol-specific eligibility, risk-group, supportive-care, monitoring, and dose-modification rules - Professional formatting makes the synthesis appear authoritative even though no controlling protocol was retrieved

If followed without verification: The output could cause severe underdosing, overdosing, omitted supportive care, or treatment delay. The exercise intentionally avoids reproducing a regimen because protocol selection and dosing must come from the controlling oncology protocol and pharmacy workflow.

Why it happened: LLM trained on general medical text, not specialized oncology protocols. Generated plausible-sounding but incorrect regimen by combining fragments from different contexts.

The lesson: Never use LLMs for medication dosing without rigorous verification against authoritative sources (protocol handbooks, institutional guidelines, pharmacy consultation).

Case 2: A Hypothetical Composite Triage Failure

Scenario (hypothetical composite, not a published adverse event): An emergency physician uses an LLM to brainstorm a differential diagnosis for a patient with sudden severe headache, photophobia, and neck stiffness.

LLM differential: 1. Migraine (most likely) 2. Tension headache 3. Sinusitis 4. Meningitis (mentioned fourth) 5. Subarachnoid hemorrhage (mentioned fifth)

The safety concern: The prompt contains features that require urgent evaluation for dangerous secondary causes. A language model’s ordering of common diagnoses is not a validated triage rule.

The problem: LLM ranked benign diagnoses (migraine, tension headache) above life-threatening emergencies (SAH, meningitis) despite classic “thunderclap headache + meningeal signs” presentation.

Why it happened: - Training data bias: Migraine is far more common than SAH in text corpora - LLMs predict based on frequency in training data, not clinical risk stratification - No understanding of “rule out worst-case-first” emergency medicine principle

The lesson: LLMs do not triage by clinical urgency or risk. Physician must apply clinical judgment to LLM suggestions.

Safe resolution: The clinician follows the applicable emergency pathway and uses patient-specific history, examination, testing, and consultation. The model output remains a brainstorming artifact, not the basis for excluding a time-sensitive diagnosis.

Case 3: A Hypothetical Composite Citation Failure

Scenario (hypothetical composite): Medical student submitted literature review using GPT-4 to generate citations supporting statements about hypertension management.

LLM-generated citations (examples): 1. “Johnson et al. (2020). ‘Intensive blood pressure control in elderly patients.’ New England Journal of Medicine 383:1825-1835.” 2. “Patel et al. (2019). ‘Renal outcomes with SGLT2 inhibitors in diabetic hypertension.’ Lancet 394:1119-1128.”

The problem: Neither paper exists. Authors, journals, years, page numbers all plausible but fabricated.

Discovery: Faculty advisor attempted to retrieve papers for detailed review. None found in PubMed, journal archives, or citation databases.

Potential consequences: - The assignment may fail institutional research-integrity or academic-conduct review - Unsupported claims can enter manuscripts, grants, teaching materials, or clinical policies - Corrective review consumes time and may require withdrawal or revision of the work

Why this matters: - False or misleading citations can violate grant, journal, institutional, or professional rules depending on context and intent - Published errors can require correction, withdrawal, or retraction - Unsupported citations in guidance can propagate misinformation into downstream decisions

The lesson: Never trust LLM-generated citations. Always verify papers exist and actually support the claims attributed to them.

AI Use in Medical Research and Publishing: What the Data Shows

Beyond individual case studies, large-scale empirical evidence reveals patterns in how physicians and researchers actually use AI for manuscript preparation, and what happens to those manuscripts.

JAMA Network Analysis (105,538 manuscripts, 2023-2025):

A cross-sectional study of manuscripts submitted to 13 JAMA Network journals found that 3.3% of authors disclosed AI use, rising from 1.71% of submissions at the start of the 27-month period to 5.97% at the end (Perlis et al., JAMA 2026).

How AI is actually being used in medical research:

  1. Language correction/refinement: 67.7% (most common use)
  2. Statistical model development: 7.3%
  3. Other data analysis: 6.3%
  4. Manuscript drafting: 5.5%
  5. Literature search/evaluation: 4.3%

The quality signal (critical finding):

Manuscripts that disclosed AI use were significantly more likely to be rejected or withdrawn: - Rejected manuscripts: 29% higher odds of AI disclosure (OR 1.29) - Withdrawn pre-review: 748% higher odds of AI disclosure (OR 8.48) - Compared to manuscripts accepted or asked to revise

What this means for physicians:

This association suggests either: - AI may degrade manuscript quality when used improperly - Authors are more likely to disclose AI when quality is already marginal - Potential bias against AI-disclosed manuscripts in peer review

The data do not prove causation, but the pattern is clear: AI disclosure correlates with lower acceptance rates.

Corpus-level writing signal: A corpus-level analysis of more than 15 million PubMed abstracts estimated that at least 13.5% of 2024 abstracts contained an excess-vocabulary signal associated with LLM-assisted processing (Kobak et al., Science Advances, 2025). This method cannot identify individual authorship or establish misconduct.

Additional patterns:

  • Authors from non-English speaking countries were 30% more likely to disclose AI use (OR 1.30)
  • Viewpoints and Letters to the Editor had higher AI disclosure rates than Original Investigations
  • JAMA Psychiatry showed highest AI disclosure among specialty journals

The transparency question:

These data reflect self-reported AI use. The true rate may be higher if some authors: - Use AI without corresponding author awareness - Choose not to disclose despite journal requirements - Use AI in ways they do not recognize as “AI use”

Clinical bottom line for research writing:

While AI tools can assist with language refinement (the most common use case), physicians should be aware that AI-disclosed manuscripts show lower acceptance rates in the current publishing environment. This underscores the importance of: - Using AI as a draft/refinement tool, not as a replacement for critical thinking - Verifying all AI-generated content rigorously - Being transparent about AI use per journal policies - Recognizing that peer reviewers may scrutinize AI-assisted work more carefully

### Case 4: The Medical Calculation Gap

The problem: Physicians routinely use clinical calculators (CHADS-VASc for stroke risk, Cockcroft-Gault for GFR, HEART score for chest pain triage, LDL calculations). These quantitative tools drive treatment decisions daily.

What the research shows: MedCalc-Bench, a benchmark from NIH researchers evaluating LLMs on 55 different medical calculator tasks across 1,000+ patient scenarios, found that the best-performing model (GPT-4 with one-shot prompting) achieved only 50.9% accuracy (Khandekar et al., NeurIPS 2024).

Three distinct failure modes:

  1. Knowledge errors (Type A): LLM does not know the correct equation or rule
    • Example: Asked to calculate CHADS-VASc score, assigns wrong points to criteria
    • Most common error in zero-shot prompting (over 50% of mistakes)
  2. Extraction errors (Type B): LLM extracts wrong parameters from patient note
    • Example: Misidentifies patient age, medication history, or lab values from clinical narrative
    • 16-31% of errors depending on model
  3. Computation errors (Type C): LLM performs arithmetic incorrectly
    • Example: Calculates LDL as 142 mg/dL when correct answer is 128 mg/dL
    • 13-17% of errors even when equation and parameters are correct

Why this matters clinically:

Medical calculations drive treatment decisions: - CHADS-VASc ≥2 → anticoagulation for atrial fibrillation - eGFR <30 → medication dose adjustments - HEART score ≥4 → admission vs. discharge decision

50.9% accuracy means LLMs are flipping a coin on tasks with direct treatment implications.

The performance gap:

Anthropic reported Claude Opus 4.5 achieves 98.1% accuracy on MedCalc-Bench. This represents substantial improvement IF independently validated. Key caveats: - Anthropic’s metric is company-reported, not peer-reviewed - Original research (GPT-4): 50.9% accuracy - 98.1% claim requires independent replication

The lesson: Never trust LLM-generated medical calculations without verification. Check all risk scores, GFR calculations, and dosing adjustments against established calculators (MDCalc, online tools, pharmacy databases).

Clinical workflow: 1. LLM suggests calculation (e.g., “Patient’s CHADS-VASc score is 4”) 2. Verify independently: Use MDCalc or manual calculation 3. If mismatch: Trust the verified calculation, not the LLM 4. Document verification in clinical note


Part 3: Ambient Clinical Documentation

Nuance DAX: Ambient Documentation AI

The problem DAX solves: Physicians spend 2+ hours per day on documentation, often completing notes after-hours. EHR documentation contributes significantly to burnout.

How DAX works: 1. Physician wears microphone during patient encounter 2. DAX records conversation (with patient consent) 3. LLM transcribes speech → converts to structured clinical note 4. Note appears in EHR for physician review/editing 5. Physician reviews, makes corrections, signs note

Evidence base:

Regulatory status: Regulatory status must be assessed for the exact functions and claims of the deployed product. A documentation function may fall outside the device definition, while a diagnostic or treatment function can raise a different question. Do not infer FDA status for an entire platform from one documentation workflow.

Clinical validation: Ambient-scribe evidence includes randomized and observational evaluations with different products, settings, outcomes, and denominators. Some studies report reduced documentation burden and favorable user experience, while others show smaller or heterogeneous effects. The documentation chapter maps those findings to their specific designs and endpoints (Rotenstein et al., 2026; Lukac et al., 2025).

Real-world deployment: - Deployment counts, retention figures, and product capabilities change rapidly - Vendor-reported adoption is evidence of use, not independent evidence of accuracy, safety, or outcome benefit - Contract-specific configuration can differ across health systems

Cost-benefit: - Use the current written contract, implementation costs, support requirements, adoption rate, and measured local time change - Do not convert repurposed clinician time into cash savings unless the workflow actually changes staffing, throughput, or compensated activity - Report sensitivity analyses rather than a universal payback period

Why this works: - Well-defined task (transcription + note structuring) - Physician review catches errors before note finalization - Integration with EHR workflow - Patient consent obtained upfront - Privacy and security controls can be configured within a covered-entity workflow, subject to organization-specific review and contracts

Limitations: - Requires patient consent (some decline) - Poor audio quality → transcription errors - Complex cases with multiple topics may require substantial editing - Subscription cost barrier for small practices

Abridge: AI-Powered Medical Conversations

Another ambient documentation platform evaluated across multiple clinical settings: - Product- and setting-specific studies report documentation and workload effects, but no single time-savings percentage represents all implementations - Focuses on primary care and specialty clinics - Generates patient-facing visit summaries automatically

The lesson: When LLMs are used for well-defined tasks with effective review and workflow integration, they can deliver value. Human review is a control whose effectiveness must be measured, not a label that makes a system safe.

Emerging: LLM-Based Clinical Copilots

Beyond documentation, an LLM copilot was integrated into Penda Health’s EHR workflow for clinician review. Earlier preprint and observational analyses are useful for hypothesis generation, but they do not establish causal patient benefit and must not be transferred between product versions. A subsequent pragmatic cluster-randomized trial included 103 clinician clusters and 9,702 encounters; the primary 14-day treatment-failure endpoint was not significantly different (2.0% versus 2.2%, adjusted odds ratio 0.77, 95% CI 0.55–1.08, P=.13) (Agweyu et al., 2026). See Primary Care AI for the product-version and endpoint boundaries.


Part 4: Appropriate and Inappropriate Clinical Use Cases

SAFE Uses (With Physician Oversight)

1. Clinical Documentation Assistance

Use cases: - Draft progress notes from dictation - Generate discharge summaries - Suggest ICD-10/CPT codes - Create procedure notes

Workflow: 1. Physician provides input (dictation, conversation recording, bullet points) 2. LLM generates structured note 3. Physician reviews every detail, edits errors, adds clinical judgment 4. Physician signs final note

The Automation Bias Trap

The safety concern: Automation bias and complacency can reduce scrutiny of machine-generated content. Goddard et al. reviewed factors associated with automation bias in decision-support use; the paper does not establish a universal three-month transition to rubber-stamping (Goddard et al., 2012).

A hypothetical erosion pattern: - Initial use: The clinician reviews every section and logs errors - Growing familiarity: The clinician begins to skim apparently reliable sections - Routine use: Signing becomes habitual and omissions are less likely to be noticed - Mature deployment: Undetected error can persist unless the organization measures review behavior and samples outputs

Counter-measures to maintain vigilance:

  1. Risk-based review: Verify every element that can materially affect diagnosis, treatment, coding, communication, or continuity of care
  2. Interface support: Make source text and generated claims easy to compare, especially medications, dates, laboratory values, and plans
  3. Red flag awareness: Know the AI’s failure modes (medication names, dosing, dates, rare conditions)
  4. Independent sampling: Audit a prespecified sample of generated notes based on use volume and risk, then adjust the sampling plan from observed performance
  5. Error tracking: Log corrected errors and near misses; a sudden fall in detected errors can reflect improvement, reporting failure, or reduced review

“Human in the loop” describes workflow position, not control effectiveness. The review must be feasible, observable, and supported by an interface that makes consequential errors detectable.

Risk mitigation: - The signing practitioner must verify that the note adequately documents the care provided, including when AI captures transcription (CMS Transmittal 12897) - Review should be designed to detect hallucinations, errors, and omissions rather than treated as a guarantee that every error will be caught - Use only institutionally approved configurations after verifying data handling, business-associate obligations where applicable, access controls, and retention - Civil liability remains fact- and jurisdiction-specific; a signature requirement does not determine every participant’s potential liability (Mello and Guha, 2024)

Evidence: Effects on documentation time and burden vary across studies and implementations (see the evidence discussion above)

2. Literature Synthesis and Summarization

Use cases: - Summarize clinical guidelines - Compare treatment options from multiple sources - Generate literature review outlines - Identify relevant studies for research questions

Workflow: 1. Provide LLM with specific question and context 2. Request summary with citations 3. Verify all citations exist and support claims 4. Cross-check medical facts against primary sources

Example prompt:

"Summarize the 2023 AHA/ACC guidelines for management
of atrial fibrillation, focusing on anticoagulation
recommendations for patients with CHADS-VASc ≥2.
Include specific drug dosing and monitoring requirements.
Cite specific guideline sections."

Risk mitigation: - Verify citations before relying on summary - Cross-check facts with original guidelines - Use as starting point, not final analysis

3. Patient Education Materials

Use cases: - Explain diagnoses in health literacy-appropriate language - Create discharge instructions - Draft procedure consent explanations - Translate medical jargon to plain language

Workflow: 1. Specify reading level, key concepts, patient concerns 2. LLM generates draft 3. Physician reviews for medical accuracy 4. Edits for cultural sensitivity, individual patient factors 5. Shares with patient

Example prompt:

"Create a patient handout about type 2 diabetes management
for a patient with 6th grade reading level. Cover: medication
adherence, blood sugar monitoring, dietary changes, exercise.
Use simple language, avoid jargon, 1-page limit."

Risk mitigation: - Fact-check all medical information - Customize to individual patient (LLM generates generic content) - Consider health literacy, cultural factors

4. Differential Diagnosis Brainstorming

Use cases: - Generate possibilities for complex cases - Identify rare diagnoses to consider - Broaden differential when stuck

Workflow: 1. Provide detailed clinical vignette 2. Request differential with reasoning 3. Treat as idea generation, not diagnosis 4. Pursue appropriate diagnostic workup based on clinical judgment

Example prompt:

"Generate differential diagnosis for 45-year-old woman
with 3 months of progressive dyspnea, dry cough, and
fatigue. Exam: fine bibasilar crackles, no wheezing.
CXR: reticular infiltrates. Consider both common and
rare etiologies. Provide likelihood and key diagnostic
tests for each."

Risk mitigation: - LLM differential is brainstorming, not diagnosis - Verify each possibility clinically plausible for patient - Pursue workup based on pretest probability, not LLM ranking

5. Medical Coding Assistance

Use cases: - Suggest ICD-10/CPT codes from clinical notes - Identify documentation gaps for proper coding - Check code appropriateness

Workflow: 1. LLM analyzes clinical note 2. Suggests codes with reasoning 3. Coding specialist or physician reviews 4. Confirms codes match care delivered and documentation

Risk mitigation: - Compliance review essential (fraudulent coding = federal offense) - Physician confirms codes represent actual care - Regular audits of LLM-suggested codes

DANGEROUS Uses (Do NOT Do)

1. Autonomous Patient Advice

Why dangerous: - Patients ask LLMs medical questions without physician involvement - LLMs provide confident answers regardless of accuracy - Patients may delay appropriate care based on false reassurance

Hypothetical failure pathway: - A patient with chest pain asks a chatbot whether the symptoms are benign - The response emphasizes a common nonurgent cause without reliable emergency escalation - The patient interprets the response as individualized triage and delays evaluation

The lesson: Patients will use LLMs for medical advice regardless of physician recommendations. Educate patients about limitations, encourage them to contact you rather than rely on AI.

Major Health AI Product Launches

Both OpenAI and Anthropic launched dedicated healthcare products in January 2026:

2. Medication Dosing Without Verification

Why dangerous: - LLMs fabricate plausible but incorrect dosages - Pediatric dosing especially error-prone - Drug interaction checking unreliable

Hypothetical near-miss: - A clinician asks a general-purpose model for dosing in impaired renal function - The response applies assumptions for normal renal function or extracts the wrong patient parameter - An independent pharmacy check identifies the mismatch before administration

The lesson: Never use LLM-generated medication dosing without verification against pharmacy databases, dose calculators, or pharmacist consultation.

Medical calculations beyond dosing:

The quantitative reasoning gap extends beyond medication dosing to all medical calculators:

  • Risk scores: CHADS-VASc, HEART score, Caprini VTE risk
  • GFR calculations: Cockcroft-Gault, MDRD equations
  • Lab-derived values: LDL calculation, anion gap
  • Clinical indices: Pneumonia severity index, Wells’ criteria

NIH research found LLMs achieve only 50.9% accuracy on medical calculation tasks, with three failure patterns: wrong equations, parameter extraction errors, and arithmetic mistakes (Khandekar et al., NeurIPS 2024).

The lesson: Verify all LLM calculations against established medical calculators (MDCalc, institutional tools, pharmacy databases). See Case 4: The Medical Calculation Gap for detailed failure modes.

3. Urgent or Emergent Clinical Decisions

Why dangerous: - Time pressure precludes adequate verification - High stakes magnify consequence of errors - Clinical judgment + experience > LLM statistical patterns

The lesson: In emergencies, rely on clinical protocols, expert consultation, established guidelines, not LLM brainstorming.

4. Generating Citations Without Verification

Why dangerous: - Citation-error rates vary widely by model and task, and real references can still be misapplied - Using fake references = academic dishonesty, research misconduct - Propagates misinformation if not caught

The lesson: Never include LLM-generated citations in manuscripts, grants, presentations without verifying papers exist and support the claims.


Part 5: Prompting Techniques and Evidence-Based Approaches

Prompt structure can change model output, but the direction and magnitude depend on the model, task, comparator, metric, and evaluation design. A scoping review of 114 prompt-engineering studies describes heterogeneous techniques and outcomes rather than one clinically validated prompting rule (Zaghir et al., 2024). Prompting can improve an evaluated task without establishing clinical safety.

Core Prompting Paradigms

Zero-Shot Prompting

The simplest approach: ask a question without examples.

"What are the first-line treatments for community-acquired pneumonia
in an otherwise healthy adult?"

When to use: Simple factual questions, initial exploration, low-stakes queries

Limitations: Less reliable for complex reasoning, nuanced clinical scenarios, or specialized domains

Few-Shot Prompting

Provide examples of desired input-output pairs before your actual question.

Example 1:
Patient: 65-year-old male, chest pain radiating to left arm, diaphoresis
Assessment: High concern for ACS, recommend immediate ECG and troponins

Example 2:
Patient: 28-year-old female, sharp chest pain worse with inspiration
Assessment: Consider pleurisy, PE, or musculoskeletal cause

Now assess:
Patient: 72-year-old female with diabetes, fatigue and jaw pain for 2 days

When to use: When you need consistent output format, domain-specific reasoning patterns, or specialized terminology

Evidence: LLMs enhanced with clinical practice guidelines via few-shot prompting showed improved performance across GPT-4, GPT-3.5 Turbo, LLaMA, and PaLM 2 compared to zero-shot baselines (Oniani et al., 2024)

Chain-of-Thought (CoT) Prompting

Request step-by-step reasoning rather than direct answers.

"A 58-year-old man presents with progressive dyspnea and bilateral
leg edema. EF is 35%. Think through this step-by-step:
1) What are the key clinical findings?
2) What is the most likely primary diagnosis?
3) What additional workup is needed?
4) What are the initial management priorities?"

When to use: Complex diagnostic reasoning, treatment planning, cases with multiple interacting factors

Evidence: Chain-of-thought prompting can change performance and produce a visible rationale, but a fluent rationale is not a faithful explanation of the model’s internal process. Evaluate the answer and the cited evidence, not the persuasiveness of the reasoning text (Savage et al., 2024).

Limits of Chain-of-Thought Prompting

CoT is not universally beneficial. Tasks requiring implicit pattern recognition, exception handling, or subtle statistical learning may show reduced performance with CoT prompting. An NEJM AI study found that reasoning-optimized models showed overconfidence and premature commitment to incorrect hypotheses in clinical scenarios requiring flexibility under uncertainty (NEJM AI, 2025).

Practical implication: Treat chain-of-thought as a task-specific prompting option, not a safety mechanism. Compare it with simpler prompts on the intended task, and do not use prompt style as a substitute for a validated triage workflow.

Structured Clinical Reasoning Prompts

Organize clinical information into predefined categories before requesting analysis.

PATIENT INFORMATION:
- Age/Sex: 45-year-old female
- Chief Complaint: Progressive fatigue x 3 months

HISTORY:
- Duration: 3 months, gradual onset
- Associated: Weight gain, cold intolerance, constipation
- PMH: Type 2 diabetes, hypertension

PHYSICAL EXAM:
- VS: BP 142/88, HR 58, afebrile
- General: Appears fatigued, dry skin, periorbital edema

LABS:
- TSH: 12.4 mIU/L (0.4-4.0)
- Free T4: 0.6 ng/dL (0.8-1.8)

Based on this structured information, provide:
1. Primary diagnosis with reasoning
2. Differential diagnoses to consider
3. Recommended next steps

Evidence: Structured templates that organize clinical information before diagnosis improve LLM diagnostic capabilities compared to unstructured narratives (Sonoda et al., 2024)

Practical Prompting Framework (R-C-T-C-F)

For clinical prompts, include these components:

Component Description Example
Role Define the LLM’s expertise level “You are an internal medicine attending…”
Context Provide relevant background “…reviewing a case for morning report…”
Task Specify exactly what you need “…generate a differential diagnosis…”
Constraints Set boundaries and requirements “…focusing on reversible causes, avoiding rare conditions…”
Format Specify output structure “…as a numbered list with likelihood estimates.”

Poor prompt:

"What's wrong with this patient?"

Effective prompt:

"You are an internal medicine attending reviewing a case for
teaching purposes. A 55-year-old woman presents with fatigue,
unintentional weight loss of 15 lbs over 3 months, and new-onset
diabetes. Generate a differential diagnosis focusing on malignancy
and endocrine causes. Format as a numbered list with brief
reasoning for each, ordered by likelihood."

What the Evidence Shows

Technique Best Use Case Evidence Quality Key Citation
Zero-shot Simple queries, exploration Moderate Baseline in most studies
Few-shot Consistent formatting, specialized domains Task-specific evidence Oniani et al., 2024
Chain-of-thought Complex reasoning, teaching Task-specific evidence with caveats Savage et al., 2024
Structured templates Diagnostic workups Moderate Sonoda et al., 2024

Common Prompting Mistakes

  1. Vague requests: “Analyze this” vs. “Calculate the CHADS-VASc score and recommend anticoagulation”
  2. Missing context: Asking about drug dosing without patient weight, renal function, or indication
  3. Overloading: Combining multiple complex tasks in one prompt (ask sequentially instead)
  4. Assuming knowledge: LLMs may not know your institution’s specific protocols or formulary
  5. Skipping verification: Even excellent prompts produce outputs requiring clinical validation

Further Reading

  • Meskó, 2023: Tutorial on prompt engineering for medical professionals (JMIR)
  • Zaghir et al., 2024: Scoping review of 114 prompt engineering studies (JMIR)

Part 6: Privacy, HIPAA, and Legal Considerations

The HIPAA Problem

HIPAA status cannot be assigned to a model name in the abstract.

What must be verified: - The exact product, account type, tenant, configuration, data flow, retention, secondary uses, security controls, and subcontractors - Whether the vendor is acting as a business associate for the intended covered-entity function and will execute a BAA when required - Whether institutional approval permits the data and use case - Whether the workflow satisfies HIPAA and any additional federal, state, contractual, research, or institutional requirements

Consequences of noncompliance: Civil, criminal, contractual, licensing, employment, and institutional consequences depend on the facts and controlling law. Monetary amounts change and should be checked against current official HHS material rather than copied from a static vendor checklist.

What NOT to enter into public ChatGPT: - Patient names, MRNs, DOB, addresses - Detailed clinical vignettes with rare diagnoses (re-identification possible) - Protected health information of any kind

Products and deployment routes requiring workflow-specific review:

  1. OpenAI for Healthcare (ChatGPT for Healthcare)
    • Enterprise product with a vendor-described healthcare offering and BAA pathway (OpenAI product page)
    • Model access, retrieval, retention, encryption, and contract terms must be verified for the purchased service
    • See Enterprise AI Platforms for details
    • On September 1, 2026, OpenAI announced that ChatGPT for Healthcare can connect supported Epic environments: an administrator configures the organization’s Epic FHIR endpoint and OAuth client, and each clinician signs in with their own Epic account. The vendor describes read-only access to authorized chart context, summarization that points back to the record, and an optional ChatGPT surface in supported EHR layouts (OpenAI, September 2026; OpenAI Help Center). The connector does not write back to the chart and is not Stanford ChatEHR. Trade reporting states that the integration is read-only and does not write data into the record (Mehta, TechCrunch, September 2026). ChatEHR is a separate institutional deployment (Shah et al., 2026); see Part 11.
    • A separate Healthcare Public Data plugin searches nine official public sources: PubMed, ClinicalTrials.gov, DailyMed, RxNorm, openFDA, CMS Coverage, CMS Open Data, Medicare Care Compare, and the NPI Registry. Those apps are read-only and do not retrieve patient-chart data. The Epic plugin is not available on individual ChatGPT for Clinicians accounts (OpenAI, ChatGPT for Clinicians; OpenAI Help Center).
    • OpenAI reported that physicians rated 99.1% of responses “safe” across 4,363 ratings and 27 connected-EHR use cases. That vendor evaluation is not independent, peer-reviewed validation (OpenAI, September 2026).
  2. ChatGPT Health
    • Dedicated health conversation space launched January 2026
    • Medical record integration (FHIR), wellness app connections (Apple Health, Function Health, Peloton)
    • Consumer health functionality does not itself establish a covered entity’s compliant workflow
    • Current evidence status: No peer-reviewed studies on clinical effectiveness; company-reported metrics only
    • See AI Tools Every Physician Should Know for detailed coverage
  3. Claude for Healthcare
    • Launched January 11, 2026 at JPM Healthcare Conference
    • The vendor describes healthcare deployment through cloud and enterprise routes; BAA availability and data-use terms are plan- and contract-specific
    • Enterprise features: native ICD-10, NPI Registry, PubMed integrations; FHIR development; prior authorization workflows
    • Consumer features (US Pro/Max): Apple Health, HealthEx medical record sync; test result explanations; appointment prep
    • Named partners: Banner Health, Novo Nordisk, Sanofi, AstraZeneca, Flatiron Health, Veeva
    • Current evidence status: No peer-reviewed clinical validation; company-reported partner list
  4. Azure OpenAI Service
    • GPT-4/GPT-5 via Microsoft Azure
    • BAA available for healthcare customers
    • Data not used for training
    • Cost: API fees (usage-based)
  5. Google Cloud Vertex AI
    • Med-PaLM 2, PaLM 2
    • BAA for healthcare
    • Enterprise controls
    • Cost: Enterprise licensing
  6. Epic Integrated LLMs
    • Built into EHR workflow
    • EHR integration can support institutional controls but does not make every configuration or downstream use compliant by design
    • Deployment accelerating 2024-2026
  7. Vendor-specific medical LLMs
    • Nuance DAX, Abridge, Glass Health
    • BAA with healthcare systems
    • Subscription models

Safe practices: - Use only institutionally approved systems after product- and workflow-specific privacy review - De-identify cases before entering into public LLMs (but de-identification is imperfect) - Institutional approval before LLM deployment - Document patient consent where appropriate

Medical Liability Landscape

Current legal framework (evolving):

Professional accountability remains: - An LLM is not a licensed practitioner - Clinicians retain duties defined by the professional standard and the facts of the case - Institutions and manufacturers can also have duties; allocation is fact- and jurisdiction-specific - A model recommendation does not independently establish reasonable clinical care

Standard of care questions: 1. Is physician negligent for NOT using available LLM tools? - Currently: No clear standard - Future: May become expected for documentation efficiency

  1. Can unsafe LLM use create liability?
    • Potentially, depending on the duty, breach, causation, harm, jurisdiction, product role, warnings, and institutional workflow
    • FDA authorization, a BAA, a consent form, or a vendor disclaimer does not decide malpractice liability by itself

The reproducibility problem:

LLMs produce different outputs for identical prompts, creating unique liability challenges. Unlike traditional medical software that produces deterministic results (same input always yields same output), LLMs use probabilistic sampling, meaning the same clinical question asked twice may generate different recommendations.

Documentation implications (Maddox et al., 2025):

  • If LLM-generated clinical note varies between runs, which version becomes the legal record?
  • Peer review of LLM-assisted decisions becomes difficult when outputs are not reproducible
  • Quality assurance audits cannot validate LLM recommendations after the fact if the system produces different outputs when tested

Defensive documentation strategies:

  1. Follow institutional policy on whether the model, version, and timestamp should be recorded
  2. Preserve material outputs and provenance when required for safety, quality, research, or legal review
  3. Document clinically relevant reasoning and verification without adding boilerplate that obscures the record
  4. Retain interaction logs only under approved privacy, security, retention, and records-management rules

Malpractice insurance: - Check policy coverage for AI-assisted care - Coverage terms vary by carrier, policy, endorsement, and use case - Ask for the carrier’s current written position on the intended workflow (Missouri Medicine, 2025) - Follow policy notice requirements and involve institutional risk management before deployment

Testing LLM consistency before deployment:

Before adopting any LLM tool for clinical use, test reproducibility:

  1. Select representative tasks and edge cases based on expected volume, risk, patient mix, and failure consequences
  2. Repeat inputs enough to characterize consequential variation, using a prespecified plan tied to the model configuration
  3. Assess variance: Do outputs differ substantively or only stylistically?
  4. Document acceptance criteria: Define material factual, omission, provenance, and workflow failures before testing
  5. Red flags: If the same prompt yields contradictory recommendations (e.g., “start beta-blocker” vs. “beta-blockers contraindicated”), do not deploy without vendor explanation

For medication dosing, diagnostic recommendations, or high-stakes decisions, reproducibility testing is essential before clinical deployment.

Risk mitigation: - Use only institutionally approved systems evaluated for the intended use - Always verify LLM outputs - Maintain human oversight for all decisions - Document verification - Obtain consent where appropriate - Monitor for errors continuously - Test reproducibility before deployment (see Physician AI Liability and Regulatory Compliance for detailed liability framework)

System Prompts as Clinical Policy

Beyond user prompts (how you ask questions), system prompts define an LLM’s underlying persona and behavioral parameters. These are typically set by vendors or IT departments, not end users, but they significantly affect clinical outputs.

Why this matters for physicians:

A Mount Sinai study tested 20 LLMs across 5 million clinical decisions (ED vignettes and discharge summaries) and found that assigning different “physician personas” (ethical orientation crossed with cognitive style) shifted affirmative action rates from 36.9% to 46.4% under identical clinical evidence. This 9.5 percentage-point swing represents a substantial change in treatment recommendations, autonomy decisions, and resource utilization, without any change in the underlying clinical facts (Klang et al., 2026, preprint).

Practical implications:

  1. System prompts function as policy settings. The same LLM can behave conservatively or liberally depending on how its persona is configured at deployment.

  2. Organizations deploying clinical LLMs should:

    • Document the system prompt configuration
    • Version control system prompts like any clinical policy
    • Audit how persona settings affect outputs in their specific use cases
    • Involve clinical leadership in persona selection decisions
  3. Individual physicians should ask:

    • “What system prompt or persona is configured for this tool?”
    • “Has the organization tested how different configurations affect clinical recommendations?”
    • “Who approved the current configuration, and when was it last reviewed?”

The governance gap: Most current LLM deployment guidance focuses on data privacy and verification workflows. System prompt configuration, which can shift clinical action thresholds by nearly 10 percentage points, receives less attention but may require equivalent governance oversight.


Part 7: Vendor Evaluation Framework

Before Adopting an LLM Tool for Clinical Practice

Questions to ask vendors:

  1. “Is this system HIPAA-compliant? Can you provide a Business Associate Agreement?”
    • Essential for any system touching patient data
    • No BAA = no patient data entry
  2. “What is the LLM training data cutoff date?”
    • Cutoff dates vary by model and version (check vendor documentation)
    • Older cutoff = more outdated medical knowledge
    • Models with web search can access current information but still require verification
  3. “What peer-reviewed validation studies support clinical use?”
    • Demand JAMA, NEJM, Nature Medicine publications
    • User satisfaction ≠ clinical validation
    • Ask for prospective studies, not just retrospective benchmarks
    • Third-party benchmark reference: The MedHELM framework (35 benchmarks, 121 clinical tasks) provides vendor-neutral model comparisons; reasoning models (DeepSeek R1, o3-mini) score highest overall, but all models score in the 0.56-0.72 range on Clinical Decision Support, the category where most LLM clinical tools operate (Bedi et al., Nature Medicine, 2025)
  4. “How did you define and measure unsupported output for this use case?”
    • Require a task-specific error taxonomy, denominator, adjudication method, uncertainty estimate, and representative test set
    • Rates cannot be transferred across summarization, question answering, reference generation, retrieval, or clinical recommendation tasks
    • Published studies use different definitions and evaluation units. Report each result with its source design rather than combining them into one platform-wide “hallucination rate” (Asgari et al., 2025; Chelli et al., 2024)
    • Stanford’s ChatEHR figures come from one institutional Nature Medicine deployment report; usage, error, and value estimates remain single-institution modeled findings from the companion technical report, not a universal benchmark (Shah et al., 2026; Shah et al., 2026, technical report)
  5. “How does the system handle uncertainty?”
    • Good LLMs express appropriate uncertainty (“I’m not certain, but…”)
    • Bad LLMs confidently hallucinate when uncertain
  6. “What verification/oversight mechanisms are built into the workflow?”
    • Best systems require physician review before acting on LLM output
    • Dangerous systems allow autonomous LLM actions
  7. “How does this integrate with our EHR?”
    • Practical integration essential for adoption
    • Clunky workarounds fail
  8. “What is the cost structure and ROI evidence?”
    • Subscription per physician? API usage fees?
    • Request time-savings data, physician satisfaction metrics
  9. “What testing validates consistency of outputs across multiple runs?”
    • Ask for reproducibility data: same input, how often does output differ?
    • Critical for clinical decisions where consistency matters (dosing, treatment recommendations)
    • If vendor has not tested, they have not validated for clinical use
  10. “How does the organization’s insurance program address this use?”
    • Coverage cannot be inferred from general statements about AI.
    • Ask the insurer and institutional risk team directly rather than relying on vendor assurances
    • Request coverage confirmation in writing before deployment
  11. “Who is liable if LLM output causes patient harm?”
    • Most vendors disclaim liability in contracts
    • Physician/institution bears risk
  12. “What data is retained, and can patients opt out?”
    • Data retention policies
    • Patient consent/opt-out mechanisms

Red Flags (Walk Away If You See These)

  1. No HIPAA compliance for clinical use (public ChatGPT marketed for medical decisions)
  2. Claims of “replacing physician judgment” (LLMs assist, do not replace)
  3. No prospective clinical validation (only bench mark exam scores)
  4. Autonomous actions without physician review (medication ordering, diagnosis without oversight)
  5. Vendor refuses to discuss hallucination rates (has not tested or hiding poor performance)

Part 8: Cost and Benefit Evaluation

What Does LLM Technology Cost?

Ambient documentation (including Nuance DAX and Abridge): - Cost: Use the executed contract, implementation, integration, support, training, and governance costs - Benefit: Measure local documentation time, after-hours work, note quality, throughput, and clinician experience - Financial effect: Repurposed time is not automatically cash savings; specify how time changes staffing, capacity, or compensated activity - Non-monetary effect: Burnout and work experience require direct measurement with an appropriate comparator

Model API deployment: - Cost: Model prices, token accounting, caching, retrieval, storage, monitoring, and support change frequently - Privacy: A BAA and approved cloud configuration may be required, but neither alone establishes a safe workflow - Full cost: Include engineering, evaluation, security, integration, monitoring, incident response, and model-change management - Comparison: Use current official prices and the actual local workload rather than a static per-note estimate

LLM clinical decision support: - Cost: Verify the current official plan and enterprise contract - Benefit: Define the precise task and measure decision quality, error, time, escalation, and downstream care - Return: Unclear without comparative local evidence and an explicit model of what resources actually change

Epic LLM integration (message drafting, note summarization): - Cost: Bundled into EHR licensing for institutions - Benefit: Incremental time savings across multiple workflows

Do These Tools Save Money?

Ambient documentation: potentially, but not universally - Some implementations reduce documentation time or burden; effects vary - Reduced after-hours charting can matter even when it does not produce cash savings - Cost-effectiveness requires local workflow, quality, utilization, and contract data - Caveat: Subscription, integration, training, and governance costs can change the result

API-based documentation assistance: MAYBE - Much cheaper than subscriptions (~$30/month vs. $400-600/month) - But requires IT infrastructure, integration effort - ROI depends on institutional technical capacity

Literature summarization: UNCLEAR - Time savings real (10 min to read guideline vs. 2 min to review LLM summary) - But risk of hallucinations means verification still required - Net time savings modest

Patient education generation: PROBABLY - Faster than writing from scratch - But requires physician review - Best for high-volume needs (discharge instructions, common diagnoses)


Part 9: The Future of Medical LLMs

Evidence-Calibrated Expectations

Likely developments:

  1. EHR-integrated LLMs expand
    • Epic, Cerner, Oracle already deploying
    • Message drafting, note summarization, coding assistance
    • HIPAA-compliant by design
  2. Multimodal medical LLMs
    • Text + images + lab data + genomics
    • “Show me this rash” + clinical history → differential diagnosis
    • Radiology report + imaging → integrated assessment
  3. Task-specific efforts to reduce unsupported output
    • Retrieval-augmented generation (LLM + medical database lookup)
    • Better uncertainty quantification
    • Improved factuality through constrained generation
  4. Prospective clinical validation
    • More randomized and pragmatic evaluations, including trials with null outcomes
    • Cost-effectiveness analyses
    • Comparative studies (LLM-assisted vs. standard care)
    • Early evidence is arriving: Google’s AMIE conversational diagnostic AI outperformed 20 primary care physicians on 30 of 32 performance axes in a randomized double-blind crossover study using 159 clinical scenarios (Tu et al., Nature, 2025); text-based controlled settings, real-world translation remains unvalidated
  5. Regulatory clarity
    • Product- and function-specific application of FDA device and CDS policy
    • State medical board policies on LLM use
    • Malpractice liability precedents
  6. Open-weight models democratizing global access
    • DeepSeek, Llama 3, Mistral offer computational efficiency for resource-constrained settings
    • A 2025 Journal of Medical Systems article documents DeepSeek deployment in 90 Chinese tertiary hospitals (Chen et al., 2025)
    • Cost: 6.71% of proprietary models (OpenAI o1) with comparable performance on medical benchmarks
    • Critical caveat: Deployment at scale without prospective clinical validation raises safety concerns
    • See AI and Global Health Equity for implementation guidance

Unlikely (despite hype):

  1. Broad autonomous diagnosis or treatment across unbounded clinical contexts
    • Too high-stakes for pure LLM decision-making
    • Human oversight will remain essential
  2. A universal guarantee of factual output
    • Fundamental to how LLMs work
    • Mitigation, not elimination, is realistic goal
  3. Replacement of the physician-patient relationship as an evidence-based goal
    • LLMs assist communication, do not replace human connection
    • Empathy, trust, shared decision-making remain human domains

Modern Clinical Decision Support: Evidence Synthesis Tools

The clinical decision support landscape has shifted dramatically with the emergence of large language model-based tools. While traditional CDS focused on rule-based alerts and drug interaction warnings, modern systems provide evidence synthesis, clinical reasoning support, and real-time guideline retrieval. Adoption has been rapid but evidence of clinical impact remains limited.

Kohane’s NEJM AI editorial “Guideline Machines” states that clinical guidelines are outdated and that using LLMs to improve them would be advantageous were it not for significant potential dangers (Kohane, 2026). He proposes focusing AI on errors of omission and generating a bias-free summary of key messages from the AI output. Named dangers are not specified in the accessible abstract; this is an editorial argument, not outcome evidence.

Clinical Evidence LLMs

Clinical evidence LLMs (OpenEvidence, UpToDate Expert AI, Doximity GPT, and peers) are decision-support with inspectable sources, not autonomous clinicians. Independent head-to-heads can overturn marketing: frontier models outperformed specialized tools on MedQA, HealthBench, and clinician-rated queries (Vishwanath et al., 2026; see independent head-to-head testing). Early OpenEvidence evaluations generally show clinically relevant, citation-grounded answers and lower fabricated-citation rates than general chatbots where measured, while often reinforcing rather than changing plans (Artsi et al., 2026; Hurt et al., 2025). Vendor adoption metrics are not accuracy or patient benefit. For trainees, use these tools as scaffolding after an independent differential, not as substitution that authors the Assessment and Plan (Giordano & Jones, 2026; Ong et al., 2026).

Product landscape:

Product Function Evidence and boundary notes
OpenEvidence Evidence synthesis from peer-reviewed literature and guidelines Vendor-reported registration and daily-use figures (including claims that roughly 40% of registered physicians use the platform daily) are activity metrics, not clinical validation (OpenEvidence, 2025)
UpToDate Expert AI Wolters Kluwer retrieval-augmented Q&A on curated clinical content Tested head-to-head against frontier LLMs and OpenEvidence; did not outperform frontier models on benchmark or clinician-rated tasks (Vishwanath et al., 2026)
Doximity GPT Physician-network LLM for clinical and administrative queries Limited independent peer-reviewed evaluation; treat vendor positioning as hypothesis until locally validated
ChatGPT for Healthcare / Enterprise General-purpose LLM with contractual PHI boundaries Covered-entity use requires an approved account type, configuration, and BAA; consumer ChatGPT is not interchangeable (OpenAI product page)
Google AI Overview Search-integrated health summaries Grouped with specialized clinical tools in clinician-rated real-query testing; no independent superiority claim (Vishwanath et al., 2026)

Keep / Trap:

  • A citation that exists and appears authoritative is not proof of clinical correctness. Real DOIs can be attached to unsupported claims (Cabral et al., 2024).
  • General-purpose chatbots fabricate references at substantial rates in controlled studies (Gravel et al., 2023; Walters, 2023). Evidence LLMs show lower fabrication where measured, but citation integrity and clinical correctness are separate dimensions.
  • Vendor-reported daily-use percentages and consultation counts are not accuracy or patient benefit.
  • Frontier general-purpose models can outperform specialty-branded tools on shared benchmarks (Vishwanath et al., 2026).
  • Early evaluations suggest OpenEvidence often reinforces existing plans rather than changing management (Artsi et al., 2026; Hurt et al., 2025).

CDS versus ambient documentation: Evidence LLMs answer clinical questions and retrieve literature. They are not ambient scribes. Products such as Abridge and DAX Copilot generate encounter documentation from audio; evaluation endpoints, privacy controls, and workflow integration differ. See AI Ambient Documentation Systems for scribe evidence and Evaluating AI Clinical Decision Support Systems for CDS evaluation.

Integration Patterns

EHR-embedded deployment: Health systems have begun integrating evidence synthesis tools as SMART on FHIR apps within Epic, allowing physicians to query evidence without leaving the EHR workflow. This addresses a critical barrier: context switching between EHR and external tools.

Point-of-care use: Unlike traditional CDS that interrupts workflow with alerts, modern tools are pull-based. Clinicians query them when needed rather than receiving unsolicited pop-ups. This reduces alert fatigue but requires clinician initiative.

Evidence Gaps and Adoption Barriers

Despite widespread use, physicians express significant concerns. According to AMA surveys, nearly half of physicians (47%) ranked increased oversight as their top regulatory priority for AI tools, and 87% cited not being held liable for AI model errors as a critical factor for adoption (AMA, February 2025). Key concerns include accuracy and misinformation risk, lack of explainability, and legal liability.

The Stanford-Harvard ARISE assessment (2026): The State of Clinical AI Report is a secondary landscape synthesis. It is useful for locating topics and studies, but product and clinical claims should be checked against primary publications, regulator records, and society guidance.

Key concern: These tools often provide synthesized information with high confidence but limited transparency about source quality, evidence strength, or knowledge cutoff dates. When asked about emerging diseases, recent guideline updates, or off-label uses, LLM-based systems may hallucinate references or conflate older evidence with current recommendations.

When Modern CDS Works: Evidence Synthesis at Scale

Successful use case: Clinical guideline synthesis

OpenEvidence excels at queries like “What are the current USPSTF recommendations for colorectal cancer screening in average-risk adults?” where: - Guidelines are publicly available and well-established - Question has a clear, documented answer - Timeliness matters but changes are infrequent - Synthesis across multiple sources adds value

Problematic use case: Rare disease diagnosis

The same tools struggle with queries like “What’s the differential diagnosis for this constellation of symptoms?” where: - Medical literature is vast and pattern-matching fails - Rare presentations require systematic reasoning, not synthesis - Hallucination risk is high when training data is sparse - Clinical judgment and experience are essential

Clinical Practice Implications

Hospital and clinic adoption: Healthcare systems increasingly use these tools for: - Point-of-care guideline lookup (e.g., “What are current AHA recommendations for NSTEMI antiplatelet therapy?”) - Drug information queries (e.g., “What’s the renal dosing adjustment for vancomycin?”) - Evidence synthesis for clinical decision-making

Risks in resource-limited settings: When CDS tools are the primary source of clinical guidance (e.g., rural hospitals without on-site specialists, solo practices), hallucinated information or outdated recommendations can propagate without detection. Unlike well-staffed academic centers where multiple physicians cross-check recommendations, single-provider settings lack redundancy.

Lessons from Comparative Evidence: Collaboration Is Not Uniform

Comparative studies do not support a universal claim that AI-assisted care outperforms either component alone. Human-model performance depends on the task, interface, user behavior, comparator, and endpoint.

  • Some radiology workflows have shown improved detection with AI assistance under defined conditions
  • Some simulated physician studies report improved treatment-reasoning scores, while another randomized diagnostic-reasoning study found no significant benefit from model access

However, collaboration is not yet optimized. Deskilling remains a real concern: if clinicians defer to AI recommendations without understanding the reasoning, diagnostic skills atrophy. The optimal collaboration pattern requires:

  1. AI provides evidence synthesis, not definitive answers
  2. Clinician integrates context: patient preferences, comorbidities, social determinants
  3. Uncertainty is explicit: AI indicates confidence and knowledge gaps
  4. Auditability: Clinician can trace recommendations to source evidence

Comparison to Traditional CDS Failures

Modern evidence synthesis tools avoid some failure modes of traditional CDS:

Epic sepsis model: - Proprietary, black-box algorithm - External validation at one health system found 33% sensitivity and 12% positive predictive value (Wong et al., 2021) - The study documented substantial false-alert burden - It did not test whether the model improved patient outcomes

OpenEvidence/modern CDS: - Pull-based (clinician-initiated) rather than push-based (unsolicited alerts) - No false positives from unsolicited alerts - Transparency varies (some cite sources, others do not) - Outcome evidence still absent

Shared risk: Both can be deployed at scale without evidence that matches the intended claim. Regulatory status depends on the exact product function, intended use, and applicable law; EHR integration or a general-purpose model label does not decide device status.

Open Questions

  1. Outcome measurement: Does AI-assisted clinical decision-making improve patient outcomes? Reduce diagnostic errors? Shorten time to treatment?

  2. Liability: When AI provides incorrect information that a clinician follows, who is liable? The tool vendor? The clinician? The health system?

  3. Equity: Do these tools work equally well for questions about diseases affecting underrepresented populations? For conditions primarily researched in high-income countries?

  4. Knowledge currency: How do these tools handle emerging evidence? COVID-19 revealed that guidelines changed weekly. LLMs trained on historical data cannot capture real-time updates without continuous retraining or retrieval-augmented generation.

For comprehensive vendor evaluation criteria, see Appendix: Vendor Evaluation Framework. For regulatory frameworks, see Evaluating AI Clinical Decision Support Systems.


### Patient-Facing AI: Unique Safety Considerations

AI systems increasingly interact directly with patients through chatbots, symptom checkers, health education platforms, and digital assistants. These applications present distinct safety challenges compared to clinician-facing tools because patients lack clinical training to identify AI errors and may act on incorrect information without professional oversight.

The Evidence Gap

The Stanford-Harvard ARISE State of Clinical AI Report (2026) found that patient-facing AI evaluation relies primarily on engagement metrics (user satisfaction, session length, return visits) rather than outcome-focused evidence. Few studies measure whether these tools improve health outcomes, reduce diagnostic errors, or facilitate appropriate care escalation (Stanford Medicine, January 2026).

A mixed-methods evaluation of Patiently AI on synthetic clinician notes combined readability scoring, expert safety review, and patient preference, and reported mean Flesch-Kincaid Grade Level falling from 10.57 to 7.61 with 87.3% of expert ratings judged clinically safe (Lamb, 2026). The sole author is the commercial developer of Patiently AI; the study used synthetic notes only, residual safety flags were 12.7% of expert ratings, and the results do not measure care outcomes or constitute an endorsement of the application (Lamb, 2026).

What’s measured: - User satisfaction scores (typically 70-85%) - Engagement rates (time spent, return visits) - Completion rates for health assessments

What’s rarely measured: - Diagnostic accuracy for patient-reported symptoms - Appropriate triage and escalation to human care - Patient safety outcomes (delayed diagnosis, inappropriate self-treatment) - Health literacy impact (do users understand AI limitations?)

Randomized Evidence: Standalone Accuracy Did Not Translate to User Benefit

A 2026 randomized study in Nature Medicine (n=1,298 UK participants) tested whether LLM access helped the general public identify medical conditions and choose appropriate care in simulated scenarios (Bean et al., 2026).

Tested alone on ten clinical scenarios, three LLMs (GPT-4o, Llama 3, Command R+) correctly identified relevant conditions in 94.9% of cases. When the same LLMs were given to participants, condition identification dropped below 34.5%, no better than a control group using internet search. Participants using LLMs were significantly less likely to identify relevant conditions than those using traditional resources (odds ratio 1.76 in favor of control, 95% CI 1.45-2.13). Disposition accuracy (choosing the right level of care) showed no significant difference between LLM and control groups.

The study identified two interaction failure points:

  1. Users provided incomplete information. In 16 of 30 sampled transcripts, initial messages contained only partial symptom descriptions. LLMs cannot compensate for what patients do not report.
  2. Users ignored correct LLM suggestions. LLMs mentioned at least one relevant condition in 65-73% of conversations, but participants failed to incorporate these into their final responses. LLMs suggested an average of 2.21 conditions per interaction, of which only 34% were correct, leaving users unable to distinguish accurate from inaccurate suggestions.

Critically, neither standard medical benchmarks (MedQA) nor simulated patient interactions predicted these failures. LLMs passed medical licensing exam questions at rates exceeding 60% for the same clinical topics where real participants scored below 20%. Simulated users (LLMs pretending to be patients) performed better and showed less variability than real humans, making them unreliable proxies for safety testing.

Clinical implication for physicians: In this randomized simulation, LLM access did not improve the tested user outcomes. The result does not estimate the accuracy of every patient-facing product or every real-world user. Conduct a standard clinical assessment rather than anchoring on an AI-generated opinion.

Risk Categories

1. Misplaced Patient Trust

Patients often cannot distinguish between: - Evidence-based health information - AI-generated plausible-sounding misinformation - General wellness advice vs. medical recommendations requiring professional oversight

Hypothetical failure mode: A patient uses a symptom checker for chest pain. The system emphasizes a common nonurgent explanation without reliable emergency escalation. The patient interprets the response as individualized triage and delays evaluation for a time-sensitive condition.

Why this happens: LLMs optimize for plausible responses, not safety. Chest pain + young patient + no risk factors → statistically likely to be benign. But rare dangerous causes require different reasoning: maximize safety, not likelihood.

2. Delayed Escalation to Professional Care

Patient-facing AI may inadvertently discourage appropriate care-seeking: - Chatbot provides reassurance for concerning symptoms - AI suggests home remedies for conditions requiring clinical evaluation - User perceives AI response as definitive medical advice

The central evaluation gap is that engagement does not establish safe triage, appropriate escalation, or improved patient outcomes.

3. Health Literacy and Informed Consent

Patients using AI health tools often do not understand: - Regulatory status depends on the exact product, function, intended use, and jurisdiction - Recommendations are not reviewed by healthcare professionals - AI can hallucinate references, statistics, or treatment guidelines - Knowledge cutoff dates mean recent information may be missing

Current state: Few patient-facing AI systems explicitly disclose limitations, error rates, or when to seek professional care instead. Terms of service often include liability disclaimers buried in legal text.

Case Example: ChatGPT for Health Education

OpenAI announced ChatGPT for Health in January 2026, positioning it as a health education tool. The system answers health questions using GPT-4-level reasoning but without clinical validation or regulatory oversight.

Intended use: General health education, wellness information, understanding diagnoses

Actual use: Patients report using it for: - Self-diagnosis of symptoms - Medication guidance - Treatment decisions (e.g., whether to seek emergency care) - Second opinions on physician recommendations

Safety gap: A Nature Medicine stress test of ChatGPT Health triage using 60 clinician-authored vignettes across 21 clinical domains (960 total responses) confirmed the system cannot reliably distinguish acuity levels. Failures concentrated at clinical extremes: 52% of gold-standard emergency cases were under-triaged to “see a doctor within 24–48 hours” rather than the emergency department, while the system correctly handled classical presentations like stroke and anaphylaxis. Family-member language minimizing symptoms shifted edge-case recommendations toward less urgent care (OR 11.7, 95% CI 3.7–36.6), and crisis intervention messaging fired inconsistently across suicidal ideation scenarios (Ramaswamy et al., 2026).

The system provides confident-sounding responses regardless of uncertainty, creating a false sense of security.

Usage patterns at scale: These behaviors are not limited to ChatGPT Health. A Microsoft Research analysis of over 500,000 de-identified health-related Copilot conversations (January 2026) found that nearly one in five involved personal symptom assessment, roughly one in seven personal-health queries were asked on behalf of someone else (driven by caregivers managing both children and aging parents), and healthcare-system navigation (finding providers, understanding insurance, accessing care) represented a substantial share of use. Personal symptom and emotional-health queries increased during evening and nighttime hours, when traditional care access is most limited (Costa-Gomes et al., 2026, Microsoft Research).

Physician Perspective: Managing Patient AI Use

Common scenario: Patient arrives with printout from ChatGPT suggesting diagnoses and treatments.

Challenges: - Correcting misinformation takes clinical time - Undermines physician-patient trust if AI contradicts physician - Patients may doctor-shop if physician disagrees with AI - Liability concerns if patient follows AI advice instead of medical recommendation

Communication strategies: - Acknowledge patient’s research initiative - Explain AI limitations (hallucinations, lack of individualization) - Review AI recommendations together, correct errors - Emphasize importance of individualized care vs. generic advice - Document patient’s AI use and physician guidance in chart

ARISE Recommendations

The ARISE report calls for:

  1. Clearer evidence requirements before widespread patient-facing deployment
  2. Stronger escalation pathways to human clinical oversight for concerning symptoms
  3. Evaluation frameworks focused on outcomes, not engagement:
    • Does AI improve health-seeking behavior for serious conditions?
    • Does AI reduce unnecessary ED visits for benign conditions?
    • Does AI correctly identify high-risk scenarios requiring immediate care?
  4. Transparency about limitations:
    • Explicit disclaimers about non-medical-device status
    • Clear guidance on when to seek professional care
    • Disclosure of knowledge cutoff dates and evidence quality

Safety Design Patterns

Pattern 1: Tiered Escalation

User query → AI assessment → Risk stratification:
- High risk (chest pain, severe headache, difficulty breathing) → Immediate redirect to 911 / emergency care
- Medium risk (persistent symptoms, worsening condition) → Prompt to schedule clinical visit within 24-48 hours
- Low risk (wellness, general information) → AI provides information with caveat to seek care if symptoms change

Pattern 2: Uncertainty Disclosure

AI response includes:
- Confidence level (high/medium/low)
- Knowledge gaps ("I don't have information about interactions with your specific medications")
- Explicit limitations ("This is educational information, not medical advice")
- Actionable next steps ("If symptoms worsen or persist beyond X days, seek professional care")

Pattern 3: Human-in-the-Loop for High Stakes

For scenarios with potential serious outcomes: - AI flags query as high-risk - Routes to nurse triage line or clinical decision support - Logs interaction for quality review - Does NOT provide definitive AI-generated recommendation without human oversight

Failure Mode: The Safety Theatre Problem

Some patient-facing AI tools include disclaimers like “This is not medical advice” while functionally operating as medical decision support tools. This creates liability protection for vendors while failing to protect patients.

Example: Symptom checker provides differential diagnosis with probabilities, recommends specific tests, suggests when to seek care vs. self-treat. Footer says “Not medical advice - consult your doctor.” Patient acts on recommendations without professional consultation because AI seemed authoritative.

The disconnect: Legal disclaimer contradicts functional design. If tool is not meant for medical decision-making, why provide diagnosis and treatment recommendations?

Regulatory Gaps

Most patient-facing health AI is not regulated as medical devices because vendors market them as “wellness” or “educational” tools rather than diagnostic systems. This creates a regulatory arbitrage:

  • FDA-cleared diagnostic AI (e.g., IDx-DR for diabetic retinopathy): Rigorous clinical validation, performance standards, post-market surveillance
  • General wellness AI (most chatbots and symptom checkers): No validation requirements, no performance standards, no adverse event reporting

The distinction depends on marketing claims, not actual use. Patients do not distinguish between regulated and unregulated tools.

Recommendations for Clinical Practice

When patients use AI health tools:

  1. Ask proactively: “Have you looked up your symptoms online or used any health apps?” normalizes discussion

  2. Review AI recommendations together: Do not dismiss outright; use as teachable moment about evidence-based medicine

  3. Document: Note patient’s AI use and your clinical guidance in chart for liability protection

  4. Educate about escalation: Teach patients red flag symptoms requiring immediate care regardless of AI reassurance

  5. Equity considerations: Does AI work equally well for:

    • Limited English proficiency patients?
    • Low health literacy populations?
    • Patients without reliable internet access or smartphones?
    • Conditions affecting underrepresented communities?

For comprehensive safety evaluation frameworks, see Clinical AI Safety and Risk Management. For equity considerations, see Medical Ethics, Bias, and Health Equity.


Part 10: Implementation Guide

Safe LLM Implementation Checklist

Pre-Implementation:

During Use:

Post-Implementation:


Part 11: Institutional LLM Deployment Evidence

While individual physicians experiment with LLMs, health systems face a distinct challenge: how to deploy LLMs at institutional scale with appropriate governance, workflow integration, and continuous monitoring. Stanford Medicine’s ChatEHR is a peer-reviewed Nature Medicine deployment report of an EHR-embedded, clinician-prompted LLM (Shah et al., 2026). Benchmark-based evaluations are insufficient for monitoring clinician-driven EHR interactions; unsupported-claim review has to run on live sessions. Usage, error, and first-year value figures remain single-institution modeled estimates from the companion technical report (Shah et al., 2026, technical report), not a universal deployment benchmark. The authors are Stanford employees, and Stanford University owns ChatEHR.

The Stanford ChatEHR Model

Stanford Health Care developed ChatEHR as an institutional capability rather than adopting external vendor solutions, enabling what they term a “build-from-within” strategy (Shah et al., 2026). The platform connects multiple LLMs (OpenAI, Anthropic, Google, Meta, DeepSeek) to the complete longitudinal patient record within the EHR.

Two deployment modes:

Mode Description Use Case
Automations Static prompt + data combinations for fixed tasks Transfer eligibility screening, surgical site infection monitoring, chart abstraction
Interactive UI Chat interface within EHR for open-ended queries Pre-visit chart review, summarization, clinical questions

The implementation hypothesis: Standalone tools can create workflow friction from manual data entry. The deployment report describes access to longitudinal records inside an institutional workflow as one design choice, alongside continuous evaluation. Other architectures can be appropriate if they satisfy the intended task, privacy, safety, and usability requirements.

Unsupported-output measures from deployment

The companion technical report reports task-specific unsupported-output measures from clinical use rather than a laboratory-only benchmark. Definitions, sampling, adjudication, and task mix determine how these figures should be interpreted.

Summarization accuracy (the most common task, 30% of queries):

Metric Rate Definition
Hallucinations per summary 0.73 Statements not found in or supported by the patient record (e.g., mentioning a procedure not documented)
Inaccuracies per summary 1.60 Statements contradicting information in the record (e.g., reporting lab value of 5.0 when record states 3.5)
Total unsupported claims 2.33 per summary Combined hallucinations + inaccuracies
Summaries with ≤1 error 50% Half of all summaries had at most one unsupported claim

Error types identified: - Temporal sequence errors (misstating when events occurred) - Numeric value confusion (labs, vitals) - Role attribution errors (misstating who performed an action) - “Gestalt” care plan confusion (e.g., conflating pulmonary workup with cardiovascular workup)

Clinical implication: Every LLM-generated clinical summary requires verification. In the evaluated sample, half of summaries had at most one unsupported claim under the authors’ definitions. The result should not be transferred to other tasks, models, settings, or versions.

Adoption and Usage Patterns

User training and adoption:

  • 1,075 clinicians completed mandatory training video
  • 99% reported the training video prepared them for use
  • Training emphasized: key features, prompting tips, and the non-negotiable requirement to verify outputs

Usage scale (first 3 months of broad deployment):

Metric Value
Total sessions 23,000+
Daily active users ~100
Tokens processed 19 billion
Sessions using external HIE data >50%
Most common task Summarization (30%)

User types: 424 physicians, 180 APPs, 151 residents, 60 fellows using the system at least once.

Response times: Most queries returned in under 20 seconds, though complex patient records with large timelines could take up to 50 seconds (with 95%+ of time spent assembling the FHIR record bundle, not LLM inference).

Value Assessment Framework

Stanford developed a structured framework to quantify LLM deployment value across three categories:

Category Definition Example
Cost savings Direct monetary reductions Avoided manual chart review labor
Time savings Decreased time on manual tasks (time repurposed, not eliminated) 120 charts/day avoided manual review = ~4 hours saved
Revenue growth Incremental revenue from new workflows or improved throughput Increased transfer throughput freeing beds

Companion-report author estimate: The authors modeled approximately $6 million in first-year value. This was an economic estimate, not an independently audited cash-savings result or a measured patient benefit.

Example automation ROI:

Automation Time Savings Revenue Impact
Transfer eligibility screening Author-estimated labor effect Modeled transfer and bed-capacity effect
Inpatient hospice identification Author-estimated hours Not directly quantified
Pre-visit chart review Author-estimated avoided review volume Time modeled as returned to clinical care

Illustrative value calculation from the companion report: The authors combined assumed users, queries, time saved, hourly value, and API cost to estimate annual value. Each input must be measured locally. Repurposed clinician time should not be presented as realized cash savings without a corresponding staffing, throughput, or compensated-activity change.

Governance and Monitoring

Stanford implemented continuous monitoring across three dimensions:

1. System integrity monitoring: - Response times, error codes, timeouts - Token usage and cost tracking - Data retrieval verification (re-extracting benchmark patients to detect upstream changes)

2. Performance monitoring: - In-workflow feedback collection (thumbs up/down) - Task categorization from interaction logs - Unsupported claims rate estimation on sample of sessions - Benchmark dataset maintenance for each automation

3. Impact monitoring: - Action rates (what proportion of flagged patients received recommended follow-up) - Documentation metrics (use of generated text in notes) - Engagement metrics (views, copies, repeat usage)

User feedback rates: ~5% of users provided feedback; of those, two-thirds was positive (thumbs up).

Lessons for Health Systems

What the Stanford deployment report illustrates:

  1. Workflow integration can support adoption. Copy-paste from standalone tools can create friction; the described EHR integration saw sustained use in one institution.

  2. Automations require different governance than interactive use. Predefined prompt + data combinations can be validated against truth sets and monitored systematically. Interactive chat requires task categorization, sampling, and ongoing quality assessment.

  3. Benchmark performance is necessary but insufficient. Stanford used MedHELM (Nature Medicine, 2026) for initial model selection. The VoR lesson is that clinician-driven EHR interactions require live-session monitoring, not benchmark selection alone; temporal confusion and numeric errors still appeared after that step (Shah et al., 2026).

  4. User training is one control. Mandatory training preceded access and users reported preparedness. This association does not establish that training caused safe behavior or durable verification.

  5. Value quantification requires structured frameworks. Time savings are “soft” (repurposed, not eliminated) and harder to measure than cost savings or revenue growth. Combining all three provides a more complete picture.

What remains challenging:

  • Verification behavior erosion: As users become accustomed to AI-generated content, verification may decline
  • Automated fact verification is still maturing (see VeriFact for emerging approaches)
  • Value of improved patient care is difficult to quantify

The Vendor-Agnostic Advantage

The “build-from-within” approach provides institutional agency that vendor-dependent deployments do not:

Factor Vendor Solution Institutional Platform
Model selection Vendor’s chosen model Best model for each task
Data governance Vendor’s terms Institution controls data
Customization Limited to vendor options Task-specific automations
Monitoring Vendor-provided dashboards Custom metrics aligned to institutional priorities
Cost Per-user or enterprise licensing Usage, engineering, infrastructure, and support costs
Continuity Vendor business risk Internal capability persists

The strategic tradeoff: Internal platforms can increase institutional control and customization, but they also create engineering, validation, monitoring, security, and maintenance obligations. Vendor platforms shift some implementation burden but introduce contractual, continuity, and configuration dependencies.


Key Takeaways

10 Principles for LLM Use in Medicine

  1. LLMs are tools, not licensed clinicians: Assign professional decision-making to accountable people and teams

  2. Unsupported output remains possible: Verify material clinical facts, patient data, calculations, and citations

  3. Privacy review is product- and workflow-specific: Do not enter protected information until the exact service and use are institutionally approved

  4. Appropriate uses: Documentation drafts, literature review, education materials (with review)

  5. Inappropriate uses: Autonomous diagnosis/treatment, medication dosing without verification, urgent decisions

  6. Accountability is not transferred to the model: Liability allocation remains fact- and jurisdiction-specific

  7. Evidence is endpoint-specific: Licensing-exam performance is not clinical utility; require a design that matches the intended claim

  8. Ambient documentation has peer-reviewed evidence: Effects depend on product and workflow, do not assume large time savings by default (Rotenstein et al., 2026; Lukac et al., 2025)

  9. Prompting is not validation: Prompt design can change output, but sources, records, calculations, and clinical decisions still require verification

  10. The unit of evaluation is the human-model workflow: Compare the full system with relevant care, not the model in isolation


Clinical Scenario: LLM Vendor Evaluation

Scenario: A Hypothetical LLM Vendor Evaluation

The pitch (hypothetical composite, not a current product profile): A vendor provides LLM-powered differential diagnosis and treatment suggestions. Marketing claims: - “Physician-level diagnostic accuracy” - “Evidence-based treatment recommendations” - “Saves 20 minutes per complex case” - A per-physician subscription with separate enterprise terms

The CMO asks for your recommendation.

Questions to ask:

  1. “What peer-reviewed validation studies support this exact product and version?”
    • Request JAMA, Annals, specialty journal publications
    • User testimonials ≠ clinical validation
  2. “Is this HIPAA-compliant? Where is the BAA?”
    • Essential for entering patient data
  3. “What is the hallucination rate?”
    • If vendor has not quantified, they have not tested properly
  4. “How does the system handle diagnostic uncertainty?”
    • Does it express appropriate uncertainty or confidently hallucinate?
  5. “What workflow oversight prevents acting on incorrect recommendations?”
    • Best systems require physician review before actions
  6. “Can we pilot with 10 physicians before hospital-wide deployment?”
    • Local validation essential
  7. “What happens if a recommendation contributes to harm?”
    • Read liability disclaimers in contract
  8. “What is actual time savings data?”
    • “20 minutes per complex case” claim: where’s the evidence?

Red Flags:

“Physician-level accuracy” without prospective validation

No discussion of hallucination rates or error modes

Marketing emphasizes speed over safety

No built-in verification mechanisms


Check Your Understanding

Scenario 1: The Medication Dosing Question

Clinical situation: You’re seeing a 4-year-old with otitis media requiring amoxicillin. You ask GPT-4 (via HIPAA-compliant API):

“What is the appropriate amoxicillin dosing for a 4-year-old child with acute otitis media?”

GPT-4 responds: “For acute otitis media in a 4-year-old, amoxicillin dosing is 40-50 mg/kg/day divided into two doses (every 12 hours). For a 15 kg child, this would be 300-375 mg twice daily.”

Question 1: Do you prescribe based on this recommendation?

Click to reveal answer

Answer: No, verify against authoritative source first.

Why:

The LLM response may be incomplete or inapplicable: - Dose selection depends on the current guideline, indication, weight, allergy history, recent antimicrobial exposure, local resistance, renal function, formulation, and institutional protocol - A plausible dose is not evidence that the correct regimen, concentration, frequency, or duration was selected for this child - The model’s training date does not reveal whether a specific answer is current or correct

The controlling sources: - The current pediatric guideline and institutional antimicrobial pathway - The exact formulary product and pharmacy dosing reference - Patient-specific history, examination, renal function, allergy assessment, and recent treatment

What you should do: 1. Check UpToDate, Lexicomp, or AAP guidelines directly 2. Confirm the indication, dose, frequency, duration, formulation, and maximum dose 3. Resolve any mismatch with the pharmacist or authoritative pathway before prescribing

The lesson: LLMs may provide outdated recommendations or miss recent guideline updates. Always verify medication dosing against current pharmacy databases or guidelines.

If an unverified dose were prescribed: - The child could receive an inappropriate dose, formulation, frequency, or duration - Treatment failure, adverse effects, or delayed reassessment could follow - The exact harm cannot be inferred from this hypothetical prompt alone


Scenario 2: The Patient Education Handout

Clinical situation: You’re discharging a patient newly diagnosed with type 2 diabetes. You use GPT-4 to generate patient education handout:

“Create a one-page patient handout for newly diagnosed type 2 diabetes, 8th-grade reading level. Cover: medications, blood sugar monitoring, diet, exercise.”

GPT-4 generates professional-looking handout with sections on metformin, glucometer use, carb counting, and walking recommendations.

Question 2: Can you give this handout to the patient as-is, or do you need to review/edit first?

Click to reveal answer

Answer: MUST review and edit before giving to patient.

Why:

Potential LLM errors to check:

  1. Medication information:
    • Is metformin dosing correct? (LLMs sometimes hallucinate dosages)
    • Are side effects accurately described?
    • Are contraindications mentioned? (metformin contraindicated in advanced CKD)
  2. Glycemic targets:
    • Are targets drawn from the current guideline and individualized for the patient?
    • Has the draft confused diagnostic thresholds, treatment goals, and urgent-action thresholds?
  3. Dietary advice:
    • Is carb counting explained clearly?
    • Any culturally inappropriate diet suggestions?
    • Consistent with current ADA nutrition guidelines?
  4. Exercise recommendations:
    • Are recommendations consistent with the patient’s current health status and treatment plan?
    • Does the draft include inappropriate universal restrictions or omit reasons to seek clinical advice?
  5. Individualization:
    • LLM generates generic handout
    • Does it fit THIS patient’s health literacy, culture, comorbidities?

What you should do: 1. Read entire handout for medical accuracy 2. Verify all medical facts (especially medication dosing, blood sugar targets) 3. Customize to patient (add their specific meds if not metformin, adjust for comorbidities) 4. Check cultural appropriateness 5. Review with patient, answer questions

The lesson: LLM-generated patient materials save time but require physician review for accuracy and individualization. Never hand out LLM content without verification.

Time assessment: - Measure drafting and review time in the actual workflow - Include time spent correcting errors, individualizing content, documenting use, and answering patient questions - A faster first draft is not a net benefit if verification burden or downstream misunderstanding increases.


Scenario 3: The Citation Fabrication

Clinical situation: You’re writing a quality improvement manuscript on reducing hospital readmissions. You ask GPT-4:

“Cite 3 recent studies (2020-2023) showing effectiveness of AI for predicting 30-day readmissions in heart failure patients.”

GPT-4 provides:

  1. “Chen et al. (2022). ‘Machine learning for heart failure readmission prediction.’ JAMA Cardiology 7:445-453.”
  2. “Rodriguez-Lopez et al. (2021). ‘Deep learning models outperform LACE index for readmission risk.’ Circulation 144:892-901.”
  3. “Patel et al. (2023). ‘Prospective validation of AI readmission tool in community hospitals.’ New England Journal of Medicine 388:1122-1131.”

Question 3: Can you include these citations in your manuscript?

Click to reveal answer

Answer: NO. You must verify each citation exists and actually supports your claim.

Why:

The LLM likely fabricated some or all of these citations. Here’s how to check:

Step 1: Search PubMed for each citation

For “Chen et al. (2022) JAMA Cardiology”: - Search: "Chen" AND "heart failure readmission" AND "machine learning" AND "JAMA Cardiology" AND 2022 - If found: Read abstract, confirm it supports your claim - If NOT found: Citation is fake

Step 2: Verify journal, volume, pages

Even if an author “Chen” published in JAMA Cardiology in 2022, check: - Is the title correct? - Is the volume/page number correct? - Does the paper actually discuss AI for HF readmissions?

Step 3: Read the actual papers

If citations exist: - Do they support the claim you’re making? - Are study methods sound? - Are conclusions being accurately represented?

Possible outcome: - One or more citations may be fabricated, incomplete, or bibliographically inconsistent - A paper can exist and still fail to support the sentence attributed to it

What you should do instead:

  1. Search PubMed yourself: ("heart failure" OR "HF") AND ("readmission" OR "rehospitalization") AND ("machine learning" OR "artificial intelligence" OR "AI") AND ("prediction" OR "risk score")

  2. Filter: Publication date 2020-2023, Clinical Trial or Review

  3. Read abstracts, select relevant papers

  4. Cite actual papers you’ve read

The lesson: Never trust an LLM-generated citation without verification. Error rates differ across systems and tasks. Confirm that the paper exists, that the author, journal, year, and DOI form one connected record, and that the source supports the exact claim.

Consequences of using fabricated citations: - Manuscript rejection - If published then discovered: Retraction - Academic dishonesty allegations - Career damage

Workflow comparison: - An unverified reference list is fast to generate but not usable evidence - Source retrieval, identity verification, and claim-level reading take additional time - The verification work is part of authorship, not optional overhead


Can physicians safely use ChatGPT with patient data?

A product is not universally HIPAA compliant in the abstract. Before entering protected health information, a covered entity should approve the exact product, account type, configuration, data uses, retention, safeguards, and contractual relationship, including a BAA when required. Consumer use and covered-entity use can have different HIPAA roles.

How do I detect LLM hallucinations in medical content?

Detection requires claim-by-claim verification. Confirm that citations exist and support the sentence, compare material clinical claims with current primary sources or authoritative guidelines, verify patient facts in the record, and check calculations independently. Retrieval and citations reduce some risks but do not prove that the synthesis is supported.

What LLM use cases are safe for clinical practice?

Lower-risk candidates are bounded, reviewable language tasks such as drafting education from approved sources, reorganizing verified text, and administrative drafting. High-risk uses include unverified dosing, urgent triage, diagnosis, treatment, and prognosis. Risk depends on intended use, verification, user, workflow, and consequence of error.

Why do LLMs hallucinate medical information?

LLMs generate token sequences from learned statistical patterns and context; their objective does not by itself verify truth. Unsupported output can arise from training data, prompting, retrieval, ambiguity, decoding, or synthesis. Retrieval and system design can reduce error, but no prompt guarantees factual clinical output.

Where can I find annual reviews of clinical AI evidence?

Secondary reports can help locate studies, but clinical conclusions should be verified against the primary publication, regulator record, or society guideline. Evidence reviews should distinguish benchmarks, simulations, observational deployments, randomized workflows, and patient outcomes.

Conclusion

Large language models can reduce language-work burden and support bounded information tasks. Clinical value depends on the full human-model workflow, not fluent output or benchmark performance alone. The safest implementations define the intended use, verify product and privacy conditions, test the exact workflow, preserve source provenance, and monitor consequential errors after deployment.

Continue Reading