The Physician-AI Partnership: Future Perspectives

Keywords

physician AI partnership, future medicine AI, will AI replace physicians, Khullar 2026, Khullar NEJM, clinical workforce AI, Jevons paradox healthcare, lump of labor AI, O-ring theory medicine, how will AI impact the clinical workforce

Physician-AI collaboration works best in narrow workflows with clear task boundaries, usable interfaces, and verification protocols. Radiologists using AI can detect more cancers in specific screening settings, ambient documentation can reduce administrative burden, and validated decision support can improve defined technical or workflow endpoints. Watson for Oncology offers a different lesson: published studies largely measured concordance with multidisciplinary recommendations, not patient benefit, and performance varied across cancers and settings. A benchmark, authorization, concordance study, workflow evaluation, and patient-outcome trial establish different things. The future of medicine must therefore be assessed task by task, not through a universal replacement or partnership forecast.

Learning Objectives

After reading this chapter, clinicians will be able to:

  • Envision realistic futures for physician-AI partnership (not replacement)
  • Distinguish genuine transformation from technological hype
  • Understand how physician roles will evolve with AI integration
  • Recognize irreplaceable human elements of medicine
  • Develop adaptive strategies for continuous learning and change
  • Advocate for patient-centered, ethical AI development
  • Lead your institutions and profession through AI transformation

Clinical Context

AI has moved from research curiosity to clinical reality. FDA maintains a periodically updated list of AI-enabled medical devices authorized for marketing in the United States (FDA AI-Enabled Medical Device List). Surveys suggest growing adoption, though precise hospital-level deployment data varies by source and definition of “clinical AI.” The central question is no longer “Will AI transform medicine?” but “How will physicians and AI work together?” This chapter pulls together lessons from throughout this handbook and charts evidence-based paths forward.

Key Applications

What Has Been Demonstrated in Defined Settings: - Autonomous diabetic-retinopathy screening for the population, camera, setting, and endpoint specified in FDA authorization and the pivotal study - AI-supported mammography screening with increased cancer detection and lower reading workload in the randomized MASAI study - Product- and workflow-specific documentation effects in randomized and observational evaluations - Selected image-acquisition, interpretation, triage, prediction, and monitoring functions under bounded study conditions

What the Evidence Does Not Establish: - General autonomous diagnosis or management across diseases, populations, and care settings - Replacement of physician judgment in complex, uncertain cases - LLM benchmark transfer: in one study, accuracy changed by 8.8 to 38.2 percentage points when the correct answer was replaced with “none of the other answers,” with reasoning-optimized models least affected. This demonstrates sensitivity to evaluation format; it does not by itself prove the presence or absence of genuine clinical reasoning (Bedi et al., 2025) - Universal benefit from adding a nominal human reviewer or an explanation layer - Transferability across sites without local acceptance testing and monitoring - AI adoption driven solely by cost-cutting rather than patient benefit

What’s Uncertain: - Timeline to general-purpose medical AI - Impact of AI on physician supply/demand (automation vs. expanding access) - Long-term effects of AI dependence on clinical skill retention - Regulation keeping pace with AI evolution (adaptive frameworks needed) - Societal acceptance of AI in high-stakes medical decisions

Critical Insights

  • Physician + AI does not automatically outperform either alone. Collaboration gains depend on task design, interface quality, user training, and verification protocols.
  • Time Horizon Reality: Deployment rates differ by task, evidence, payment, interoperability, regulation, and organizational capacity. Historical analogies do not produce a reliable calendar.
  • Skill Shift: Automation of routine tasks elevates importance of uniquely human skills (empathy, ethical reasoning, communication, advocacy)
  • Hype Cycle: Post-Watson skepticism transitioning to evidence-based optimism for narrow applications
  • Physician Agency: Future shaped by choices physicians make today (governance participation, evidence demands, advocacy)

Clinical Bottom Line

The defensible future is conditional partnership, not automatic partnership. AI can support pattern recognition, data processing, drafting, and standardized tasks, while physicians remain responsible for deciding whether an output is relevant, resolving uncertainty, integrating patient goals, communicating risk, supervising consequential actions, and escalating system failure. Action items: (1) invest in AI literacy, (2) engage in institutional governance, (3) measure unaided and assisted performance, (4) preserve independent reasoning and patient relationships, and (5) demand evidence matched to the claimed benefit.

Essential Reading

  • Topol (2019). “Deep Medicine: How Artificial Intelligence Can Make Healthcare Human Again.” (optimistic vision of physician-AI partnership)
  • Obermeyer & Emanuel (2016). “Predicting the Future: Big Data, Machine Learning, and Clinical Medicine.” NEJM 375:1216-1219 (balanced assessment of AI capabilities)
  • Rajkomar et al. (2019). “Machine Learning in Medicine.” NEJM 380:1347-1358 (technical foundations for clinicians)
  • Emanuel & Wachter (2019). “Artificial Intelligence in Health Care: Will the Value Match the Hype?” JAMA 321:2281-2282 (skeptical counterpoint)
  • WHO (2021). “Ethics and Governance of Artificial Intelligence for Health: WHO Guidance.” (global policy framework)

Introduction

Watson for Oncology was promoted as a broad treatment-support system, but its published evidence was dominated by retrospective concordance studies. A 2021 meta-analysis of nine studies and 2,463 patients estimated 81.52% overall concordance when both “Recommended” and “For consideration” options counted as agreement, with meaningful variation by cancer, stage, and patient characteristics (Jie et al., 2021). Concordance with a local multidisciplinary team is not a test of causal patient benefit. Meanwhile, randomized mammography screening evidence shows that a specific human-AI workflow can improve a defined endpoint under study conditions. The contrast is not that one technology failed and another category succeeded. It is that the scope of the claim must match the design and endpoint of the evidence.


Part 1: IBM Watson for Oncology and the Limits of Concordance

What the case teaches: A system can agree with experts in many retrospective cases without establishing that its recommendations improve treatment selection, survival, quality of life, safety, or cost. Technology alone does not transform medicine. Evidence, implementation, and clinical accountability do.

The Promise (2011-2013)

Historical background: - IBM Watson wins Jeopardy! (2011), demonstrating natural language processing capabilities - IBM then promoted Watson-branded health products for literature processing, clinical decision support, imaging, and data analytics - Watson for Oncology was developed with expert input from Memorial Sloan Kettering and presented treatment options rather than independently treating patients

The commercial pitch, not established clinical evidence: - Watson reads all medical literature (millions of articles) - Analyzes patient-specific data (genomics, medical history, imaging) - Recommends personalized treatment plans superior to human oncologists - “AI will democratize world-class cancer care: every patient gets Memorial Sloan Kettering expertise”

Media hype: - Forbes: “IBM Watson: The Smartest Doctor in the Room?” - Wall Street Journal: “IBM Supercomputer Watson Could Help Cure Cancer” - The Jeopardy! result became a powerful analogy for medicine even though question answering and cancer-treatment selection are different tasks

The Reality (2013-2018)

Implementation and evidence challenges:

Problem 1: Training data mismatch - Expert curation tied the product to particular guidelines, practices, and available treatment options - Local formularies, insurance coverage, patient preferences, molecular testing, and regional practice can make apparently discordant recommendations clinically understandable - Agreement with one expert group cannot establish that the group, the product, or the local comparator chose the outcome-optimal treatment

Problem 2: Natural language processing failures - Clinical decision support depends on accurate extraction of diagnosis, stage, prior treatment, biomarkers, comorbidity, performance status, and patient preference - Any missing or incorrectly entered variable can change the displayed options - Published concordance reports generally assessed the treatment output after cases were entered; they did not establish end-to-end reliability of record extraction in routine care

Problem 3: Lack of validation - The best-known literature measured retrospective concordance, not randomized patient outcomes - Concordance definitions differed. Some counted only “Recommended” options; others combined “Recommended” and “For consideration” - A meta-analysis of nine studies and 2,463 cases reported 81.52% concordance under the broader definition, with variation by cancer, stage, age, performance status, and tumor biology (Jie et al., 2021) - A Chinese multicancer study likewise found that concordance differed by cancer type and stage, demonstrating limited transferability rather than a universal product score (Zhou et al., 2019) - Concordance cannot tell whether the product or comparator was clinically correct when they disagreed, and it cannot estimate treatment benefit.

Problem 4: Physician resistance - Clinicians need to inspect the source guidelines, evidence, version, missing inputs, contraindications, and reason for disagreement - A ranked option without accessible evidence is difficult to contest, especially when the recommendation conflicts with local practice or patient goals - Usability and professional acceptance therefore require more than an agreement percentage

Problem 5: Workflow disruption - End-to-end value depends on who enters and verifies data, how options appear in the tumor-board workflow, whether recommendations are current, and how disagreements are resolved - A product may be technically accurate yet fail if it duplicates work, obscures provenance, or cannot accommodate local treatment constraints - Workflow time must be measured in the intended setting rather than inherited from vendor demonstrations or other hospitals

What the Published Numbers Establish

Evidence Measured result Defensible interpretation
Nine-study meta-analysis, 2,463 cases 81.52% overall concordance when “Recommended” and “For consideration” both counted Retrospective agreement under a broad definition, not patient benefit (Jie et al., 2021)
Metastatic NSCLC retrospective study 85.16% concordance with a medical team Agreement differed by pathology and mutation; no randomized outcome comparison (You et al., 2020)
Colorectal cancer retrospective study, 175 cases 66.9% concordance Associations with survival are vulnerable to confounding because treatment was not randomized by Watson agreement (Liu et al., 2021)
Breast cancer concordance literature Results varied by age, stage, receptor status, setting, and definition Local transferability and comparator choice matter (Somashekhar et al., 2018)

IBM announced in January 2022 that Francisco Partners would acquire healthcare data and analytics assets from the Watson Health business. The official announcement did not disclose the financial terms and described a portfolio broader than Watson for Oncology, including MarketScan, Micromedex, imaging software, and other assets (IBM, 2022). It therefore cannot support a precise loss, write-off, or product-specific return calculation.

The Durable Consequences

1. Concordance is no longer enough. Agreement with an expert panel can be useful for quality review, but it does not substitute for safety, workflow, utility, or outcome evidence.

2. Product identity matters. The Watson Health name covered multiple assets and products. Evidence for one cannot be transferred to another, and a business transaction involving a portfolio does not measure the clinical value of a particular oncology system.

3. Local variation is part of the result. Differences by cancer, stage, population, guideline, and setting are not merely implementation noise. They identify the boundary of the evidence.

4. Overselling creates an evidence debt. When marketing claims exceed the measured endpoint, clinicians and patients must later disentangle what the system actually did from what the brand implied. That repair consumes trust and institutional capacity even when the precise market-wide effect cannot be quantified.

The Lesson for Physicians

Why Watson failed (and what this teaches about AI future):

1. Hype ≠ Evidence - Jeopardy! win demonstrated natural language processing, NOT medical reasoning - Extrapolating from narrow AI success to complex medical application = fallacy - Demand evidence: Prospective trials, peer-reviewed publications, real-world performance data

2. Training data quality > quantity - Watson drew on a curated knowledge base and expert input, but breadth of literature access did not resolve local treatment constraints or establish causal benefit - “Reading” literature ≠ understanding causality, applying to individual patients - Question data sources: Whose data? What population? Validated how?

3. Explainability matters - Oncologists rejected black-box recommendations lacking transparent reasoning - Trust requires understanding WHY AI recommends specific treatment - Insist on transparency: “Show me the evidence. Explain the reasoning. Let me override if I disagree.”

4. Workflow integration critical - Technology succeeds when it REDUCES physician burden, fails when it increases it - Measure data-entry, verification, disagreement-resolution, and follow-up time in the intended workflow - A documentation or triage system should not inherit a time-saving claim from another product or site

5. Physician acceptance non-negotiable - Cannot force-feed AI to resistant physicians → passive sabotage, workarounds - Co-design with end-users from start, not after-the-fact retrofitting

Questions to ask about ANY new AI clinical tool:

Red flags identified by this case: - No prospective validation trials - Black-box algorithm without explainability - Trained on narrow dataset (single institution) - Vendor hype exceeds published evidence - Physicians not involved in design/validation - Increases workflow burden

Evidence and implementation features to seek: - Study design matched to the claim, including comparative prospective evaluation when patient or workflow benefit is asserted - Transparent reasoning (“AI recommends X because Y”) - Diverse training data (multiple institutions, demographics) - Peer-reviewed publications in major journals - Physician co-design and pilot testing - Reduces physician burden or improves efficiency

Current status: IBM announced the sale of specified Watson Health data and analytics assets to Francisco Partners in 2022, without disclosing financial terms (IBM, 2022). This corporate event should not be represented as proof that FDA changed its regulatory pace or that the entire clinical-AI field suffered a measurable multiyear delay. The defensible lesson is narrower and stronger: evidence must remain attached to the exact product, intended use, population, comparator, and endpoint.


Part 2: Radiology Progress Without a Universal Partnership Claim

What the evidence teaches: Some radiology workflows have randomized evidence for specific screening or process endpoints. Others have null or harmful findings. The unit of evidence is the product, version, task, population, comparator, workflow, and endpoint, not “radiology AI” as one intervention.

The Clinical and Operational Problem

Imaging services face increasing study volume, uneven subspecialty access, time-sensitive findings, repetitive measurements, and persistent risks from interruption, fatigue, and communication delay. These pressures make detection, prioritization, acquisition guidance, segmentation, report drafting, and quality monitoring plausible targets. They do not establish that any particular tool is accurate, useful, cost-effective, or safe.

The FDA maintains a periodically updated list of AI-enabled medical devices, with radiology representing the largest specialty category (FDA AI-Enabled Medical Devices). Device count measures authorization activity, not adoption, daily use, errors prevented, lives saved, or return on investment.

The Intervention Families

AI-assisted radiology workflows include products from multiple vendors. Evidence must remain product-specific.

Critical-finding triage - AI pre-screens imaging studies (CT head, CT chest, etc.) - Flags critical findings: Intracranial hemorrhage (ICH), pulmonary embolism (PE), pneumothorax, large vessel occlusion (LVO) - Sends immediate alerts to radiologist + clinical team - Goal: Reduce time-to-diagnosis for time-sensitive conditions

AI-assisted interpretation and screening - AI can highlight findings, assign scores, support single or double reading, or route low-risk examinations under a prespecified protocol - Radiologist reviews AI assessment, signs final report - Goal: Increase radiologist efficiency, reduce fatigue on routine studies

Acquisition, measurement, reporting, and quality support - AI highlights suspicious findings (calcifications, nodules, fractures) radiologist might miss - Acquisition guidance can help an operator capture a limited examination; segmentation and measurement tools can structure quantitative work; generative systems can draft report text - These functions have different users, failure modes, regulatory records, and outcome pathways - The phrase “second set of eyes” describes an interface concept, not an evidence grade.

The Evidence

Stroke workflow, Viz.ai: In a stepped-wedge cluster-randomized trial involving 243 patients treated with endovascular thrombectomy, activation of the Viz.ai workflow reduced adjusted door-to-groin time by 11.2 minutes and CT-to-treatment time by 9.8 minutes. The trial did not show a significant improvement in 90-day functional independence (Martinez-Gutierrez et al., 2023). An 82-patient pre/post study found no statistically significant overall door-to-groin change, although an off-hours subgroup improved (Figurelle et al., 2023). These findings cannot be assigned to RapidAI, Brainomix, or every stroke center.

Mammography screening, MASAI: The randomized MASAI trial assigned 105,934 women to AI-supported screening or standard double reading. A protocol-defined analysis reported 6.4 versus 5.0 cancers detected per 1,000 screened women and a 44.2% reduction in screen-reading workload, without a statistically significant increase in recall or false-positive rates (Lång et al., 2025). Full two-year follow-up reported sensitivity of 80.5% versus 73.8% at the same 98.5% specificity, while reductions in interval and aggressive cancers did not reach statistical significance in the noninferiority design (Gommers et al., 2026). These results support the tested Transpara workflow, not a class-wide cancer-detection percentage.

Chest-radiograph pathway, LungIMPACT: A prospective multicenter randomized study of 93,326 chest radiographs found that AI worklist prioritization did not reduce time to chest CT or time to lung-cancer diagnosis (Woznitza et al., 2026). A favorable prioritization score does not guarantee a changed patient pathway.

Historical mammography CAD: In a large observational analysis, earlier CAD use was associated with increased recall and no improvement in cancer detection or diagnostic accuracy (Lehman et al., 2015). Newer systems require their own evaluation and should not inherit either benefit or harm solely from the category name.

Partnership model evidence remains task-specific:

Task Radiologist Only AI Only Radiologist + AI Best Performer
Stroke triage Faster selected process intervals in one Viz.ai randomized workflow No significant 90-day functional-independence benefit Process effect, not established functional-outcome benefit
Mammography screening Higher cancer detection and lower reading workload in MASAI Product and screening-program specific Supports the tested configuration
Chest-radiograph prioritization Worklist intervention tested across a patient pathway Null time-to-CT and time-to-diagnosis results Technical prioritization may not change care
Human-AI interaction Correct and incorrect advice can change reader performance Effect depends on interface, task, expertise, and error correlation Collaboration is a separate intervention

Key insight: Radiologist + AI can improve performance in selected settings, but the effect is not automatic. The strongest implementations preserve independent review, show uncertainty clearly, and make AI outputs easy to verify.

Why Some Radiology Workflows Have Stronger Evidence Than Watson

Success Factor Radiology AI Watson Oncology
Task scope Narrow, well-defined (detect ICH, PE, tumors) Broad, complex (recommend cancer treatment)
Validation Some products have randomized or prospective workflow evidence; others do not Literature dominated by retrospective concordance
Inspectability Some outputs localize a finding or quantify a measurement Treatment options require evidence, contraindication, and preference review
Workflow fit Selected studies directly measure process or workload Published concordance does not measure end-to-end workflow value
Physician role Role varies from acquisition and triage to interpretation and monitoring Clinician still needed to assess treatment relevance and disagreement
Failure mode False positives, false negatives, alert delay, automation bias, pathway bottlenecks Incorrect option, missing context, stale knowledge, local nontransferability

Scale, Adoption, and Impact Require Separate Data

The number of FDA-authorized devices is public, but there is no complete public denominator for hospitals using each product, examinations interpreted with assistance, daily user exposure, or model version. Adoption surveys use different definitions and response frames. Vendor customer counts are commercial context, not a census.

Likewise, device authorization and technical performance do not support national estimates of errors prevented, lives saved, burnout reduction, workforce productivity, job satisfaction, or return on investment. Those outcomes require their own denominators, counterfactuals, and economic methods. Local leaders should measure at least:

  1. eligible studies and actual system use;
  2. time to review, escalation, and final action;
  3. false-positive, false-negative, override, and abstention patterns;
  4. patient-pathway and patient-outcome endpoints appropriate to the use;
  5. subgroup performance and access effects;
  6. total implementation, monitoring, downtime, and exit costs;
  7. clinician workload and experience, including work shifted to other staff.

The Lesson for Physicians

Radiology AI success teaches:

1. Partnership must be tested - AI as “second set of eyes,” not replacement - Human makes final decision, AI augments - Result: The combined workflow is its own intervention and can help, fail, or harm

2. Narrow, well-defined tasks succeed first - Detect ICH on CT (clear yes/no, well-validated) vs. recommend cancer treatment (complex, value-laden) - Start with pattern recognition, expand cautiously to reasoning

3. Workflow integration essential - AI must REDUCE physician burden or ADD clear value - Radiology AI speeds critical alerts (value), reduces reading time (efficiency)

4. Transparency builds trust - Radiologists see WHERE AI detected finding (bounding boxes, heatmaps) - Can assess whether AI correct or false positive - Trust develops iteratively as radiologists observe AI performance

5. Physician autonomy preserved - Radiologist can override AI anytime - No forced acceptance of AI recommendations - Autonomy = adoption

Applicability to other specialties:

Applications with radiology-like evaluability: - Dermatology (skin lesion detection) - Pathology (slide analysis, tumor classification) - Ophthalmology (diabetic retinopathy, glaucoma screening) - Cardiology (ECG interpretation, echo measurements)

Applications with harder causal and value questions: - Complex treatment decisions (oncology, critical care) - Diagnostic reasoning with high uncertainty - Goals-of-care discussions - Behavioral health

Current status: Radiology has the largest concentration of FDA-authorized AI devices and a growing body of randomized, prospective, observational, and human-factors studies. The evidence is heterogeneous. The ACR-SIIM Practice Parameter for Imaging AI, approved in May 2026 and scheduled to take effect October 1, 2026, emphasizes governance, inventory, local acceptance testing, monitoring, privacy, and continuous quality improvement (ACR, 2026). This is a model of disciplined implementation, not proof that every radiology AI tool is essential or beneficial.


Part 3: Separating Demonstrated Capability from Forecast

The Hype Cycle: Where We Are

The Gartner labels below are an interpretive lens, not a measured trajectory for all medical AI. Different products and tasks occupy different evidence states at the same time.

Peak of Inflated Expectations (2016-2018): - IBM Watson hype - “AI will replace radiologists in 5 years” (Geoff Hinton, 2016) - broad promises that question answering or image classification would transfer directly to clinical reasoning and replacement

Trough of Disillusionment (2018-2020): - Watson failure - poor reporting and high risk of bias in many COVID-19 imaging models - documented bias and transfer failures in risk prediction and dermatology

Slope of Enlightenment (2020-2024): - Evidence base growing across prospective trials, regulatory authorizations, and postdeployment studies - Understanding what works (narrow tasks, augmentation) vs. does not (complex reasoning, replacement) - Realistic expectations emerging

Possible productivity phase, product by product: - Selected narrow applications can become routine after authorization, local acceptance testing, workflow integration, payment, and monitoring - Other products remain experimental, are withdrawn, or show null patient-pathway results - No calendar establishes when a category reaches a universal plateau

A Three-Tier Framework for Future Claims

Demonstrated: Supported under defined study or authorization conditions, such as a specified screening endpoint, image-acquisition task, simulated benchmark, documentation workflow, or process interval.

Theoretical: Plausible but not generally established, such as persistent burnout reduction, autonomous coordination of multiple care steps, earlier detection across multimodal longitudinal data, or large-scale redistribution of physician work.

Beyond current evidence: General physician replacement, autonomous management of complex multimorbidity across settings, universal human-AI superiority, fixed percentages of future specialty work automated, and guaranteed savings or error reductions.

A forecast becomes more credible when it names the task, action, evidence needed, accountable user, failure mode, and result that would falsify it.

What AI Actually Does Well

Pattern recognition in defined data: - Imaging: Support detection, segmentation, measurement, prioritization, or acquisition for specified findings - Genomics: Classify variants or predict selected molecular or treatment-related endpoints under validated conditions - ECG: Support rhythm classification or defined risk-prediction tasks

Processing vast information quickly: - Literature search: Retrieve and organize candidate studies quickly, with human verification of identity and relevance - Drug interactions: Check 50+ medications simultaneously - Guideline adherence: Flag deviations from protocols

Standardized, repetitive tasks: - Measurements: Tumor volumes, ejection fractions, bone density - Screening: Diabetic retinopathy, cervical cancer, colorectal polyps - Documentation: Generate draft notes from ambient audio

Prediction from large datasets: - Risk scores: Cardiac events, sepsis, readmission (WHEN properly validated) - Resource allocation: Predict ICU capacity, staffing needs - Population health: Identify high-risk patients for outreach

Triage and prioritization: - Critical findings: ICH, PE, pneumothorax alerts - Emergency department: Acuity scoring, fast-track routing - Inbox management: Flag urgent messages, auto-respond to routine

What AI Struggles With

Novel situations: - Rare diseases (limited training data) - Unusual presentations (out-of-distribution) - Example: AI trained on adult pneumonia fails on pediatric pneumonia

Context and nuance: - Patient circumstances omitted or poorly represented in the input (frailty, preferences, goals) - Social determinants (housing, food security, family support) - Example: AI recommends aggressive chemotherapy for 85-year-old who wants comfort care

Causal reasoning: - Correlation ≠ causation - Predictive association alone does not identify a causal mechanism - Example: AI associates peripheral edema with heart failure (correct) but also with warfarin use (spurious correlation in training data)

Uncertainty and ambiguity: - Some systems are overconfident or fail to abstain appropriately - An uncertainty score is useful only if calibrated and actionable in the target workflow - Example: AI predicts 73% mortality but patient survives (confidence score misleading)

Moral and ethical judgment: - An optimization objective does not supply moral legitimacy or resolve competing values - Fairness metrics encode choices but cannot decide which distribution of benefit and burden is just - Example: AI optimizes for hospital profit, conflicts with patient benefit

Adapting to distribution shift: - Performance can change across populations, devices, prevalence, protocols, and time - The magnitude and direction must be measured for the specific system and setting - Example: AI trained in U.S. academic hospitals fails in rural community hospital


Part 4: How Physician Roles May Evolve

Tasks Likely to Receive Continued Automation Support

1. Documentation and administrative work

Current state: - Ambient systems can draft notes and support selected coding, inbox, and administrative tasks - Randomized and observational evidence shows variable effects by product, specialty, user, and workflow

Plausible future state, not a dated prediction: - AI generates comprehensive notes from natural conversation - Automated coding, billing, prior authorization - Inbox management: Routine messages auto-responded - Net time, note quality, cognitive burden, safety, and work shifted to staff must be measured locally

2. Routine image support

Current: Radiologist reads all studies

Future: - Some workflows may route low-risk studies, draft reports, or support screening, subject to evidence, authorization, and local controls - Radiologist focuses on abnormal findings, complex cases - “AI-first read”: Radiologist verifies AI report rather than reading from scratch - Throughput is not automatically beneficial if false alerts, verification, downstream testing, or staff work increase

3. Screening and monitoring support

Current: Physician-dependent screening (diabetic retinopathy, cervical cancer, colonoscopy interpretation)

Future: - Autonomous screening for specified authorized uses, plus AI-assisted screening and monitoring for other tasks - Physician reviews only AI-flagged abnormalities - Home monitoring (wearables + AI) reduces in-person visits for stable chronic disease - Access improvement: Screening available in primary care, community clinics without specialists

Tasks Requiring Continued Physician Leadership

1. The diagnostic process beyond pattern recognition

What AI does: Detects patterns, calculates probabilities What humans do: - Generate comprehensive differential (including rare, unusual) - Strategic hypothesis testing (order tests iteratively, revise based on results) - Contextualizing (patient’s unique history, risk factors, prior diagnoses) - Recognizing when something does not fit (clinical gestalt, “This doesn’t make sense”)

Example: - 45-year-old man with chest pain - AI: 15% probability ACS, 60% GERD, 20% musculoskeletal, 5% other - Physician notes: Patient anxious, recent job loss, family history of early MI, atypical quality of pain - Decision: Admit for cardiac workup despite only 15% AI probability (clinical gestalt > algorithm)

2. Navigating uncertainty and ambiguity

Medicine is irreducibly uncertain. AI provides probabilities, but physician must decide: - Treat presumptively or wait for more data? - Pursue aggressive workup or watchful waiting? - Balance risk of missing diagnosis vs. harm from overtesting?

Example: - Patient with possible appendicitis - AI: 65% probability appendicitis - Question: Operate now (avoid perforation risk) or observe (avoid unnecessary surgery)? - Physician reasoning: Considers patient’s age, comorbidities, reliability for follow-up, OR availability, surgical risk → decision depends on context AI does not fully capture

3. Communication, empathy, and relationship-building

Tasks that require an accountable human relationship: - Conveying bad news with compassion - Eliciting goals of care in end-of-life discussions - Building trust with skeptical patients - Motivating behavior change - Providing comfort in suffering

Why this matters MORE with AI: - As AI handles technical tasks, physician’s humanistic skills become differentiator - Patients seek connection, not just information - Paradox: AI may make medicine MORE human-centered by freeing physicians from administrative burdens

4. Ethical and moral reasoning

Clinical decisions involve values, not just facts: - Balancing autonomy vs. beneficence (patient refuses life-saving treatment) - Allocating scarce resources justly (who gets ICU bed when only one available?) - Defining futile care vs. preserving hope - Navigating conflicts (patient, family, team disagree on goals)

AI can inform decisions by estimating selected outcomes, but a model output does not decide what is right. Moral and professional accountability cannot be delegated merely by placing a system in the workflow.

5. Advocacy for patients

Against systems, insurers, institutions: - Fighting prior authorization denials - Challenging unfair policies - Addressing social determinants of health - Protecting vulnerable patients from exploitation

AI optimizes within systems; physicians advocate to CHANGE systems. This moral dimension does not disappear with AI.


Part 5: An Explicitly Hypothetical Vision of Medicine in 2035

A Day in the Life: Dr. Sarah Chen, Family Physician

Hypothetical future scenario. Dr. Chen, every patient, every product output, every metric, every time saving, and every clinical outcome below is fictional. The scenario organizes design questions; it does not forecast what will exist in 2035 or claim that an AI system caused any benefit. Any deployed version would require product-specific evidence, authorization where applicable, privacy review, workflow validation, monitoring, and accountable human control.

7:00 AM: Pre-clinic preparation - Dr. Chen reviews schedule on tablet. AI has triaged overnight inbox: - Routine prescription refills: Drafted for protocol-based clinician review (15 illustrative requests) - Urgent messages flagged: Chest pain (patient scheduled same-day), abnormal lab (patient called for f/u) - Information requests: AI drafted responses, awaiting Dr. Chen’s review/edit (8 messages) - Illustrative time assumption: 30 minutes saved relative to a fictional 45-minute baseline

8:00 AM: Mr. Garcia, diabetes follow-up - AI displays summary: - Home glucose readings (from continuous monitor): Average 165 mg/dL, 28% time-in-range (target >70%) - Medication adherence: 92% (from smart pill bottle) - Recent A1C: 8.2% (up from 7.4% last visit) - AI flags: “Glucose control worsening despite good adherence. Consider intensification or check for infection, stress, other causes.” - Dr. Chen examines Mr. Garcia, discusses recent job stress (laid off 3 months ago) - Decision: Refers to behavioral health (stress management), adjusts medications, schedules follow-up 1 month - Documentation: Dr. Chen speaks naturally while examining patient. AI generates draft note, she reviews/edits post-visit (2 min vs. 8 min pre-AI)

9:30 AM: Ms. Johnson, fatigue workup - Chief complaint: 3 months fatigue, weight loss - Dr. Chen orders labs, ECG. While talking with patient, results return: - ECG: AI interprets as normal sinus rhythm (Dr. Chen reviews strip, confirms) - Labs: Anemia (Hgb 9.2), low iron, low ferritin - AI generates differential: Iron deficiency anemia → GI bleeding (most likely), nutritional deficiency, menorrhagia - AI suggests: Colonoscopy, upper endoscopy, OB-GYN referral if menorrhagia - Dr. Chen’s decision: Agrees with GI workup, orders colonoscopy. AI assists scheduling (finds earliest available appointment, sends prep instructions tailored to patient’s health literacy level). - AI sends patient education: “Understanding Anemia” (personalized to Ms. Johnson’s reading level, preferred language, prior questions)

11:00 AM: Telemedicine. Mrs. Lee, diabetic retinopathy screening - Mrs. Lee had retinal photos at local pharmacy (AI-interpreted, flagged moderate non-proliferative DR) - Dr. Chen reviews images with Mrs. Lee via video - Refers to ophthalmology (AI auto-schedules appointment, arranges transportation via social services integration) - Intended access improvement: A locally validated, authorized workflow could move an initial screening step closer to the patient; the scenario does not establish access, follow-up completion, or visual outcomes

12:00 PM: Multidisciplinary tumor board - 5 complex cancer cases discussed - For each patient, AI compiles: - Imaging (tumor size, location, metastases) - Pathology (tumor grade, molecular markers) - Genomics (mutations, predicted treatment response) - Prior treatments and responses - AI proposes evidence-based treatment options based on NCCN guidelines, clinical trial matches - Team discusses: Oncologist, surgeon, radiologist, Dr. Chen. AI provides decision support, but team makes final decisions considering patient values, goals, comorbidities, preferences.

2:00 PM: Quality improvement meeting - AI dashboard shows practice metrics: - Vaccination rates: Overall 78%, but disparities noted (Hispanic patients 65%, white patients 85%) - Cancer screening: Colorectal 71%, mammography 82%, cervical 88% - Chronic disease control: Diabetes A1C <8%: 68%, BP <140/90: 72% - Team discusses interventions: - AI identifies barriers (language, transportation) and suggests outreach strategies - AI predicts which patients most likely to respond to reminders, home visits, community health worker engagement - Equity focus: Target resources to close disparities

3:30 PM: Complex patient. Mrs. Thompson - 78-year-old with CHF, COPD, CKD Stage 4, diabetes, polypharmacy (14 medications) - Home monitoring: Weight scale, pulse oximeter, BP cuff (all AI-connected) - Alert yesterday: Weight up 3 lbs in 2 days, oxygen saturation trending down (94% → 91%) - AI flags: “Early CHF decompensation predicted. Recommend medication adjustment.” - Dr. Chen calls Mrs. Thompson: - Exam over video: Increased dyspnea, mild pedal edema - Decision: Increase furosemide dose, schedule in-person visit tomorrow - Scenario result: Outpatient management is attempted. A single fictional case cannot establish that hospitalization was avoided because the counterfactual is unobserved.

4:30 PM: Administrative time - Reviews AI-generated draft notes (minor edits, signs in 15 min vs. 60 min pre-AI) - Prior authorizations: AI drafted justifications with supporting evidence (guidelines, peer-reviewed studies), Dr. Chen reviews/signs (5 min vs. 30 min pre-AI) - AI system alert: Imaging algorithm showing performance drift on recent chest X-rays (sensitivity declining 94% → 88%) - Dr. Chen escalates to IT for investigation, recommends pausing AI auto-reporting until resolved

5:00 PM: Reflection - Illustrative time assumption: 90 minutes across documentation, inbox, and administrative work - Used for: Extra time with complex patients (Mrs. Thompson call), less rushed visits, earlier departure (work-life balance) - AI value-add: Flagged early decompensation (Mrs. Thompson), identified quality improvement targets (vaccination disparities), streamlined routine tasks - Human value-add: Diagnostic reasoning, behavioral health insight, goals-of-care discussions, advocacy

Dr. Chen’s fictional reflection: “The tools reduced selected administrative work and surfaced issues for review. But diagnosis, decisions, and relationships still required accountable clinical work. Whether the workflow reduces burden or improves care must be measured, not presumed.”


Check Your Understanding

The three exercises below are hypothetical. They are decision prompts, not reports of actual patients, product failures, legal outcomes, or institutional policies.

Scenario 1: AI Recommends Treatment You Disagree With

You’re a hospitalist. 72-year-old man with community-acquired pneumonia (CAP), admitted for IV antibiotics.

Fictional AI clinical decision support system (integrated into the EHR) recommends: - Antibiotic: A regimen represented by the exercise as guideline-concordant; actual treatment depends on current guidance, allergy, renal and hepatic function, microbiology, local resistance, and individual risk - Duration: Predicted 5-day course based on “expected time to clinical stability” - Discharge: AI predicts “low risk for treatment failure, suitable for early discharge”

You review patient: - Vital signs improving, but still febrile (101.2°F on Day 3) - CXR: Dense right lower lobe consolidation, small parapneumonic effusion - Patient feels weak, not eating well - Your clinical gestalt: Patient improving but not ready for discharge

AI system generates discharge order set (Day 4) with plan for oral antibiotics at home.

Question 1: Do you follow AI recommendation for Day 4 discharge?

Answer: NO. Override AI recommendation.

Clinical reasoning: - AI prediction based on population data (average CAP patient) - Individual patient not yet clinically stable: - Persistent fever (not afebrile for 24+ hours) - Poor oral intake (risk of dehydration, medication non-adherence) - Small effusion (could worsen) - Clinical judgment: Patient needs 1-2 more days of IV antibiotics, monitored environment

Documentation (critical for medico-legal protection): > “AI clinical decision support recommended discharge Day 4 with transition to oral antibiotics. However, patient not yet clinically stable: persistent fever (101.2°F), poor oral intake, small parapneumonic effusion on CXR. Clinical judgment: Continue IV antibiotics, reassess in 24-48 hours. Override AI recommendation in favor of individualized care plan.”

Question 2: What if hospital administration pressures you to follow AI (for cost savings)?

Response: > “I understand AI suggests discharge, but patient is not clinically ready. Following AI recommendation would risk treatment failure, readmission, and patient harm. My clinical judgment is that this patient needs continued hospitalization. If administration disagrees, I’m happy to discuss with medical director, but I cannot discharge a patient I believe unsafe to discharge, even if AI says otherwise. My medical license and patient’s well-being take precedence over cost savings.”

Key principle: The treating clinician remains accountable for the clinician’s own decision and documentation, while institutional, developer, and other duties remain fact- and jurisdiction-specific. AI is decision support, not an independent legal authority.

Scenario 2: AI Fails to Detect Critical Finding

You’re an emergency physician. 55-year-old woman presents with headache (sudden onset, “worst headache of my life”), nausea, photophobia.

You order: CT head non-contrast

Fictional imaging triage system integrated into PACS analyzes the CT and returns: - AI finding: “No intracranial hemorrhage detected” - AI priority: “Routine” (not flagged as critical)

You review CT personally: See subtle hyperdensity in left Sylvian fissure concerning for subarachnoid hemorrhage (SAH)

You order: CT angiography (confirms ruptured aneurysm), neurosurgery consult, patient to OR for clipping

Question 1: What went wrong with AI?

AI failure mode: - Small, subtle SAH (difficult even for humans to detect) - AI trained on larger, obvious hemorrhages - False negative (AI missed diagnosis)

Why you caught it: - Clinical suspicion (“worst headache of life” = SAH until proven otherwise) - Did not rely solely on AI, reviewed images personally - Key: Maintained independent clinical reasoning, did not outsource thinking to algorithm

Question 2: What should you do after patient stabilized?

Report AI failure through institutional channels: 1. Incident report: Document AI false negative, patient outcome, your corrective action 2. Quality committee: Escalate to AI governance committee for investigation 3. Root cause analysis: Was this one-off failure or systematic issue? 4. Potential actions: - Retrain AI on subtle SAH cases - Adjust sensitivity/specificity threshold (accept more false positives to reduce false negatives) - Add clinical context (AI should flag “worst headache of life” + negative CT for physician review regardless of AI finding)

Question 3: Are you liable if you missed AI’s error (and patient harmed)?

Liability cannot be determined from this vignette. Relevant questions include the applicable standard of care, the intended use and warnings, the clinician’s and radiologist’s roles, institutional workflow, what each party knew, causation, and jurisdiction.

If the images were independently reviewed and the finding was still missed: The result alone does not establish negligence. Standard of care, reasonableness, causation, and expert evidence remain necessary.

If the workflow required independent review but the output was accepted without it: That fact could be important, but liability still cannot be declared categorically. The intended user, assigned review duties, labeling, policy, and causal pathway matter.

Lesson: AI is decision support, not decision-maker. Always maintain independent clinical reasoning, especially for high-stakes decisions.

Scenario 3: Patient Refuses AI Involvement in Care

You’re a primary care physician. 68-year-old woman scheduled for screening mammography.

Patient: “I don’t want any AI reading my mammogram. I heard AI makes mistakes, has biases, and I don’t trust it. I want a human doctor to read it.”

Your radiology department: Uses AI-assisted mammography interpretation (AI flags suspicious lesions, radiologist makes final determination). Standard workflow, all screening mammograms use AI.

Question 1: Can patient refuse AI?

Legal answer: Unclear, evolving area.

Patient autonomy argument: - Informed consent principle: Patients have right to refuse medical interventions - If AI is “intervention,” patient can refuse - Similar to refusing specific surgeon, medication, test

Healthcare system argument: - AI is behind-the-scenes tool (like digital image processing, CAD systems) - Not “treatment” requiring consent - Operational workflow decision, not patient choice

Practical answer: Most hospitals have not yet established formal policies on patient consent for AI involvement. Physicians should:

  1. Explain AI role (educate patient): > “I understand your concern. In this hypothetical workflow, the AI highlights findings and changes how examinations are assigned, while a radiologist remains responsible for the final report. A large randomized Swedish screening study found higher cancer detection and lower reading workload for one AI-supported configuration, but that result does not establish the performance of our local product. The radiology team should be able to explain the exact system, evidence, monitoring, and available alternatives.”

  2. Validate concerns, provide evidence: > “Your concern about bias is fair. Some systems perform differently across populations or settings. FDA authorization would establish only the authorized intended use, not local equity or outcome benefit. The department should verify the exact authorization record, conduct local acceptance testing, and monitor performance and access across relevant groups. If you would like, a radiologist or patient advocate can discuss the workflow with you.”

  3. Offer compromise if patient still refuses: > “If you prefer, I can request that the radiologist read your mammogram without AI assistance. It may take longer to schedule (fewer appointments available for non-AI reads), and the radiologist won’t have AI’s second opinion, but it’s your choice. I want you to feel comfortable.”

Question 2: What if patient still refuses and hospital says “AI is non-optional”?

Ethical tension: - Patient autonomy (right to refuse) - Healthcare system efficiency (workflow standardization)

Physician role: Advocate for patient

Options: 1. Request exception: Ask radiology to accommodate patient preference (AI-free read) 2. External referral: Refer patient to imaging center without AI 3. Escalate: If hospital refuses accommodation, involve patient advocate, ethics committee

Key principle: Patient autonomy should be respected when feasible. If refusal would cause substantial delay or unavailability of care, discuss trade-offs with patient. Do not force AI on refusing patient without serious justification.


Part 6: A Call to Action

For Individual Physicians

1. Invest in AI literacy - Understand basics: How AI works, what it can/cannot do, when to trust vs. question - Resources: Online courses (Coursera, edX), professional society webinars (AMA, specialty societies), journal articles - Goal: Be informed consumer of AI, not passive recipient

2. Engage with AI at your institution - Join AI governance committees (ensure physician voice in decisions) - Participate in pilots, provide honest feedback - Advocate for evidence, transparency, equity monitoring

3. Maintain core clinical skills - Do not outsource thinking to algorithms - Practice clinical reasoning independent of AI (“What would I do if system crashed?”) - Teach trainees to think first, use AI second

4. Strengthen patient relationships - Use AI time savings for deeper engagement, not just volume - Discuss AI use transparently (“I’m using AI to help interpret your ECG, but I review it personally”) - Reaffirm commitment and responsibility (“I’m accountable for your care, not the algorithm”)

5. Advocate for patient-centered AI - Demand validation before institutional adoption - Insist on equity monitoring (performance across demographics) - Oppose AI that harms patients, even if profitable for hospital

For Medical Educators

1. Integrate AI into curriculum - Medical school: AI foundations, ethics, data science - Residency: Specialty-specific AI, hands-on training with real tools - CME: Continuous updates (AI evolves rapidly)

2. Teach AI-augmented clinical reasoning - How to interpret AI outputs (probabilities, confidence scores) - When to trust vs. override AI - Document reasoning when disagreeing with AI

3. Preserve core clinical skills - Physical exam, history-taking, bedside teaching remain essential - Do not let trainees become AI-dependent (skills atrophy)

4. Model ethical AI use - Show trainees how to question AI, advocate for patients, prioritize human judgment - Discuss failures openly (learn from mistakes)

For Healthcare Leaders

1. Establish robust AI governance - Physician-led oversight committees (not just IT, administrators, vendors) - Mandatory validation on local population before deployment - Continuous monitoring (performance, equity, safety)

2. Align incentives with patient benefit - Do not adopt AI solely for cost-cutting - Require evidence of clinical benefit (outcomes, not just efficiency) - Invest in AI that reduces disparities, improves care

3. Support workforce adaptation - Training, protected time for learning - Workflow redesign (optimize human-AI collaboration) - Address resistance constructively (engage physicians as partners, not impose top-down)

4. Ensure transparency and accountability - Patients informed when AI used - Clear documentation of AI recommendations, physician decisions - Incident reporting for AI errors, near-misses - Liability frameworks (who’s responsible when AI errs?)


Questions About the Future of AI in Medicine

Will AI replace physicians?

No evidence supports replacing physicians as a profession. Whether a physician-AI team outperforms either member alone depends on the task, interface, training, review conditions, and whether their errors are complementary.

Economic history also pushes back on a simple “AI shrinks the clinical workforce” story. In a Perspective, Khullar (2026) argues that efficiency gains can expand demand (Jevons paradox), that work is not a fixed lump (so new care modes can create clinician demand), and that automating tasks is not the same as automating jobs when high-stakes care still depends on interlocking human skills (O-ring / tasks ≠ skills). That is economic reasoning and historical analogy: not a forecast that every specialty’s headcount will rise, or that no roles will change or disappear.

How much time does ambient documentation AI save?

Randomized and observational studies report product- and workflow-specific changes in documentation time, cognitive burden, and note quality. No single daily time-saving estimate applies across products, specialties, institutions, or users.

Does AI improve diagnostic accuracy?

Some systems improve specified technical, workflow, or screening endpoints under defined conditions. In the randomized MASAI screening study, AI-supported mammography increased cancer detection and reduced reading workload, while other human-AI configurations have shown no benefit or caused new errors.

What happened to IBM Watson for Oncology?

Watson for Oncology was evaluated mainly through concordance with multidisciplinary recommendations, and concordance varied by cancer, stage, population, and setting. Those studies did not establish improved patient outcomes, illustrating why agreement, clinical utility, and patient benefit must remain separate claims.

Can LLMs reason like physicians?

No single study establishes that an LLM reasons like a physician. Performance changes when question format, answer choices, context, or workflow changes, so benchmark accuracy should not be treated as proof of transferable clinical reasoning.

Conclusion: Partnership, Not Replacement

The deepest lesson from the evidence across foundations, applications, ethics, failures, and future directions:

AI is a tool. Medicine is a profession. Tools serve professions, not vice versa.

AI can change healthcare by supporting earlier detection, treatment selection, access, error reduction, and relief from selected administrative burdens. Some of these benefits have been demonstrated for specific systems and endpoints; none should be generalized without matched evidence.

But AI cannot replace what makes medicine meaningful: human connection, moral commitment to alleviating suffering, presence in patients’ most vulnerable moments. These are irreplaceable. Not because AI lacks technological sophistication, but because they’re fundamentally human.

The strongest direction is conditional partnership. Physician + AI works best when the workflow makes verification easier, not harder. Physicians practicing in increasingly AI-enabled settings will need both: - Technological competence: Use AI effectively, interpret critically, recognize limitations - Humanistic excellence: Communicate with empathy, reason ethically, build trust, advocate fiercely

This is harder than pure technical or pure humanistic medicine. It requires analytical rigor AND compassionate presence, data fluency AND narrative understanding, algorithmic precision AND moral wisdom.

Medicine can do this. Clinicians enter medicine to help people. AI is another tool toward that goal: powerful, imperfect, and requiring wise use.

The challenge: Integrate AI without losing medicine’s soul. Embrace efficiency without sacrificing empathy. Leverage algorithms without abdicating judgment.

The opportunity: Medicine has incorporated antibiotics, imaging, genomics, transplantation, and many other technologies while preserving healing’s humanistic core. AI should be judged by whether it strengthens that work under real clinical conditions.

The responsibility: Shape this future actively. Physicians can lead institutions, advocate in communities, teach the next generation, demand better evidence from vendors and policymakers, and remain accountable to patients.

The vision: Future physicians should be able to say: “AI improved defined aspects of medicine, and the profession remained accountable to evidence, equity, patient choice, and human care.”

That’s the future worth working toward.

The physician-AI partnership must be built for patients, professional integrity, and the future of healing.