The Future of AI in Medicine: Physician Roles and Evidence

The future of clinical AI will be determined task by task. Evidence supports selected screening, interpretation, documentation, and decision-support uses, but no result justifies a general claim about physician replacement or universal collaboration benefit. Each system must remain bound to its intended use, evidence, workflow, accountability, and monitoring plan.

Learning Objectives

After reading this chapter, clinicians will be able to:

  • Distinguish demonstrated capabilities from theoretical and unsupported forecasts
  • Evaluate human-AI collaboration as a measurable intervention
  • Assess claims about task automation and workforce replacement
  • Identify durable physician responsibilities in AI-enabled care
  • Define evidence and governance conditions for future adoption
  • No universal forecast: Effects will differ across tasks, specialties, institutions, and patient populations.
  • Collaboration is conditional: Human-AI performance depends on interface, expertise, training, review conditions, and complementary error patterns.
  • Evidence must match the claim: Technical performance, workflow effect, patient benefit, equity, cost, and workforce impact require different studies.
  • Roles will change: Selection, supervision, contestability, communication, monitoring, and incident review are clinical work.
  • Skill preservation is measurable: Automation bias is documented; generalized longitudinal deskilling remains insufficiently studied.
  • Governance is continuous: Version control, local acceptance testing, monitoring, and decommissioning do not end at procurement.

History Sets the Boundary

Watson for Oncology was promoted ahead of sufficient clinical-outcome evidence. Internal examples included unsafe or incorrect recommendations, and concordance studies did not establish improved patient outcomes. Its history demonstrates that literature processing, agreement with an expert panel, and clinical benefit are separate claims. The source review is maintained in IBM Watson for Oncology.

Autonomous diabetic retinopathy screening provides a different example. FDA granted De Novo authorization for IDx-DR, now LumineticsCore, for a specified population, setting, camera workflow, and diagnostic endpoint (FDA DEN180001). The prospective pivotal study evaluated sensitivity and specificity for more-than-mild diabetic retinopathy, not every eye disease or long-term vision outcomes (Abràmoff et al., 2018).

The contrast is not failure versus success in the abstract. It is an unbounded clinical promise versus a specified intended use supported by matched evidence.

Radiology: Progress Without a Universal Partnership Claim

Radiology has the largest concentration of FDA-authorized AI devices, but device count is not a measure of clinical benefit. Evidence varies across detection, triage, acquisition, interpretation, workflow, and outcomes. The specialty’s current evidence is reviewed in Diagnostic Imaging and Radiology.

Earlier mammography computer-aided detection demonstrates why a human plus AI configuration is not automatically superior. In a large observational analysis, CAD use was associated with increased recall and no improvement in cancer detection or diagnostic accuracy (Lehman et al., 2015). Newer systems require their own evaluation and should not inherit benefit from the category name.

The ACR-SIIM Practice Parameter for Imaging AI emphasizes governance, inventory, local acceptance testing, monitoring, privacy, and continuous quality improvement. It was approved in May 2026 and is scheduled to take effect October 1, 2026. These controls are necessary precisely because performance is product-, task-, version-, site-, and workflow-specific.

A Three-Tier Framework for Future Claims

Demonstrated

Capabilities supported under defined study or authorization conditions include:

  • narrow autonomous diagnostic outputs for specified intended uses;
  • image-acquisition guidance for limited examinations;
  • language-model performance on defined question-answering or summarization tasks;
  • task completion by agents in simulated or benchmark environments;
  • workflow effects for specific documentation and triage systems.

These findings should retain their population, comparator, endpoint, and version.

Theoretical

Plausible but not generally established possibilities include:

  • large-scale redistribution of routine work from clinicians to AI systems;
  • persistent reduction in burnout through automation;
  • improved continuity through longitudinal multimodal assistants;
  • earlier detection across multiple data streams;
  • clinical agents coordinating several steps of care with bounded autonomy.

These hypotheses need comparative prospective studies, economic analysis, equity assessment, and monitoring.

Beyond current evidence

Claims not supported as general realities include:

  • replacement of physicians as a profession;
  • autonomous management of complex, multimorbid patients across settings;
  • universal superiority of physician plus AI over either alone;
  • fixed percentages of specialty work that will be automated by a future year;
  • guaranteed time savings, cost savings, or error reductions across institutions;
  • artificial systems assuming the moral, fiduciary, and legal role of a physician.

Human-AI Performance Is a Study Question

Human-AI collaboration should be evaluated as its own intervention. A model can outperform a clinician alone while degrading the clinician’s performance when combined, or the reverse. Important variables include:

  • whether the human forms an independent assessment first;
  • how uncertainty and abstention are shown;
  • whether AI errors are correlated with human errors;
  • user expertise and training;
  • time available for review;
  • alert prevalence and false-positive burden;
  • whether disagreement is easy to investigate;
  • whether the human can stop or reverse the action.

A nominal human in the loop is not a safety control unless that person can detect error, has authority to intervene, and is given adequate time and information.

Potential Changes to Physician Work

Forecasts are more defensible at the task level than the occupation level.

Tasks likely to receive continued automation support

  • drafting and structuring documentation
  • retrieving and organizing records or literature
  • measurements and segmentation in images
  • preliminary prioritization and routing
  • repetitive administrative preparation
  • surveillance for prespecified signals

Tasks likely to require continued physician leadership

  • resolving uncertainty and conflicting evidence
  • integrating patient goals, comorbidity, and social context
  • choosing when a model is relevant or out of scope
  • communicating risk and negotiating care decisions
  • supervising high-consequence actions
  • identifying system-level harm and triggering remediation
  • establishing professional and organizational accountability

The distinction is not uniquely human intuition versus machine calculation. Many tasks combine both. What matters is whether responsibility, evidence, and authority are explicit.

Skills, Automation Bias, and Independent Competence

Automation bias and inappropriate reliance are documented human-factors risks. Direct evidence that routine AI exposure causes generalized, irreversible physician deskilling is more limited. Training programs should therefore measure rather than presume:

  1. unaided performance when independent competence is required;
  2. assisted performance in the actual workflow;
  3. ability to recognize AI failure and out-of-scope use;
  4. quality of disagreement resolution;
  5. performance after model or interface changes.

Lower unaided performance observed after AI use does not by itself prove causality. Case mix, baseline ability, supervision, fatigue, and assessment validity also require investigation.

Patient Relationship, Consent, and Contestability

Patient expectations will vary. Some will prefer AI-supported access or faster results; others will object to data use or algorithmic involvement. Whether specific disclosure or consent is legally required depends on jurisdiction, intervention, data flow, institutional policy, and standard clinical consent principles.

Regardless of the legal minimum, institutions should be able to explain:

  • what the system does and does not do;
  • whether it generates, prioritizes, recommends, or acts;
  • who reviews its output;
  • what data it uses and retains;
  • how errors or complaints are handled;
  • whether an alternative pathway is available.

Contestability matters. Patients and clinicians need a credible way to challenge an output, obtain human review, correct records, and seek redress.

Governance Conditions for Future Adoption

Future systems should not be adopted because they resemble an imagined endpoint. They should meet current, auditable conditions:

Domain Minimum condition
Scope Exact intended use and prohibited uses
Evidence Design and endpoint match the claimed benefit
Regulation Current status verified from the official record
Workflow Accountable human authority and tested escalation
Technical Version, inputs, prompts, retrieval, and tools recorded
Safety Failure modes, downtime, rollback, and stop rules tested
Equity Access and subgroup effects measured with uncertainty
Monitoring Use, overrides, performance, incidents, and drift reviewed
Economics Local costs and benefits measured rather than assumed
Exit Contractual and operational decommissioning plan

Liability is not governed by a universal rule that the clinician always bears responsibility or that the developer does. Allocation depends on jurisdiction and facts and may involve multiple actors (Mello & Guha, 2024).

Decision Questions for Leaders

Before approving a future-facing system, leaders should answer:

  1. What patient or workforce problem is being addressed, and what is the baseline?
  2. What exact action will change because of the system?
  3. What evidence supports that action in the intended setting?
  4. What evidence is missing, and how will it be generated?
  5. Who can stop an unsafe action or deployment?
  6. What result would justify expansion, restriction, or withdrawal?

This framework replaces fictional success stories with testable conditions.

Conclusion

The most credible future for clinical AI is neither automatic replacement nor guaranteed partnership. It is a sequence of bounded systems that earn wider use through evidence, useful workflow design, professional accountability, and monitoring.

Medicine should adopt AI where it improves a defined clinical or operational outcome, preserve independent judgment where failure is consequential, and withdraw systems when evidence or controls no longer support use.