History of AI in Medicine: From MYCIN to Foundation Models

In a 1979 blinded evaluation, expert reviewers judged 65% of MYCIN’s antimicrobial recommendations acceptable, a result comparable with the human prescribers evaluated in the same study (Yu et al., 1979). MYCIN was not deployed in routine patient care. The distinction between controlled performance and clinical adoption has remained central across every later generation of medical AI.

Learning Objectives

After reading this chapter, clinicians should be able to:

  • identify major transitions from expert systems to machine learning and foundation models;
  • distinguish demonstrations, concordance studies, regulatory authorization, workflow evidence, and patient outcomes;
  • explain why adoption and reimbursement do not establish clinical benefit;
  • recognize recurring failures involving transportability, workflow, maintenance, and overclaiming; and
  • apply historical lessons to current clinical AI evaluation.
  • Early systems demonstrated narrow reasoning and behavioral simulation without establishing clinical treatment benefit.
  • MYCIN showed that expert-reviewed recommendations could perform well in a controlled evaluation while remaining undeployed.
  • Computer-aided detection showed that widespread adoption can precede evidence of improved diagnostic performance.
  • Autonomous diabetic-retinopathy screening paired a narrow intended use with prospective evidence and a defined regulatory pathway.
  • IBM Watson for Oncology showed why marketing, literature processing, and concordance must not be treated as patient-outcome evidence.
  • Foundation models broaden task coverage, but they do not remove the need for intended-use, workflow, and outcome evaluation.

Milestones and Evidence Boundaries

Period Development Evidence boundary
1950s–1960s Symbolic AI and early conversational programs Behavioral demonstrations were not clinical efficacy studies
1960s–1980s DENDRAL, MYCIN, INTERNIST-I, and other expert systems Hand-built knowledge worked in bounded tasks but created maintenance and workflow problems
1990s–2000s Statistical learning and computer-aided detection Adoption outpaced proof of improved clinical outcomes in some applications
2010s Deep learning for imaging and autonomous screening External and prospective validation became more prominent, but generalizability remained task-specific
2020s Foundation models, multimodal systems, and agents Breadth and fluency increased faster than evidence for autonomous clinical use

Early Symbolic and Conversational Systems

DENDRAL used expert knowledge and search to infer molecular structures from mass-spectrometry data. Its narrow scientific success helped establish the expert-system approach (Lindsay et al., 1993).

ELIZA used pattern matching to simulate conversation. The published program demonstrated how a simple system could produce language that users interpreted as meaningful, not that it understood or treated mental illness (Weizenbaum, 1966). PARRY later simulated aspects of paranoid thought and language. Psychiatrists were asked to distinguish program transcripts from patient transcripts, an experiment in behavioral simulation rather than diagnosis or therapy (Colby et al., 1972).

Human attribution of understanding appeared decades before modern language models. The historical lesson is to evaluate what a system actually did, not the mental state that fluent output appears to imply.

Expert Systems and MYCIN

MYCIN used a hand-built rule base to recommend antimicrobial therapy for serious bacterial infections (Shortliffe et al., 1975). Its blinded evaluation compared the acceptability of its recommendations with recommendations from human prescribers (Yu et al., 1979). It did not test patient outcomes or routine deployment.

Historical accounts identify several translation barriers:

  • manual data entry outside normal workflow;
  • the burden of maintaining a large rule base as evidence and resistance patterns changed;
  • uncertainty about accountability and regulatory fit;
  • limited computing and integration infrastructure; and
  • lack of prospective evidence in actual care.

The evidence does not support assigning a precise causal share to each barrier. MYCIN’s durable lesson is that technical evaluation, clinical workflow, governance, maintenance, and outcome evidence are separate requirements.

Other systems occupied different niches. INTERNIST-I evaluated computer-assisted diagnosis across internal medicine (Miller et al., 1982). DXplain was developed as a differential-diagnosis support system (Barnett et al., 1987). Their existence should not be rewritten as evidence of autonomous diagnosis or improved patient outcomes.

Data-Driven Learning and Computer-Aided Detection

The shift from manually encoded rules to statistical learning allowed models to estimate patterns from digitized records and images. It also introduced new dependence on dataset construction, labels, and transportability.

Mammography computer-aided detection illustrates the difference between adoption and benefit. In a large observational analysis, CAD use was associated with increased recall and no improvement in cancer detection or diagnostic accuracy (Lehman et al., 2015).

A human-plus-AI configuration is an intervention that requires evidence. It cannot inherit benefit from either component’s standalone performance.

Google Flu Trends provides a related data lesson. The original system used search queries to estimate influenza activity (Ginsberg et al., 2009). It later substantially overestimated influenza-like illness, and Lazer and colleagues identified changing search behavior, platform dynamics, and modeling choices as important concerns (Lazer et al., 2014).

Deep Learning and Narrow Clinical Authorization

Deep learning increased performance on image-recognition tasks and reduced the need for manually designed features. The 2016 diabetic-retinopathy study by Gulshan and colleagues evaluated a deep neural network on retinal image datasets and reported performance at specified operating points (Gulshan et al., 2016). That research model was not itself an FDA authorization.

In 2018, FDA granted De Novo authorization to IDx-DR, now LumineticsCore, for a specified autonomous diabetic-retinopathy screening use (FDA DEN180001). The prospective pivotal study evaluated sensitivity, specificity, imageability, and workflow under defined conditions, not treatment, vision preservation, or every retinal disease (Abràmoff et al., 2018).

This sequence matters: research performance, product configuration, prospective evidence, intended use, and authorization were distinct records.

IBM Watson for Oncology

Watson for Oncology was promoted as a literature-informed cancer decision-support system. Reporting based on internal documents described unsafe and incorrect recommendations in some training cases (Ross and Swetlitz, 2018). Published studies later reported variable concordance with local treatment decisions across cancers and settings (Jie et al., 2021).

Concordance is not an outcome. It can reflect local practice, the reference panel, available therapies, and how disagreement is adjudicated. These studies did not establish improved survival, toxicity, quality of life, or cost.

Watson’s history is evidence against transferring a broad technical demonstration into an unbounded clinical promise.

Foundation Models and Clinical Breadth

Foundation models expanded the number of tasks addressable through prompting or adaptation. Med-PaLM, for example, was evaluated on medical question-answering datasets and clinician-rated long-form answers (Singhal et al., 2023). Those evaluations established performance under specified conditions, not independent diagnosis, treatment, or longitudinal care.

Later systems added images, audio, retrieval, tools, and multi-step action. Each addition changes the intervention and its failure modes. Benchmark breadth does not eliminate the need to specify the model version, input, prompt, tools, user, authority, comparator, and endpoint.

Skill, Reliance, and Human Performance

Three outcomes should remain distinct:

  • Automation bias: assistance changes a decision during the assisted task, particularly when advice is wrong.
  • Deskilling: a previously acquired skill deteriorates after repeated reliance.
  • Never-skilling: a learner fails to acquire independent competence.

Incorrect recommendations have reduced performance in assisted imaging and vignette studies, supporting concern about automation bias. Evidence for inevitable longitudinal loss of physician expertise is substantially weaker. A 2025 observational endoscopy study reported lower non-AI adenoma detection after exposure to AI-assisted colonoscopy, but later prospective studies did not reproduce an inevitable decline (Budzyń et al., 2025; Okumura et al., 2026; Pedersen et al., 2026).

Current evidence supports measuring independent and assisted performance separately. It does not support assuming that all clinical AI use causes deskilling.

Recurring Translation Tests

Across eras, the same questions recur:

  1. What exact task was studied?
  2. Was the comparison against clinicians, usual care, or a reference standard?
  3. Did the study measure technical performance, workflow, or patient outcomes?
  4. Does the evaluated configuration match the deployed product and version?
  5. Were the population, site, devices, and prevalence relevant to deployment?
  6. Who maintains knowledge, data, thresholds, and interfaces after change?
  7. Can users detect failure, abstain, escalate, and recover during downtime?

Clinical Conclusions

  • Medical AI history contains important successes, but technical performance has repeatedly been overextended into clinical claims.
  • Adoption, concordance, regulatory authorization, and patient benefit are different forms of evidence.
  • Narrow intended uses are easier to validate and govern than general clinical promises.
  • Human-AI performance depends on the configured workflow, not a universal rule that assistance helps.
  • New architectures change capabilities, but they do not remove the requirements for source fidelity, clinical evaluation, accountability, and monitoring.