History of AI in Medicine: From MYCIN to Foundation Models

In 1979, eight independent experts reviewed antimicrobial regimens proposed by Stanford’s MYCIN and nine human prescribers for ten meningitis cases. They rated 65% of MYCIN’s regimens acceptable, compared with 42.5%–62.5% for the five faculty specialists, but MYCIN never entered routine patient care (Yu et al., 1979). That distinction between controlled task performance and clinical adoption is one of the central lessons of 70 years of AI in medicine.

Learning Objectives

After reading this chapter, you will be able to:

  • Distinguish genuine clinical breakthroughs from recurring hype cycles (MYCIN, IBM Watson, diagnostic AI)
  • Recognize why technical excellence does not guarantee clinical adoption
  • Identify patterns separating successful medical AI deployments from research prototypes
  • Understand FDA regulation evolution and current frameworks
  • Evaluate whether today’s foundation models represent a fundamental departure from prior medical AI
  • Apply historical lessons about clinical integration, medical liability, and physician trust

The Clinical Context: AI has experienced 70 years of boom-bust cycles, from 1970s expert systems to today’s foundation models, with each wave promising to reshape medicine but delivering narrow, often fragile applications. Understanding this history is essential for distinguishing genuine breakthroughs from marketing hype and avoiding expensive implementation failures.

The Cautionary Tales:

  • MYCIN (1972–1979): Stanford’s expert system inferred likely organisms and recommended antimicrobials (Shortliffe et al., 1975). In a blinded evaluation of ten meningitis cases, experts rated 65% of MYCIN’s regimens acceptable, compared with 42.5%–62.5% for the five faculty specialists (Yu et al., 1979). Yet MYCIN was never used in routine clinical care. Contemporary accounts identify technical, workflow, governance, and accountability barriers, but the evidence does not assign precise causal shares to each barrier. Key lesson: Controlled performance does not establish adoption or patient benefit.

  • IBM Watson for Oncology (2012–2018): After Watson’s Jeopardy! success, IBM promoted oncology decision support across multiple hospitals. Investigative reporting documented unsafe recommendations and the cancellation of MD Anderson’s $62 million implementation (Ross & Swetlitz, 2018; Strickland, 2019). A meta-analysis of nine nonrandomized studies measured concordance with multidisciplinary recommendations, not patient outcomes, and found variation across cancers and settings (Jie et al., 2021). Key lesson: agreement with clinicians is not proof of clinical utility.

  • Google Flu Trends (2008-2015): Published in Nature (Ginsberg et al., 2009), GFT used search queries to predict influenza activity 1-2 weeks ahead of CDC surveillance. After initial success and widespread adoption, GFT failed dramatically in the 2012-2013 flu season, predicting more than double the proportion of influenza-like illness doctor visits that the CDC reported, and it ran high for 100 of 108 weeks between August 2011 and September 2013 (Lazer et al., 2014). Why? Algorithm updates, changing search behavior, overfitting to spurious correlations. Discontinued 2015. Key lesson: Black-box algorithms fail when they cannot adapt to distribution shift.

The Pattern That Repeats: 1. Breakthrough technology demonstrates impressive capabilities in controlled settings 2. Overpromising: “AI will transform medicine / replace radiologists / eliminate diagnostic errors” 3. Pilot studies succeed with carefully curated datasets 4. Reality: Deployment reveals liability concerns, workflow disruption, integration challenges, trust gaps 5. Disillusionment when technology falls short of marketing claims 6. Eventual integration into narrow, well-defined applications (if evidence supports it)

What Actually Works in Clinical Medicine:

Diabetic retinopathy screening (IDx-DR, FDA De Novo authorization DEN180001 in 2018) (FDA De Novo summary; Abràmoff et al., 2018) - Specific, well-defined task with clear ground truth - Prospective validation in real clinical settings - Autonomous operation without physician interpretation - Currently deployed in primary care, endocrinology clinics

Computer-aided detection (CAD) in mammography (Lehman et al., 2015) - Augments radiologist interpretation, does not replace it - Integrated into radiology workflow - In a large observational comparison of 625,625 digital screening examinations, CAD did not improve sensitivity, specificity, or cancer detection rate - Illustrates that adoption and reimbursement can outpace evidence of benefit

Sepsis prediction alerts (Epic Sepsis Model, others) (Wong et al., 2021) - High-stakes problem with clear intervention pathway - Alerts clinicians to deteriorating patients - But: High false positive rates remain problematic - Ongoing debate about clinical benefit vs. alert fatigue

AI-assisted pathology (Paige Prostate, FDA De Novo authorization DEN200080 in 2021) (FDA decision letter; Raciti et al., 2020) - Flags suspicious regions for pathologist review - Reduces interpretation time - Maintains human-in-the-loop oversight

What Does Not Work (Documented Failures):

IBM Watson for Oncology - Unsafe recommendations, poor real-world performance (Ross & Swetlitz, 2018) Epic Sepsis Model at Michigan Medicine - 33% sensitivity (missed 67% of sepsis cases), 12% PPV (88% false positives) (Wong et al., 2021) Skin cancer apps lacking validation - Many show poor performance outside training distributions (Freeman et al., 2020) Autonomous diagnostic systems without appropriate workflow and contingency planning - Authorization is indication-specific, and safe implementation still requires defined responsibilities, escalation pathways, and handling of ungradable or out-of-scope cases

The Critical Insight for Physicians:

Technical metrics (accuracy, AUC-ROC, sensitivity, specificity) do not predict clinical utility. The hardest problems deploying medical AI are: - Medical liability: Who’s responsible when AI fails? - FDA regulation: Which devices require clearance? How much evidence is enough? - Clinical workflow integration: Does this fit how we actually practice? - Physician trust: Will clinicians follow AI recommendations? - Patient acceptance: Are patients comfortable with algorithmic decisions?

Why This Time Might Be Different:

  • Data availability: EHRs, genomics, imaging archives, wearables, multi-omic datasets
  • Computational power: Cloud computing, GPUs, TPUs make complex models feasible
  • Algorithmic breakthroughs: Transfer learning, foundation models (GPT-4, Med-PaLM 2), few-shot learning
  • Regulatory maturity: FDA has frameworks for AI/ML-based medical devices
  • Clinical acceptance: Younger physicians trained alongside AI tools show higher adoption

Yet fundamental challenges persist: Explainability (black-box models), fairness (algorithmic bias), reliability (distribution shift), deployment barriers, and the irreducible complexity of clinical medicine.

The Clinical Bottom Line:

Be skeptical of vendor claims. Match evidence requirements to the claim: retrospective validation can assess model performance, prospective workflow studies can assess use in practice, and comparative studies with patient-relevant endpoints are needed for outcome-benefit claims. Prioritize patient safety over efficiency. Liability remains fact-specific and jurisdiction-specific; FDA authorization does not determine civil liability. Start with narrow, well-defined problems rather than general diagnostic systems. Center physician and patient perspectives in AI development and deployment.

History shows that many promising medical AI systems do not progress from technical validation to durable clinical use. Learning where translation fails matters as much as celebrating successful deployments.

Introduction

Artificial intelligence in medicine is not new. The field has experienced multiple waves of excitement and bitter disillusionment over seven decades. Each cycle promised to reshape clinical practice. Each fell short. The scale of recent activity is unprecedented: AI-related medical publications increased 36-fold between 2000 and 2022, from 1,614 to over 58,000 annually (Shi et al., 2023).

So why should physicians believe that this time is different?

Understanding AI’s history in medicine is not just academic curiosity. It’s essential for navigating today’s hype, identifying genuinely transformative applications, and avoiding expensive failures that harm patients or waste resources. The patterns repeat: breathless promises, pilot studies that look impressive, deployment challenges nobody anticipated, and eventual disillusionment when technology does not match marketing.

But history also reveals what works. Successful medical AI applications often share common traits: they solve specific, well-defined clinical problems; they are evaluated in the intended population and workflow; and they define what the clinician must do when the system is wrong, unavailable, or out of scope. Prospective evaluation strengthens claims about use in practice, while patient-benefit claims require comparative evidence with patient-relevant endpoints.

timeline
    title AI in Medicine: Breakthroughs and Setbacks (1950-2025)
    1950s-1960s : Birth of Medical AI
                : Turing Test (1950)
                : Dartmouth Conference (1956)
                : DENDRAL (1965): Molecular structure
    1970s-1980s : Expert Systems Era
                : MYCIN (1972): matched experts in evaluation, never deployed
                : INTERNIST-I (1974): Internal medicine
                : First AI Winter (late 1980s): Funding collapses
    1990s-2000s : Machine Learning Era
                : FDA clears first CAD systems (1998)
                : Google Flu Trends (2008): Initial success
                : EHR adoption accelerates
    2010s : Deep Learning Revolution
          : AlexNet (2012): ImageNet breakthrough
          : Watson for Oncology (2013): $62M MD Anderson failure
          : Google Flu Trends discontinued (2015)
          : IDx-DR (2018): First autonomous AI clearance
    2020s : Foundation Model Era
          : GPT-3 (2020), ChatGPT (2022)
          : Epic Sepsis Model criticism (2021): 33% sensitivity
          : Med-PaLM 2 passes USMLE (2023)
          : GPT-4 medical reasoning (2023)
          : FDA publishes AI-enabled medical devices list
          : WHO LMM guidance published (2025)
Figure 4.1: Timeline of AI development in medicine from the 1950s to 2020s, showing major breakthroughs and setbacks. Each era brought different approaches: expert systems, machine learning, deep learning, and foundation models. The cyclical pattern of hype and disillusionment has repeated multiple times, yet each wave built on lessons from previous attempts. Notable failures (Watson, Epic Sepsis Model criticism) are as instructive as successes.

The Birth of Medical AI (1950s–1960s)

The Turing Test and Medical Diagnosis

In 1950, Alan Turing asked a question that still haunts medical AI discussions: “Can machines think?”

What made Turing brilliant was skipping the philosophy entirely. He did not care about defining consciousness or machine sentience. He wanted a practical test. If you cannot tell whether you’re talking to a machine or a human, does it really matter which it is? Judge by outputs, not internal mechanisms.

We still evaluate medical AI this way. Can the algorithm’s diagnostic accuracy match or exceed expert physicians? Turns out that’s the easy part. What MYCIN taught us (painfully) is that matching expert performance does not mean anyone will actually use your system.

Historical Context

When Turing wrote his paper, computers were room-sized calculators used for mathematical computations and code-breaking (Turing himself had led cryptanalysis at Bletchley Park during World War II). The idea that machines might diagnose diseases or recommend treatments seemed like science fiction. Yet Turing explicitly discussed medical diagnosis as a potential application of machine intelligence.

The Dartmouth Conference (1956)

The field of AI was formally born at Dartmouth College (McCarthy et al., 1955) in summer 1956, when John McCarthy, Marvin Minsky, Claude Shannon, and other luminaries gathered for a two-month workshop with an audacious premise:

“Every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it.”

Their timeline predictions proved wildly optimistic. Some participants predicted machines would achieve human-level intelligence within a generation. Instead, we got symbolic AI systems that could play checkers but could not recognize a cat, a task any toddler performs effortlessly.

This early overconfidence established a pattern physicians still see today: brilliant computer scientists underestimating how much of medical expertise is tacit, contextual, and embodied. Diagnosing a patient is not just applying rules. It requires intuition built from thousands of cases, cultural competency, noticing what’s not documented in the chart, and integrating ambiguous, incomplete, sometimes contradictory information under uncertainty and time pressure.

These tacit competencies remain difficult to encode algorithmically.

Early Medical AI Attempts

The first attempts to apply AI to medicine emerged in the 1960s:

  • DENDRAL (1965) (Lindsay et al., 1993): Stanford built a system that identified molecular structures from mass spectrometry data. It worked, but only in a highly constrained domain with clear rules.

  • Pattern recognition for cancer diagnosis: Early computer vision systems tried to identify cancerous cells from microscope images. Results were mixed, and 1960s computational power was not remotely sufficient for complex image analysis.

These early efforts revealed a fundamental challenge that still persists: medicine deals with messy, incomplete data about extraordinarily complex biological systems. Unlike chess (deterministic rules, perfect information) or mathematical theorem proving (formal logic), medical diagnosis involves uncertainty, missing data, biological variability, comorbidities, and patient preferences.

The First Chatbots: ELIZA and the Computer Therapist Debate

While DENDRAL tackled molecular structure, another Stanford researcher pursued a different question: could computers conduct psychotherapy?

ELIZA (1964-1966): MIT’s Joseph Weizenbaum created ELIZA, a program that simulated conversation by pattern-matching user input and reflecting it back. The most famous script, DOCTOR, mimicked a Rogerian psychotherapist:

User: I am unhappy.
ELIZA: Do you think coming here will help you not to be unhappy?
User: I need some help, that much seems certain.
ELIZA: What would it mean to you if you got some help?

Weizenbaum intended ELIZA as a demonstration of superficiality, showing how easy it was to simulate understanding without possessing it. Instead, he was disturbed to find users confiding intimate details to the program and attributing empathy to a simple pattern-matcher. This phenomenon became known as the “ELIZA effect,” and it directly anticipates current concerns about users forming emotional attachments to AI chatbots.

Colby’s Computer Psychotherapy Proposal (1966): The same year ELIZA appeared, Stanford psychiatrist Kenneth Colby published a remarkable paper proposing that computers could deliver scripted psychotherapy at scale (Colby et al., 1966). Colby recognized that timesharing computers could support multiple simultaneous therapy sessions, potentially addressing the shortage of mental health professionals, an argument that sounds strikingly familiar today.

This created an enduring philosophical divide:

  • Colby (optimist): Computers could augment therapists, expand access, and serve as teaching tools for psychiatric trainees
  • Weizenbaum (skeptic): Delegating emotional care to machines was “immoral,” regardless of technical feasibility

Weizenbaum later wrote Computer Power and Human Reason (1976), explicitly rejecting Colby’s proposal: “I reject Colby’s proposal that computers be installed as psychotherapists, not on the grounds that such a project might be technically infeasible, but on the grounds that it is immoral.”

This debate, now 60 years old, remains unresolved. Today’s chatbot therapy applications (Woebot, Wysa) embody Colby’s vision while facing Weizenbaum’s critique. The Psychiatry and Behavioral Health chapter examines contemporary evidence for chatbot therapy.

The Expert Systems Era (1970s–1980s): MYCIN’s Promise and Failure

The 1970s brought a new approach: if we cannot make machines think like humans, maybe we can capture human expertise in formal IF-THEN rules.

MYCIN: Technical Success, Clinical Failure

In 1972, Edward Shortliffe began developing MYCIN (Shortliffe et al., 1975) at Stanford, an expert system that inferred likely bacterial causes of serious infection and recommended antimicrobial therapy. This was a seemingly well-bounded test case for AI in medicine:

Why MYCIN was promising:

  • Well-defined clinical problem: Identify causative bacteria and select appropriate antibiotics
  • Clear expertise: Infectious disease specialists followed identifiable reasoning patterns
  • Life-or-death stakes: Sepsis kills quickly; correct antibiotic choice dramatically affects mortality
  • Knowledge-intensive task: Success requires knowing hundreds of drug-bug interactions, resistance patterns, patient-specific factors

How MYCIN worked:

MYCIN used backward chaining through approximately 600 IF-THEN rules:

IF:
  1) Patient is immunocompromised host
  AND
  2) Site of infection is gastrointestinal tract
  AND
  3) Gram stain is gram-negative-rod
THEN:
  Evidence (0.7) that organism is E. coli

The system conducted structured consultations by asking questions, applying rules, and (crucially) explaining its reasoning. This explainability was unprecedented. MYCIN could answer “Why do you believe this?” and “How did you reach that conclusion?”, something today’s deep learning systems struggle to do convincingly.

The results were stunning:

The best-known blinded evaluation compared MYCIN’s antimicrobial choices with prescriptions from nine clinicians for ten meningitis cases (Yu et al., 1979). Eight independent evaluators with expertise in meningitis management reported:

  • 65% of MYCIN’s regimens were deemed acceptable
  • The five faculty specialists’ individual acceptability ratings ranged from 42.5% to 62.5%
  • MYCIN did not fail to cover a treatable pathogen in the ten-case evaluation

This was evidence about regimen acceptability in a small controlled case set, not diagnostic accuracy, patient outcomes, or safe autonomous deployment.

Papers celebrated MYCIN as a breakthrough. It was featured in medical journals, computer science conferences, and popular media. Stanford Medicine showcased it as the future of clinical decision support.

The devastating reality:

Despite this technical and clinical success, MYCIN was never deployed in routine clinical care. Not once. Not in a clinical trial. Not even in a supervised pilot study.

Why did MYCIN remain a research system? Historical accounts repeatedly identify the following barriers. They are useful translation hypotheses, but available sources do not quantify how much each barrier contributed:

Why MYCIN Failed: Lessons for Today’s Physician AI

Medical Liability: Who would be responsible if MYCIN recommended the wrong antibiotic and a patient was harmed? The physician who followed its advice? Stanford University? The programmers? In the 1970s, these questions lacked a mature software-specific clinical and regulatory framework. Current liability analysis remains fact- and jurisdiction-specific rather than resolved by a universal rule (see Liability and Legal Considerations).

Regulatory uncertainty: Software-specific medical-device pathways and evidence expectations were far less developed than they are today. Whether and how a consultation program such as MYCIN would be regulated was unsettled.

Technical Integration Barriers: Getting MYCIN’s recommendations required a separate computer terminal, manual data entry, and stepping outside normal clinical workflow. In the 1970s, hospitals were just beginning to adopt electronic systems. MYCIN could not integrate with existing infrastructure.

Physician trust and accountability: MYCIN could expose the rules it used, but explanation did not resolve whether a clinician could validate the underlying knowledge base or safely act on its advice during care.

Knowledge Maintenance Burden: Medical knowledge evolves. Keeping 600 hand-coded rules updated proved impractical. When new antibiotics became available or resistance patterns changed, updating MYCIN required programmer intervention.

Organizational fit and incentives: Dedicated computing resources, specialist roles, local governance, and the absence of a clear payment model all complicated adoption.

The lesson for today’s physicians:

MYCIN proves that technical excellence does not guarantee clinical adoption. The hardest problems deploying medical AI are rarely algorithmic. They’re legal, regulatory, social, organizational, and workflow-related.

Two decades after MYCIN’s failure, the Institute of Medicine formalized this insight in To Err is Human (Kohn et al., 2000): medical errors are primarily system failures, not individual failures. MYCIN failed not because of bad algorithms or bad physicians, but because the healthcare system lacked the infrastructure, regulatory frameworks, and organizational culture to deploy it safely. The IOM’s “good people working in bad systems” framing explains MYCIN’s fate and predicts similar challenges for today’s AI systems.

This lesson remains profoundly relevant. Today’s deep learning models vastly exceed MYCIN’s capabilities. Yet they face identical deployment challenges: liability uncertainty, integration complexity, physician trust barriers, and unclear value proposition relative to existing clinical workflows.

Other Expert Systems: Similar Patterns

The 1980s saw dozens of medical expert systems, most following MYCIN’s pattern of impressive demonstrations but minimal clinical impact:

  • INTERNIST-I/CADUCEUS (1974-1985) (Miller et al., 1982): Diagnosed diseases across internal medicine (~1000 diseases, ~3,500 manifestations). Could rival internists in complex cases. Never clinically deployed.

  • DXplain (1984-present) (Barnett et al., 1987): Differential diagnosis support system. Still used today for education and clinical decision support, but as a reference tool, not autonomous diagnostic system. Rare success story due to modest scope and human-in-the-loop design.

  • ONCOCIN (1981): Guided cancer chemotherapy protocols. Impressive demonstrations, minimal clinical adoption.

PARRY: The Psychiatric Counterpart to MYCIN

While MYCIN diagnosed infections, Kenneth Colby pursued a parallel path at Stanford. In 1972, the same year Shortliffe began MYCIN, Colby created PARRY, a chatbot simulating a patient with paranoid schizophrenia. Unlike ELIZA’s simple reflection, PARRY implemented a model of paranoid cognition with beliefs, fears, and defensive conversational strategies.

PARRY underwent a restricted, Turing-like validation experiment. Psychiatrists reviewed teleprinter conversations involving the simulation and patients with paranoid illness, and their transcript classifications were near chance (Colby et al., 1972). The experiment tested behavioral indistinguishability within a narrow transcript task, not diagnostic accuracy, therapeutic benefit, empathy, or safety. In 1972, network pioneer Vint Cerf arranged for PARRY and ELIZA to converse over ARPANET, an early chatbot-to-chatbot exchange between an artificial therapist and an artificial patient.

PARRY demonstrated that expert systems could model personality and affect, not just diagnostic rules. But like MYCIN, it remained a research curiosity. Colby’s vision of AI as a psychiatric teaching tool and therapeutic adjunct would wait 50 years for computational power, ubiquitous connectivity, and a mental health access crisis to make it viable.

Why Expert Systems Failed

By the late 1980s, the expert systems approach hit fundamental limits:

  1. Brittleness: Systems worked perfectly within their narrow domain but failed dramatically on edge cases. A single unexpected finding could derail the entire diagnostic process.

  2. Knowledge acquisition bottleneck: Extracting rules from expert physicians was extraordinarily time-consuming and incomplete. Experts often could not articulate their reasoning explicitly.

  3. Combinatorial explosion: Real-world medicine requires thousands of interacting rules. Managing complexity became unworkable.

  4. Maintenance burden: Medical knowledge evolves continuously. Hand-coded rules became outdated, requiring constant programmer intervention.

  5. Lack of learning: Expert systems could not improve from experience. Every new case provided no feedback to enhance future performance.

  6. Deployment realities: As MYCIN demonstrated, technical performance did not address liability, regulation, integration, or trust.

The “AI Winter” of the late 1980s and 1990s arrived. Funding dried up. Researchers left the field. Companies removed “AI” from marketing materials. Medical AI seemed like a failed experiment.

The lesson: Rule-based approaches could not capture the complexity, uncertainty, and nuance of real clinical practice. Medicine needed a fundamentally different approach.

The Machine Learning Revolution (1990s–2000s)

From Rules to Data

The 1990s brought a decisive shift: instead of encoding expert rules manually, why not let algorithms learn patterns directly from data?

Machine learning, particularly supervised learning, offered a solution. Show an algorithm thousands of examples (X-rays labeled normal vs. abnormal, patient data labeled survived vs. died), and it learns to recognize patterns without explicit rule programming.

Key developments enabling this shift:

  • Increased computational power: Moore’s Law made previously impractical computations feasible
  • Digitization of medical data: Picture Archiving and Communication Systems (PACS), early electronic health records
  • Algorithmic advances: Support Vector Machines (SVMs), decision trees, ensemble methods

Computer-Aided Detection (CAD) in Radiology

The first commercially successful medical AI applications emerged in radiology:

Computer-Aided Detection (CAD) for mammography:

  • FDA began approving CAD systems in the late 1990s
  • Designed to augment radiologist interpretation by flagging suspicious regions
  • Became widely adopted in breast cancer screening

Initial promise: Retrospective studies suggested CAD could reduce false negatives (missed cancers).

Reality check: Prospective studies showed mixed results (Lehman et al., 2015). CAD increased recall rates (more women called back for additional imaging) without consistently improving cancer detection rates. Some studies suggested CAD reduced radiologist specificity without improving sensitivity.

Current status: Second-generation AI systems using deep learning show more promise. The lesson: first-generation medical AI often underperforms expectations in real-world deployment.

FDA Begins Regulating Medical AI

The 1990s-2000s saw FDA establish regulatory frameworks for software-based medical devices:

  • 1990s: First CAD systems cleared through 510(k) pathway (substantial equivalence to existing devices)
  • 2000s: FDA establishes guidance for computer-assisted detection devices
  • Challenge: Rapid AI evolution outpaced regulatory frameworks designed for static medical devices

This regulatory evolution continues today, with FDA developing new frameworks for continuously learning AI systems.

Why This Era Mattered

The machine learning revolution established principles still guiding medical AI:

  • Data-driven approaches could achieve good performance without explicit rule encoding
  • Narrow, well-defined tasks (e.g., detecting lung nodules on CT) worked better than general diagnosis
  • Human-AI configuration became an empirical question: assistance can improve, worsen, or leave performance unchanged depending on the task, interface, threshold, user, and workflow
  • Prospective validation essential: Retrospective performance did not guarantee real-world utility
  • Workflow integration remained challenging despite better algorithms

The limitation: Traditional machine learning required careful feature engineering (hand-crafted measurements and patterns). Algorithms could not learn complex representations directly from raw data.

The deep learning revolution would change that.

The Deep Learning Revolution (2010s): Imaging AI Comes of Age

AlexNet and the ImageNet Moment (2012)

In 2012, a deep convolutional neural network called AlexNet (Krizhevsky et al., 2012) won the ImageNet visual recognition challenge by a massive margin, halving the error rate of competing approaches.

This was not just incremental progress. It was a fundamental break from prior approaches. Deep learning could learn complex features directly from raw pixels without manual feature engineering. Suddenly, computer vision problems that had resisted decades of effort became tractable.

Implications for medical imaging:

Medical images (X-rays, CTs, MRIs, pathology slides, dermatology photos, retinal fundus images) are fundamentally visual pattern recognition problems. If deep learning could master general image classification, perhaps it could learn to detect pneumonia, identify cancers, or diagnose diabetic retinopathy.

Medical Imaging AI Explosion

The mid-2010s saw an explosion of medical imaging AI research:

Papers proliferated. Venture capital flooded into medical AI startups. Headlines proclaimed “AI Will Replace Radiologists.”

The First FDA-Authorized Autonomous AI (2018): IDx-DR

On April 11, 2018, FDA granted De Novo authorization to IDx-DR (DEN180001), the first system authorized to provide a diabetic-retinopathy screening decision without requiring an eye-care specialist to interpret the images (Abràmoff et al., 2018). Its authorized scope was specific: detecting more-than-mild diabetic retinopathy in adults with diabetes who had not previously been diagnosed with diabetic retinopathy, using images acquired with the specified camera.

Why IDx-DR succeeded where MYCIN failed:

  • Narrow, well-defined task: Detect referable diabetic retinopathy (yes/no decision)
  • Clear clinical need: Primary care physicians need retinal screening but lack ophthalmology expertise
  • Prospective validation: 900-patient clinical trial in primary care settings (not just retrospective analysis)
  • Autonomous operation: Primary care staff operate the system without specialist interpretation
  • Regulatory clarity: FDA had developed frameworks for AI-based medical devices
  • Reimbursement: CPT codes established for AI-assisted diabetic retinopathy screening

Clinical bottom line: IDx-DR illustrates several conditions that can support medical AI deployment: 1. Well-defined, high-value clinical problem 2. Prospective validation in real-world settings 3. Clear regulatory pathway 4. Workflow integration 5. Reimbursement model 6. Physician and patient acceptance

The “AI Will Replace Radiologists” Debate

Geoffrey Hinton (deep learning pioneer) stated at a 2016 conference: “It’s quite obvious that we should stop training radiologists.”

This sparked fierce debate. Would AI replace radiologists? Should medical students avoid imaging specialties?

What actually happened:

  • AI did not eliminate the radiology workforce
  • AI systems entered specific radiology workflow tasks
  • Radiologists increasingly use AI as assistive tools
  • AI handles some straightforward screening tasks
  • Complex cases still require radiologist expertise
  • New roles emerged: radiologists curating datasets, validating AI, interpreting edge cases

The lesson for physicians:

The historical record does not support a single replacement narrative. Some tasks are automated, others are redistributed, and many tools create new work for verification, exception handling, quality assurance, and patient communication.

Fear replacement less. Focus on effective integration more.

Deep Learning’s Limitations in Medicine

Despite these advances, deep learning revealed serious limitations:

Interpretability problem: Neural networks do not ordinarily provide a faithful, clinically sufficient account of the internal basis for a prediction. A radiologist can articulate reasoning (“spiculated mass in the upper lobe with associated lymphadenopathy suggests malignancy”). A generated explanation or heat map may help review, but it is not automatically a causal account of the model’s computation.

Data hunger: Deep learning requires massive labeled datasets (thousands to millions of examples). Curating these datasets requires enormous physician time.

Brittleness to distribution shift: Models trained on data from Hospital A often perform poorly at Hospital B due to different imaging equipment, patient populations, disease prevalence, or documentation practices.

Adversarial vulnerability: Tiny, imperceptible changes to input images can fool deep learning models completely, a major patient safety concern (Finlayson et al., 2019).

Fairness and bias: AI systems inherit biases from training data. If training data under-represents certain populations, model performance may be worse for those groups (Obermeyer et al., 2019).

These limitations drive ongoing research and FDA regulatory evolution.

IBM Watson for Oncology: The Highest-Profile Failure

No discussion of medical AI history would be complete without IBM Watson for Oncology, arguably the most expensive, highest-profile medical AI failure to date.

The Promise

In 2011, IBM’s Watson defeated human champions on Jeopardy!, demonstrating impressive natural language processing capabilities. IBM launched Watson for Oncology in 2013, partnering with Memorial Sloan Kettering to train the system. IBM promised Watson would reshape medicine by:

  • Analyzing millions of medical journal articles
  • Synthesizing complex patient data from EHRs
  • Recommending personalized, evidence-based cancer treatments
  • Augmenting oncologist decision-making

Major cancer centers worldwide partnered with IBM: MD Anderson Cancer Center, Memorial Sloan Kettering, hospitals across India, China, and other countries. IBM invested billions. Expectations soared.

The Reality

Watson for Oncology became a prominent example of the gap between commercial ambition and clinical evidence. The strongest public record combines investigative reporting with retrospective concordance studies, rather than randomized patient-outcome evidence (Ross & Swetlitz, 2018):

Unsafe recommendations: Internal documents revealed Watson recommended treatments that would have harmed patients. In one case, Watson suggested administering chemotherapy to a patient with severe bleeding, a contraindicated, potentially fatal recommendation (Ross & Swetlitz, 2018).

Variable concordance: A meta-analysis of nine nonrandomized studies and 2,463 patients assessed agreement between Watson for Oncology and multidisciplinary teams. Overall concordance was 81.52% when both “recommended” and “for consideration” categories counted as concordant, with substantial variation by cancer, stage, and patient characteristics (Jie et al., 2021). Concordance measures agreement with a comparator, not treatment effectiveness or patient benefit.

Geographic inappropriateness: Watson recommended treatments unavailable in certain regions, ignoring local resource constraints and formulary restrictions. While some studies showed high concordance (93% in one Indian breast cancer study), Watson’s recommendations often included drugs not covered by local insurance or unavailable in certain formularies (Somashekhar et al., 2018).

Training data issues: Watson was trained primarily on synthetic cases created by Memorial Sloan Kettering physicians, not real-world patient data. It learned institutional preferences, not universal evidence (Strickland, 2019).

MD Anderson debacle: MD Anderson Cancer Center spent $62 million on Watson implementation before canceling the project in 2016, concluding it was not ready for clinical use (Strickland, 2019).

Widespread abandonment: By 2019, multiple health systems had stopped using Watson. IBM sold Watson Health assets in 2021.

What Went Wrong: Lessons for Physicians

Critical Lessons from Watson’s Failure
  1. Marketing ≠ Clinical Validation: Jeopardy! success does not translate to clinical competence. Demand prospective clinical trials, not just impressive demonstrations in controlled settings.

  2. Reviewability matters: A recommendation is difficult to govern when clinicians cannot inspect its evidence basis, applicability, and uncertainty efficiently.

  3. Training data matters immensely: Synthetic cases created by one institution do not represent the diversity of real clinical practice.

  4. Clinical participation must extend beyond labeling: Clinicians, patients, informaticians, safety teams, and operational leaders must shape intended use, evaluation, escalation, and monitoring.

  5. Geographic and institutional context matters: Treatment recommendations must account for local resources, formularies, patient populations, and practice patterns.

  6. Commercial momentum is not evidence: Procurement and deployment can proceed faster than independent clinical validation.

  7. Transparency and reporting: Watson’s failures remained largely hidden until investigative journalists uncovered them. Medical AI needs transparent reporting of failures.

The clinical bottom line:

Watson’s failure demonstrates why physicians must evaluate AI systems with the same rigor applied to new pharmaceuticals: demand prospective clinical trials, transparent reporting, independent validation, and post-market surveillance.

Do not accept vendor claims without evidence.

The Foundation Model Era (2020–Present)

Large Language Models Arrive

November 2022 brought ChatGPT, introducing millions of people (including physicians) to large language models (LLMs) (Brown et al., 2020). These “foundation models” could:

  • Answer medical questions with apparently sophisticated reasoning
  • Draft clinical notes
  • Explain complex concepts
  • Translate between languages
  • Write code
  • Summarize literature

Unlike narrow AI systems designed for specific tasks, LLMs demonstrated general capabilities across diverse domains.

Medical-Specific Foundation Models

Recognizing both promise and risks of general LLMs in medicine, researchers developed medical-specific models.

How Post-Training Shapes Medical LLM Behavior

Medical LLM performance cannot be attributed to RLHF alone. Modern post-training may combine supervised instruction tuning, domain adaptation, human or AI preference data, safety examples, and verifier-based reinforcement learning. Each method optimizes against its own demonstrations, labels, rubric, or reward. Clinician feedback can shape response behavior, but evaluator preference is not proof of clinical accuracy, safety, or patient benefit. Open post-training recipes illustrate this diversity, although their reported gains do not establish transfer to clinical work (Lambert et al., 2024, preprint; Lee et al., 2024).

Med-PaLM is an important correction to the simplified RLHF story. The published model used instruction prompt tuning, a parameter-efficient method based on medical-domain exemplars, rather than clinician-rated RLHF. Its authors attributed performance differences to model scale, prompting, instruction tuning, and medical instruction prompt tuning, while emphasizing that the system remained inferior to clinicians on important dimensions (Singhal et al., 2023). Clinical performance also depends on the base model, data, retrieval and tools, evaluation design, and deployment context.

Med-PaLM (Google, 2023):

  • Adapted from Flan-PaLM using medical instruction prompt tuning
  • Achieved passing scores on USMLE-style questions (Singhal et al., 2023)
  • Improved clinical accuracy compared to general LLMs

GPT-4 in medicine (OpenAI, 2023):

  • Demonstrated strong performance on medical licensing exams
  • Used by physicians for literature review and summarization, clinical reasoning support, patient education

Challenges remain:

  • Hallucinations: LLMs confidently generate plausible but incorrect information
  • Bias: Inherit biases from training data
  • Liability: Duties and allocation of responsibility remain fact-specific and jurisdiction-specific
  • Privacy: Patient data security concerns
  • Validation: How to validate general-purpose models across diverse clinical scenarios?

Current State: Promise and Uncertainty

As of August 2026, medical AI presents a mixed picture:

Undeniable progress: - FDA maintains a growing, non-exhaustive list of AI-enabled devices authorized through pathway-specific review (see the FDA AI-enabled medical-device list) - Radiology, pathology, dermatology, and ophthalmology include authorized and deployed tools, but adoption and evidence vary by product, institution, and use case - Clinical decision support systems are deployed in major EHR systems - Foundation models demonstrate unprecedented versatility

Persistent challenges: - Most AI systems remain narrow, fragile, and context-dependent - Integration into clinical workflow remains difficult - Liability rules vary across jurisdictions and continue to develop alongside clinical use - Bias and fairness concerns are well-documented but incompletely solved - Physician trust varies widely by specialty, generation, and prior experience

The fundamental question:

Are today’s foundation models genuinely different from previous AI waves, or are we in another hype cycle that will end in disillusionment?

Evidence suggests: Both. Foundation models represent real algorithmic breakthroughs. But transforming clinical practice requires solving the same deployment challenges that defeated MYCIN and Watson: integration, liability, trust, validation, and proving clinical benefit beyond technical performance metrics.

Key Lessons from History

Seven decades of medical AI teach clear lessons:

Historical Lessons for Physicians
  1. Technical performance ≠ clinical adoption (MYCIN)
  2. Marketing hype ≠ clinical validity (Watson)
  3. Retrospective validation ≠ prospective utility (Google Flu Trends)
  4. Narrow, well-defined tasks work best (IDx-DR)
  5. Human-AI performance must be tested, not assumed (CAD in radiology)
  6. Physician involvement is essential (Watson’s failure)
  7. Regulatory frameworks evolve slowly (ongoing FDA challenges)
  8. Liability concerns shape adoption (MYCIN’s legal uncertainty)
  9. Workflow integration is harder than algorithmic development (expert systems)
  10. Patient safety must remain paramount (Watson’s unsafe recommendations)

The path forward:

Learn from failures. Demand evidence. Integrate thoughtfully. Maintain physician oversight. Prioritize patients.

History does not repeat, but it rhymes. The physicians who understand AI’s history will navigate its future most effectively.

Skill Degradation in the Age of AI: Historical Lessons

One important risk of medical AI is cognitive deskilling, the possible erosion of clinical expertise when clinicians repeatedly defer observation, interpretation, or judgment to an automated system.

The concern is plausible and has emerging empirical support, but its magnitude is not established across medicine. The evidence must distinguish three related problems: automation bias (overweighting a system’s recommendation), deskilling (loss of a previously acquired capability), and never-skilling (failing to acquire that capability during training).

Parallel Examples: GPS Navigation and Mental Arithmetic

Consider GPS navigation. Before smartphones, drivers developed spatial reasoning and mental maps of their environments. They could navigate without turn-by-turn directions, improvise when routes were blocked, and maintain geographic awareness.

Research has associated greater lifetime GPS use with poorer spatial memory during self-guided navigation, including longitudinal associations within the study cohort (Dahmani & Bohbot, 2020). This supports concern about cognitive offloading, but an association in navigation does not prove that clinical AI will cause the same effect.

Calculators provide a familiar analogy. A user can obtain a correct computation without maintaining the mental arithmetic needed to estimate whether the output is plausible. The analogy clarifies the risk, but it is not direct evidence about physician performance.

Neither GPS nor calculators are inherently harmful. They illustrate the design question: which capabilities can be offloaded safely, and which must remain available for supervision, exception handling, and system failure?

Documented Skill Atrophy in Medical AI

Medical AI is following the same trajectory, with concerning evidence emerging from radiology and gastroenterology.

Computer-Aided Detection (CAD) in Radiology:

Early CAD systems for mammography promised to reduce missed cancers by flagging suspicious regions for radiologist review. Their clinical value proved inconsistent. A large community-practice analysis found no improvement in diagnostic accuracy with CAD (Lehman et al., 2015). Researchers also examined whether prompts altered how radiologists searched images and weighted evidence.

Experimental and observational work described interaction effects between CAD prompts and radiologist decisions, including errors related to inappropriate reliance (Alberdi et al., 2004). That study supports an automation-bias concern. It did not longitudinally demonstrate that radiologists’ pattern-recognition ability atrophied.

One proposed mechanism is attentional redistribution: when an algorithm highlights suspicious regions, a reader may devote less attention to unflagged areas. Whether repeated exposure produces durable skill loss depends on the interface, task, training, feedback, and opportunities for independent practice.

The Gastroenterology Study: Physicians Became Significantly Worse

Recent gastroenterology evidence shows both a risk signal and important uncertainty. In a retrospective multicenter observational study at four Polish centers, the adenoma detection rate during non-AI colonoscopy declined from 28.4% (226/795) before routine AI exposure to 22.4% (145/648) afterward, an absolute difference of 6.0 percentage points (Budzyń et al., 2025). The before-and-after design supports an association, not proof that AI exposure caused the decline.

The observed decline is a credible deskilling signal, but the design does not establish causation or inevitable harm.

Two later studies did not reproduce inevitable deterioration. A prospective single-center observational study found that non-AI detection performance was maintained after CADe implementation (Okumura et al., 2026). A prospective multicenter pragmatic trial involving 13 endoscopists and 5,013 colonoscopies found a temporary CADe-associated improvement among inexperienced endoscopists, but no significant upskilling or deskilling after CADe removal (Pedersen et al., 2026). The defensible conclusion is conditional: deskilling can occur, but it is not an established universal consequence of AI assistance.

Training the Next Generation: Medical Students Who Never Develop Foundational Skills

The skill question becomes especially important in medical education. If students learn alongside tools that generate differential diagnoses, treatment suggestions, and documentation, educators must decide when independent performance is required and how competence will be assessed without assistance.

Dhruv Khullar, writing in The New Yorker, described medical students’ discomfort about presenting AI-generated reasoning as their own (Khullar, 2025). This is a useful account of learner experience, not a controlled estimate of educational harm.

This raises profound questions about what medical training should accomplish in an age of clinical AI:

  • If students can query AI for differential diagnoses, should they still memorize disease presentations?
  • If algorithms interpret ECGs more accurately than cardiologists, what ECG interpretation skills should medical students master?
  • If AI drafts clinical notes, how will students learn to integrate complex information into coherent clinical narratives?
  • If radiology AI detects pathology automatically, what level of image interpretation competency should physicians maintain?

The risk is not that students use AI tools. It’s that they never develop robust foundational skills because AI assistance is always available. Then, when AI fails, is unavailable, or encounters edge cases beyond its training, there’s no physician expertise to fall back on.

Historical Lesson: Every Technology Changed What Physicians Knew

This pattern is not new. Each wave of medical technology transformed physician capabilities:

Calculators and clinical scoring systems: Physicians no longer perform complex risk calculations mentally. They input variables into validated scoring systems. Mental arithmetic skills declined, but risk stratification improved.

Electronic health records (EHRs): Physicians rely on structured templates, autocomplete, and copy-forward functionality. Documentation became standardized but less personalized. Narrative skills atrophied. Many physicians struggle to write coherent clinical summaries without EHR templates.

Automated laboratory analyzers: Physicians stopped performing manual blood counts and chemical assays. Lab interpretation skills improved (access to more data), but understanding of underlying physiology declined. Few physicians today could estimate hemoglobin from examining a peripheral blood smear.

Advanced imaging (CT, MRI): Physical examination skills declined as imaging became routine for diagnosis. Studies show physicians miss physical findings that would have been obvious to previous generations, because they order imaging instead of examining patients thoroughly.

Each technology brought benefits and redistributed work. The effects on competence varied by task, specialty, training, and implementation.

AI may produce a broader shift because it can affect pattern recognition, diagnostic reasoning, documentation, and decision-making under uncertainty at the same time.

The Unsolved Tension: Competence vs. Dependency

There’s no simple resolution to this tension.

Refusing to use AI to preserve traditional skills would harm patients if AI genuinely improves outcomes. But uncritical adoption creates dangerous dependencies and erodes the expertise needed when technology fails.

Medical educators, licensing bodies, and specialty societies are grappling with fundamental questions:

  • What clinical skills are truly foundational, requiring mastery regardless of AI availability?
  • What tasks can we safely delegate to algorithms, accepting that human competency will decline?
  • How do we maintain physician capability to recognize and override algorithmic errors?
  • What backup competencies must physicians preserve for when AI systems fail?
  • How do we train medical students to use AI effectively without becoming entirely dependent on it?

No consensus exists. Medical schools are experimenting with different approaches: some integrate AI early in training, others restrict its use until students demonstrate independent competency. Some specialties (radiology, pathology) are redesigning curricula around AI-augmented practice. Others (surgery, primary care) emphasize skills that remain distinctly human.

The historical lesson is clear: Technology always changes physician capabilities. The question is not whether AI will transform medical expertise, but how we manage that transformation to preserve the irreplaceable elements of clinical judgment while embracing genuine improvements.

Physicians who understand this historical pattern, and who treat skill preservation as a measurable implementation objective, will be better positioned to maintain critical competencies while integrating AI thoughtfully.

The long-term effects of AI on physician expertise remain uncertain. A safer response is to define the capabilities clinicians must retain, test performance with and without assistance, monitor override behavior and error recovery, and design periodic independent practice where the consequences justify it.

Questions About the Medical AI Timeline

Why did IBM Watson Health fail?

Watson for Oncology was evaluated mainly through agreement with multidisciplinary treatment recommendations, not randomized patient-outcome trials. Concordance varied by cancer type and setting, reports documented unsafe recommendations, and major implementations were discontinued.

When was AI first used in medicine?

Medical AI research began before the 1970s. DENDRAL analyzed molecular structures in the 1960s, ELIZA simulated psychotherapy dialogue in 1966, and MYCIN began development in 1972. Earlier statistical prediction models were influential precursors but are not identical to later artificial-intelligence systems.

When was AI first used in healthcare?

Research systems entered health-related domains in the 1960s, while MYCIN became a landmark clinical expert system in the 1970s. In a 1979 blinded evaluation, experts rated 65% of MYCIN’s antimicrobial regimens acceptable for ten meningitis cases, but MYCIN was not deployed in routine patient care.

What is the difference between machine learning and deep learning in healthcare?

Machine learning is the broader category of algorithms that learn patterns from data. Deep learning is a machine-learning approach based on multilayer neural networks and is used with both structured and unstructured data. Its modern resurgence accelerated medical imaging research after the 2012 AlexNet result.

Will AI replace doctors?

No historical record can prove that AI will never replace any physician task. Current evidence supports task-specific automation and human-AI collaboration, while adoption depends on clinical benefit, workflow, regulation, reliability, and accountability.

What was the first chatbot used in medicine?

ELIZA’s DOCTOR script simulated a Rogerian psychotherapist in 1966. PARRY, described by Kenneth Colby and colleagues in 1972, modeled paranoid processes. A restricted transcript-identification experiment found that psychiatrists did not reliably distinguish the simulation from patient transcripts; it was not a test of diagnosis or treatment effectiveness.