timeline
title AI in Medicine: Breakthroughs and Setbacks (1950-2025)
1950s-1960s : Birth of Medical AI
: Turing Test (1950)
: Dartmouth Conference (1956)
: DENDRAL (1965): Molecular structure
1970s-1980s : Expert Systems Era
: MYCIN (1972): matched experts in evaluation, never deployed
: INTERNIST-I (1974): Internal medicine
: First AI Winter (late 1980s): Funding collapses
1990s-2000s : Machine Learning Era
: FDA clears first CAD systems (1998)
: Google Flu Trends (2008): Initial success
: EHR adoption accelerates
2010s : Deep Learning Revolution
: AlexNet (2012): ImageNet breakthrough
: Watson for Oncology (2013): $62M MD Anderson failure
: Google Flu Trends discontinued (2015)
: IDx-DR (2018): First autonomous AI clearance
2020s : Foundation Model Era
: GPT-3 (2020), ChatGPT (2022)
: Epic Sepsis Model criticism (2021): 33% sensitivity
: Med-PaLM 2 passes USMLE (2023)
: GPT-4 medical reasoning (2023)
: FDA publishes AI-enabled medical devices list
: WHO LMM guidance published (2025)
History of AI in Medicine: From MYCIN to Foundation Models
In 1979, eight independent experts reviewed antimicrobial regimens proposed by Stanford’s MYCIN and nine human prescribers for ten meningitis cases. They rated 65% of MYCIN’s regimens acceptable, compared with 42.5%–62.5% for the five faculty specialists, but MYCIN never entered routine patient care (Yu et al., 1979). That distinction between controlled task performance and clinical adoption is one of the central lessons of 70 years of AI in medicine.
After reading this chapter, you will be able to:
- Distinguish genuine clinical breakthroughs from recurring hype cycles (MYCIN, IBM Watson, diagnostic AI)
- Recognize why technical excellence does not guarantee clinical adoption
- Identify patterns separating successful medical AI deployments from research prototypes
- Understand FDA regulation evolution and current frameworks
- Evaluate whether today’s foundation models represent a fundamental departure from prior medical AI
- Apply historical lessons about clinical integration, medical liability, and physician trust
Introduction
Artificial intelligence in medicine is not new. The field has experienced multiple waves of excitement and bitter disillusionment over seven decades. Each cycle promised to reshape clinical practice. Each fell short. The scale of recent activity is unprecedented: AI-related medical publications increased 36-fold between 2000 and 2022, from 1,614 to over 58,000 annually (Shi et al., 2023).
So why should physicians believe that this time is different?
Understanding AI’s history in medicine is not just academic curiosity. It’s essential for navigating today’s hype, identifying genuinely transformative applications, and avoiding expensive failures that harm patients or waste resources. The patterns repeat: breathless promises, pilot studies that look impressive, deployment challenges nobody anticipated, and eventual disillusionment when technology does not match marketing.
But history also reveals what works. Successful medical AI applications often share common traits: they solve specific, well-defined clinical problems; they are evaluated in the intended population and workflow; and they define what the clinician must do when the system is wrong, unavailable, or out of scope. Prospective evaluation strengthens claims about use in practice, while patient-benefit claims require comparative evidence with patient-relevant endpoints.
The Birth of Medical AI (1950s–1960s)
The Turing Test and Medical Diagnosis
In 1950, Alan Turing asked a question that still haunts medical AI discussions: “Can machines think?”
What made Turing brilliant was skipping the philosophy entirely. He did not care about defining consciousness or machine sentience. He wanted a practical test. If you cannot tell whether you’re talking to a machine or a human, does it really matter which it is? Judge by outputs, not internal mechanisms.
We still evaluate medical AI this way. Can the algorithm’s diagnostic accuracy match or exceed expert physicians? Turns out that’s the easy part. What MYCIN taught us (painfully) is that matching expert performance does not mean anyone will actually use your system.
When Turing wrote his paper, computers were room-sized calculators used for mathematical computations and code-breaking (Turing himself had led cryptanalysis at Bletchley Park during World War II). The idea that machines might diagnose diseases or recommend treatments seemed like science fiction. Yet Turing explicitly discussed medical diagnosis as a potential application of machine intelligence.
The Dartmouth Conference (1956)
The field of AI was formally born at Dartmouth College (McCarthy et al., 1955) in summer 1956, when John McCarthy, Marvin Minsky, Claude Shannon, and other luminaries gathered for a two-month workshop with an audacious premise:
“Every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it.”
Their timeline predictions proved wildly optimistic. Some participants predicted machines would achieve human-level intelligence within a generation. Instead, we got symbolic AI systems that could play checkers but could not recognize a cat, a task any toddler performs effortlessly.
This early overconfidence established a pattern physicians still see today: brilliant computer scientists underestimating how much of medical expertise is tacit, contextual, and embodied. Diagnosing a patient is not just applying rules. It requires intuition built from thousands of cases, cultural competency, noticing what’s not documented in the chart, and integrating ambiguous, incomplete, sometimes contradictory information under uncertainty and time pressure.
These tacit competencies remain difficult to encode algorithmically.
Early Medical AI Attempts
The first attempts to apply AI to medicine emerged in the 1960s:
DENDRAL (1965) (Lindsay et al., 1993): Stanford built a system that identified molecular structures from mass spectrometry data. It worked, but only in a highly constrained domain with clear rules.
Pattern recognition for cancer diagnosis: Early computer vision systems tried to identify cancerous cells from microscope images. Results were mixed, and 1960s computational power was not remotely sufficient for complex image analysis.
These early efforts revealed a fundamental challenge that still persists: medicine deals with messy, incomplete data about extraordinarily complex biological systems. Unlike chess (deterministic rules, perfect information) or mathematical theorem proving (formal logic), medical diagnosis involves uncertainty, missing data, biological variability, comorbidities, and patient preferences.
The First Chatbots: ELIZA and the Computer Therapist Debate
While DENDRAL tackled molecular structure, another Stanford researcher pursued a different question: could computers conduct psychotherapy?
ELIZA (1964-1966): MIT’s Joseph Weizenbaum created ELIZA, a program that simulated conversation by pattern-matching user input and reflecting it back. The most famous script, DOCTOR, mimicked a Rogerian psychotherapist:
User: I am unhappy.
ELIZA: Do you think coming here will help you not to be unhappy?
User: I need some help, that much seems certain.
ELIZA: What would it mean to you if you got some help?
Weizenbaum intended ELIZA as a demonstration of superficiality, showing how easy it was to simulate understanding without possessing it. Instead, he was disturbed to find users confiding intimate details to the program and attributing empathy to a simple pattern-matcher. This phenomenon became known as the “ELIZA effect,” and it directly anticipates current concerns about users forming emotional attachments to AI chatbots.
Colby’s Computer Psychotherapy Proposal (1966): The same year ELIZA appeared, Stanford psychiatrist Kenneth Colby published a remarkable paper proposing that computers could deliver scripted psychotherapy at scale (Colby et al., 1966). Colby recognized that timesharing computers could support multiple simultaneous therapy sessions, potentially addressing the shortage of mental health professionals, an argument that sounds strikingly familiar today.
This created an enduring philosophical divide:
- Colby (optimist): Computers could augment therapists, expand access, and serve as teaching tools for psychiatric trainees
- Weizenbaum (skeptic): Delegating emotional care to machines was “immoral,” regardless of technical feasibility
Weizenbaum later wrote Computer Power and Human Reason (1976), explicitly rejecting Colby’s proposal: “I reject Colby’s proposal that computers be installed as psychotherapists, not on the grounds that such a project might be technically infeasible, but on the grounds that it is immoral.”
This debate, now 60 years old, remains unresolved. Today’s chatbot therapy applications (Woebot, Wysa) embody Colby’s vision while facing Weizenbaum’s critique. The Psychiatry and Behavioral Health chapter examines contemporary evidence for chatbot therapy.
The Expert Systems Era (1970s–1980s): MYCIN’s Promise and Failure
The 1970s brought a new approach: if we cannot make machines think like humans, maybe we can capture human expertise in formal IF-THEN rules.
MYCIN: Technical Success, Clinical Failure
In 1972, Edward Shortliffe began developing MYCIN (Shortliffe et al., 1975) at Stanford, an expert system that inferred likely bacterial causes of serious infection and recommended antimicrobial therapy. This was a seemingly well-bounded test case for AI in medicine:
Why MYCIN was promising:
- Well-defined clinical problem: Identify causative bacteria and select appropriate antibiotics
- Clear expertise: Infectious disease specialists followed identifiable reasoning patterns
- Life-or-death stakes: Sepsis kills quickly; correct antibiotic choice dramatically affects mortality
- Knowledge-intensive task: Success requires knowing hundreds of drug-bug interactions, resistance patterns, patient-specific factors
How MYCIN worked:
MYCIN used backward chaining through approximately 600 IF-THEN rules:
IF:
1) Patient is immunocompromised host
AND
2) Site of infection is gastrointestinal tract
AND
3) Gram stain is gram-negative-rod
THEN:
Evidence (0.7) that organism is E. coli
The system conducted structured consultations by asking questions, applying rules, and (crucially) explaining its reasoning. This explainability was unprecedented. MYCIN could answer “Why do you believe this?” and “How did you reach that conclusion?”, something today’s deep learning systems struggle to do convincingly.
The results were stunning:
The best-known blinded evaluation compared MYCIN’s antimicrobial choices with prescriptions from nine clinicians for ten meningitis cases (Yu et al., 1979). Eight independent evaluators with expertise in meningitis management reported:
- 65% of MYCIN’s regimens were deemed acceptable
- The five faculty specialists’ individual acceptability ratings ranged from 42.5% to 62.5%
- MYCIN did not fail to cover a treatable pathogen in the ten-case evaluation
This was evidence about regimen acceptability in a small controlled case set, not diagnostic accuracy, patient outcomes, or safe autonomous deployment.
Papers celebrated MYCIN as a breakthrough. It was featured in medical journals, computer science conferences, and popular media. Stanford Medicine showcased it as the future of clinical decision support.
The devastating reality:
Despite this technical and clinical success, MYCIN was never deployed in routine clinical care. Not once. Not in a clinical trial. Not even in a supervised pilot study.
Why did MYCIN remain a research system? Historical accounts repeatedly identify the following barriers. They are useful translation hypotheses, but available sources do not quantify how much each barrier contributed:
Medical Liability: Who would be responsible if MYCIN recommended the wrong antibiotic and a patient was harmed? The physician who followed its advice? Stanford University? The programmers? In the 1970s, these questions lacked a mature software-specific clinical and regulatory framework. Current liability analysis remains fact- and jurisdiction-specific rather than resolved by a universal rule (see Liability and Legal Considerations).
Regulatory uncertainty: Software-specific medical-device pathways and evidence expectations were far less developed than they are today. Whether and how a consultation program such as MYCIN would be regulated was unsettled.
Technical Integration Barriers: Getting MYCIN’s recommendations required a separate computer terminal, manual data entry, and stepping outside normal clinical workflow. In the 1970s, hospitals were just beginning to adopt electronic systems. MYCIN could not integrate with existing infrastructure.
Physician trust and accountability: MYCIN could expose the rules it used, but explanation did not resolve whether a clinician could validate the underlying knowledge base or safely act on its advice during care.
Knowledge Maintenance Burden: Medical knowledge evolves. Keeping 600 hand-coded rules updated proved impractical. When new antibiotics became available or resistance patterns changed, updating MYCIN required programmer intervention.
Organizational fit and incentives: Dedicated computing resources, specialist roles, local governance, and the absence of a clear payment model all complicated adoption.
The lesson for today’s physicians:
MYCIN proves that technical excellence does not guarantee clinical adoption. The hardest problems deploying medical AI are rarely algorithmic. They’re legal, regulatory, social, organizational, and workflow-related.
Two decades after MYCIN’s failure, the Institute of Medicine formalized this insight in To Err is Human (Kohn et al., 2000): medical errors are primarily system failures, not individual failures. MYCIN failed not because of bad algorithms or bad physicians, but because the healthcare system lacked the infrastructure, regulatory frameworks, and organizational culture to deploy it safely. The IOM’s “good people working in bad systems” framing explains MYCIN’s fate and predicts similar challenges for today’s AI systems.
This lesson remains profoundly relevant. Today’s deep learning models vastly exceed MYCIN’s capabilities. Yet they face identical deployment challenges: liability uncertainty, integration complexity, physician trust barriers, and unclear value proposition relative to existing clinical workflows.
Other Expert Systems: Similar Patterns
The 1980s saw dozens of medical expert systems, most following MYCIN’s pattern of impressive demonstrations but minimal clinical impact:
INTERNIST-I/CADUCEUS (1974-1985) (Miller et al., 1982): Diagnosed diseases across internal medicine (~1000 diseases, ~3,500 manifestations). Could rival internists in complex cases. Never clinically deployed.
DXplain (1984-present) (Barnett et al., 1987): Differential diagnosis support system. Still used today for education and clinical decision support, but as a reference tool, not autonomous diagnostic system. Rare success story due to modest scope and human-in-the-loop design.
ONCOCIN (1981): Guided cancer chemotherapy protocols. Impressive demonstrations, minimal clinical adoption.
PARRY: The Psychiatric Counterpart to MYCIN
While MYCIN diagnosed infections, Kenneth Colby pursued a parallel path at Stanford. In 1972, the same year Shortliffe began MYCIN, Colby created PARRY, a chatbot simulating a patient with paranoid schizophrenia. Unlike ELIZA’s simple reflection, PARRY implemented a model of paranoid cognition with beliefs, fears, and defensive conversational strategies.
PARRY underwent a restricted, Turing-like validation experiment. Psychiatrists reviewed teleprinter conversations involving the simulation and patients with paranoid illness, and their transcript classifications were near chance (Colby et al., 1972). The experiment tested behavioral indistinguishability within a narrow transcript task, not diagnostic accuracy, therapeutic benefit, empathy, or safety. In 1972, network pioneer Vint Cerf arranged for PARRY and ELIZA to converse over ARPANET, an early chatbot-to-chatbot exchange between an artificial therapist and an artificial patient.
PARRY demonstrated that expert systems could model personality and affect, not just diagnostic rules. But like MYCIN, it remained a research curiosity. Colby’s vision of AI as a psychiatric teaching tool and therapeutic adjunct would wait 50 years for computational power, ubiquitous connectivity, and a mental health access crisis to make it viable.
Why Expert Systems Failed
By the late 1980s, the expert systems approach hit fundamental limits:
Brittleness: Systems worked perfectly within their narrow domain but failed dramatically on edge cases. A single unexpected finding could derail the entire diagnostic process.
Knowledge acquisition bottleneck: Extracting rules from expert physicians was extraordinarily time-consuming and incomplete. Experts often could not articulate their reasoning explicitly.
Combinatorial explosion: Real-world medicine requires thousands of interacting rules. Managing complexity became unworkable.
Maintenance burden: Medical knowledge evolves continuously. Hand-coded rules became outdated, requiring constant programmer intervention.
Lack of learning: Expert systems could not improve from experience. Every new case provided no feedback to enhance future performance.
Deployment realities: As MYCIN demonstrated, technical performance did not address liability, regulation, integration, or trust.
The “AI Winter” of the late 1980s and 1990s arrived. Funding dried up. Researchers left the field. Companies removed “AI” from marketing materials. Medical AI seemed like a failed experiment.
The lesson: Rule-based approaches could not capture the complexity, uncertainty, and nuance of real clinical practice. Medicine needed a fundamentally different approach.
The Machine Learning Revolution (1990s–2000s)
From Rules to Data
The 1990s brought a decisive shift: instead of encoding expert rules manually, why not let algorithms learn patterns directly from data?
Machine learning, particularly supervised learning, offered a solution. Show an algorithm thousands of examples (X-rays labeled normal vs. abnormal, patient data labeled survived vs. died), and it learns to recognize patterns without explicit rule programming.
Key developments enabling this shift:
- Increased computational power: Moore’s Law made previously impractical computations feasible
- Digitization of medical data: Picture Archiving and Communication Systems (PACS), early electronic health records
- Algorithmic advances: Support Vector Machines (SVMs), decision trees, ensemble methods
Computer-Aided Detection (CAD) in Radiology
The first commercially successful medical AI applications emerged in radiology:
Computer-Aided Detection (CAD) for mammography:
- FDA began approving CAD systems in the late 1990s
- Designed to augment radiologist interpretation by flagging suspicious regions
- Became widely adopted in breast cancer screening
Initial promise: Retrospective studies suggested CAD could reduce false negatives (missed cancers).
Reality check: Prospective studies showed mixed results (Lehman et al., 2015). CAD increased recall rates (more women called back for additional imaging) without consistently improving cancer detection rates. Some studies suggested CAD reduced radiologist specificity without improving sensitivity.
Current status: Second-generation AI systems using deep learning show more promise. The lesson: first-generation medical AI often underperforms expectations in real-world deployment.
FDA Begins Regulating Medical AI
The 1990s-2000s saw FDA establish regulatory frameworks for software-based medical devices:
- 1990s: First CAD systems cleared through 510(k) pathway (substantial equivalence to existing devices)
- 2000s: FDA establishes guidance for computer-assisted detection devices
- Challenge: Rapid AI evolution outpaced regulatory frameworks designed for static medical devices
This regulatory evolution continues today, with FDA developing new frameworks for continuously learning AI systems.
Why This Era Mattered
The machine learning revolution established principles still guiding medical AI:
- Data-driven approaches could achieve good performance without explicit rule encoding
- Narrow, well-defined tasks (e.g., detecting lung nodules on CT) worked better than general diagnosis
- Human-AI configuration became an empirical question: assistance can improve, worsen, or leave performance unchanged depending on the task, interface, threshold, user, and workflow
- Prospective validation essential: Retrospective performance did not guarantee real-world utility
- Workflow integration remained challenging despite better algorithms
The limitation: Traditional machine learning required careful feature engineering (hand-crafted measurements and patterns). Algorithms could not learn complex representations directly from raw data.
The deep learning revolution would change that.
The Deep Learning Revolution (2010s): Imaging AI Comes of Age
AlexNet and the ImageNet Moment (2012)
In 2012, a deep convolutional neural network called AlexNet (Krizhevsky et al., 2012) won the ImageNet visual recognition challenge by a massive margin, halving the error rate of competing approaches.
This was not just incremental progress. It was a fundamental break from prior approaches. Deep learning could learn complex features directly from raw pixels without manual feature engineering. Suddenly, computer vision problems that had resisted decades of effort became tractable.
Implications for medical imaging:
Medical images (X-rays, CTs, MRIs, pathology slides, dermatology photos, retinal fundus images) are fundamentally visual pattern recognition problems. If deep learning could master general image classification, perhaps it could learn to detect pneumonia, identify cancers, or diagnose diabetic retinopathy.
Medical Imaging AI Explosion
The mid-2010s saw an explosion of medical imaging AI research:
- 2016: Dermatology AI matches dermatologists in skin cancer classification (Esteva et al., 2017)
- 2016: Deep learning detects diabetic retinopathy from retinal images (Gulshan et al., 2016)
- 2017: Radiology AI detects pneumonia from chest X-rays (Rajpurkar et al., 2017)
- 2018: Pathology AI analyzes prostate biopsies (Nagpal et al., 2019)
Papers proliferated. Venture capital flooded into medical AI startups. Headlines proclaimed “AI Will Replace Radiologists.”
The “AI Will Replace Radiologists” Debate
Geoffrey Hinton (deep learning pioneer) stated at a 2016 conference: “It’s quite obvious that we should stop training radiologists.”
This sparked fierce debate. Would AI replace radiologists? Should medical students avoid imaging specialties?
What actually happened:
- AI did not eliminate the radiology workforce
- AI systems entered specific radiology workflow tasks
- Radiologists increasingly use AI as assistive tools
- AI handles some straightforward screening tasks
- Complex cases still require radiologist expertise
- New roles emerged: radiologists curating datasets, validating AI, interpreting edge cases
The lesson for physicians:
The historical record does not support a single replacement narrative. Some tasks are automated, others are redistributed, and many tools create new work for verification, exception handling, quality assurance, and patient communication.
Fear replacement less. Focus on effective integration more.
Deep Learning’s Limitations in Medicine
Despite these advances, deep learning revealed serious limitations:
Interpretability problem: Neural networks do not ordinarily provide a faithful, clinically sufficient account of the internal basis for a prediction. A radiologist can articulate reasoning (“spiculated mass in the upper lobe with associated lymphadenopathy suggests malignancy”). A generated explanation or heat map may help review, but it is not automatically a causal account of the model’s computation.
Data hunger: Deep learning requires massive labeled datasets (thousands to millions of examples). Curating these datasets requires enormous physician time.
Brittleness to distribution shift: Models trained on data from Hospital A often perform poorly at Hospital B due to different imaging equipment, patient populations, disease prevalence, or documentation practices.
Adversarial vulnerability: Tiny, imperceptible changes to input images can fool deep learning models completely, a major patient safety concern (Finlayson et al., 2019).
Fairness and bias: AI systems inherit biases from training data. If training data under-represents certain populations, model performance may be worse for those groups (Obermeyer et al., 2019).
These limitations drive ongoing research and FDA regulatory evolution.
IBM Watson for Oncology: The Highest-Profile Failure
No discussion of medical AI history would be complete without IBM Watson for Oncology, arguably the most expensive, highest-profile medical AI failure to date.
The Promise
In 2011, IBM’s Watson defeated human champions on Jeopardy!, demonstrating impressive natural language processing capabilities. IBM launched Watson for Oncology in 2013, partnering with Memorial Sloan Kettering to train the system. IBM promised Watson would reshape medicine by:
- Analyzing millions of medical journal articles
- Synthesizing complex patient data from EHRs
- Recommending personalized, evidence-based cancer treatments
- Augmenting oncologist decision-making
Major cancer centers worldwide partnered with IBM: MD Anderson Cancer Center, Memorial Sloan Kettering, hospitals across India, China, and other countries. IBM invested billions. Expectations soared.
The Reality
Watson for Oncology became a prominent example of the gap between commercial ambition and clinical evidence. The strongest public record combines investigative reporting with retrospective concordance studies, rather than randomized patient-outcome evidence (Ross & Swetlitz, 2018):
Unsafe recommendations: Internal documents revealed Watson recommended treatments that would have harmed patients. In one case, Watson suggested administering chemotherapy to a patient with severe bleeding, a contraindicated, potentially fatal recommendation (Ross & Swetlitz, 2018).
Variable concordance: A meta-analysis of nine nonrandomized studies and 2,463 patients assessed agreement between Watson for Oncology and multidisciplinary teams. Overall concordance was 81.52% when both “recommended” and “for consideration” categories counted as concordant, with substantial variation by cancer, stage, and patient characteristics (Jie et al., 2021). Concordance measures agreement with a comparator, not treatment effectiveness or patient benefit.
Geographic inappropriateness: Watson recommended treatments unavailable in certain regions, ignoring local resource constraints and formulary restrictions. While some studies showed high concordance (93% in one Indian breast cancer study), Watson’s recommendations often included drugs not covered by local insurance or unavailable in certain formularies (Somashekhar et al., 2018).
Training data issues: Watson was trained primarily on synthetic cases created by Memorial Sloan Kettering physicians, not real-world patient data. It learned institutional preferences, not universal evidence (Strickland, 2019).
MD Anderson debacle: MD Anderson Cancer Center spent $62 million on Watson implementation before canceling the project in 2016, concluding it was not ready for clinical use (Strickland, 2019).
Widespread abandonment: By 2019, multiple health systems had stopped using Watson. IBM sold Watson Health assets in 2021.
What Went Wrong: Lessons for Physicians
Marketing ≠ Clinical Validation: Jeopardy! success does not translate to clinical competence. Demand prospective clinical trials, not just impressive demonstrations in controlled settings.
Reviewability matters: A recommendation is difficult to govern when clinicians cannot inspect its evidence basis, applicability, and uncertainty efficiently.
Training data matters immensely: Synthetic cases created by one institution do not represent the diversity of real clinical practice.
Clinical participation must extend beyond labeling: Clinicians, patients, informaticians, safety teams, and operational leaders must shape intended use, evaluation, escalation, and monitoring.
Geographic and institutional context matters: Treatment recommendations must account for local resources, formularies, patient populations, and practice patterns.
Commercial momentum is not evidence: Procurement and deployment can proceed faster than independent clinical validation.
Transparency and reporting: Watson’s failures remained largely hidden until investigative journalists uncovered them. Medical AI needs transparent reporting of failures.
The clinical bottom line:
Watson’s failure demonstrates why physicians must evaluate AI systems with the same rigor applied to new pharmaceuticals: demand prospective clinical trials, transparent reporting, independent validation, and post-market surveillance.
Do not accept vendor claims without evidence.
The Foundation Model Era (2020–Present)
Large Language Models Arrive
November 2022 brought ChatGPT, introducing millions of people (including physicians) to large language models (LLMs) (Brown et al., 2020). These “foundation models” could:
- Answer medical questions with apparently sophisticated reasoning
- Draft clinical notes
- Explain complex concepts
- Translate between languages
- Write code
- Summarize literature
Unlike narrow AI systems designed for specific tasks, LLMs demonstrated general capabilities across diverse domains.
Medical-Specific Foundation Models
Recognizing both promise and risks of general LLMs in medicine, researchers developed medical-specific models.
How Post-Training Shapes Medical LLM Behavior
Medical LLM performance cannot be attributed to RLHF alone. Modern post-training may combine supervised instruction tuning, domain adaptation, human or AI preference data, safety examples, and verifier-based reinforcement learning. Each method optimizes against its own demonstrations, labels, rubric, or reward. Clinician feedback can shape response behavior, but evaluator preference is not proof of clinical accuracy, safety, or patient benefit. Open post-training recipes illustrate this diversity, although their reported gains do not establish transfer to clinical work (Lambert et al., 2024, preprint; Lee et al., 2024).
Med-PaLM is an important correction to the simplified RLHF story. The published model used instruction prompt tuning, a parameter-efficient method based on medical-domain exemplars, rather than clinician-rated RLHF. Its authors attributed performance differences to model scale, prompting, instruction tuning, and medical instruction prompt tuning, while emphasizing that the system remained inferior to clinicians on important dimensions (Singhal et al., 2023). Clinical performance also depends on the base model, data, retrieval and tools, evaluation design, and deployment context.
Med-PaLM (Google, 2023):
- Adapted from Flan-PaLM using medical instruction prompt tuning
- Achieved passing scores on USMLE-style questions (Singhal et al., 2023)
- Improved clinical accuracy compared to general LLMs
GPT-4 in medicine (OpenAI, 2023):
- Demonstrated strong performance on medical licensing exams
- Used by physicians for literature review and summarization, clinical reasoning support, patient education
Challenges remain:
- Hallucinations: LLMs confidently generate plausible but incorrect information
- Bias: Inherit biases from training data
- Liability: Duties and allocation of responsibility remain fact-specific and jurisdiction-specific
- Privacy: Patient data security concerns
- Validation: How to validate general-purpose models across diverse clinical scenarios?
Current State: Promise and Uncertainty
As of August 2026, medical AI presents a mixed picture:
Undeniable progress: - FDA maintains a growing, non-exhaustive list of AI-enabled devices authorized through pathway-specific review (see the FDA AI-enabled medical-device list) - Radiology, pathology, dermatology, and ophthalmology include authorized and deployed tools, but adoption and evidence vary by product, institution, and use case - Clinical decision support systems are deployed in major EHR systems - Foundation models demonstrate unprecedented versatility
Persistent challenges: - Most AI systems remain narrow, fragile, and context-dependent - Integration into clinical workflow remains difficult - Liability rules vary across jurisdictions and continue to develop alongside clinical use - Bias and fairness concerns are well-documented but incompletely solved - Physician trust varies widely by specialty, generation, and prior experience
The fundamental question:
Are today’s foundation models genuinely different from previous AI waves, or are we in another hype cycle that will end in disillusionment?
Evidence suggests: Both. Foundation models represent real algorithmic breakthroughs. But transforming clinical practice requires solving the same deployment challenges that defeated MYCIN and Watson: integration, liability, trust, validation, and proving clinical benefit beyond technical performance metrics.
Key Lessons from History
Seven decades of medical AI teach clear lessons:
- Technical performance ≠ clinical adoption (MYCIN)
- Marketing hype ≠ clinical validity (Watson)
- Retrospective validation ≠ prospective utility (Google Flu Trends)
- Narrow, well-defined tasks work best (IDx-DR)
- Human-AI performance must be tested, not assumed (CAD in radiology)
- Physician involvement is essential (Watson’s failure)
- Regulatory frameworks evolve slowly (ongoing FDA challenges)
- Liability concerns shape adoption (MYCIN’s legal uncertainty)
- Workflow integration is harder than algorithmic development (expert systems)
- Patient safety must remain paramount (Watson’s unsafe recommendations)
The path forward:
Learn from failures. Demand evidence. Integrate thoughtfully. Maintain physician oversight. Prioritize patients.
History does not repeat, but it rhymes. The physicians who understand AI’s history will navigate its future most effectively.
Skill Degradation in the Age of AI: Historical Lessons
One important risk of medical AI is cognitive deskilling, the possible erosion of clinical expertise when clinicians repeatedly defer observation, interpretation, or judgment to an automated system.
The concern is plausible and has emerging empirical support, but its magnitude is not established across medicine. The evidence must distinguish three related problems: automation bias (overweighting a system’s recommendation), deskilling (loss of a previously acquired capability), and never-skilling (failing to acquire that capability during training).
Documented Skill Atrophy in Medical AI
Medical AI is following the same trajectory, with concerning evidence emerging from radiology and gastroenterology.
Computer-Aided Detection (CAD) in Radiology:
Early CAD systems for mammography promised to reduce missed cancers by flagging suspicious regions for radiologist review. Their clinical value proved inconsistent. A large community-practice analysis found no improvement in diagnostic accuracy with CAD (Lehman et al., 2015). Researchers also examined whether prompts altered how radiologists searched images and weighted evidence.
Experimental and observational work described interaction effects between CAD prompts and radiologist decisions, including errors related to inappropriate reliance (Alberdi et al., 2004). That study supports an automation-bias concern. It did not longitudinally demonstrate that radiologists’ pattern-recognition ability atrophied.
One proposed mechanism is attentional redistribution: when an algorithm highlights suspicious regions, a reader may devote less attention to unflagged areas. Whether repeated exposure produces durable skill loss depends on the interface, task, training, feedback, and opportunities for independent practice.
The Gastroenterology Study: Physicians Became Significantly Worse
Recent gastroenterology evidence shows both a risk signal and important uncertainty. In a retrospective multicenter observational study at four Polish centers, the adenoma detection rate during non-AI colonoscopy declined from 28.4% (226/795) before routine AI exposure to 22.4% (145/648) afterward, an absolute difference of 6.0 percentage points (Budzyń et al., 2025). The before-and-after design supports an association, not proof that AI exposure caused the decline.
The observed decline is a credible deskilling signal, but the design does not establish causation or inevitable harm.
Two later studies did not reproduce inevitable deterioration. A prospective single-center observational study found that non-AI detection performance was maintained after CADe implementation (Okumura et al., 2026). A prospective multicenter pragmatic trial involving 13 endoscopists and 5,013 colonoscopies found a temporary CADe-associated improvement among inexperienced endoscopists, but no significant upskilling or deskilling after CADe removal (Pedersen et al., 2026). The defensible conclusion is conditional: deskilling can occur, but it is not an established universal consequence of AI assistance.
Training the Next Generation: Medical Students Who Never Develop Foundational Skills
The skill question becomes especially important in medical education. If students learn alongside tools that generate differential diagnoses, treatment suggestions, and documentation, educators must decide when independent performance is required and how competence will be assessed without assistance.
Dhruv Khullar, writing in The New Yorker, described medical students’ discomfort about presenting AI-generated reasoning as their own (Khullar, 2025). This is a useful account of learner experience, not a controlled estimate of educational harm.
This raises profound questions about what medical training should accomplish in an age of clinical AI:
- If students can query AI for differential diagnoses, should they still memorize disease presentations?
- If algorithms interpret ECGs more accurately than cardiologists, what ECG interpretation skills should medical students master?
- If AI drafts clinical notes, how will students learn to integrate complex information into coherent clinical narratives?
- If radiology AI detects pathology automatically, what level of image interpretation competency should physicians maintain?
The risk is not that students use AI tools. It’s that they never develop robust foundational skills because AI assistance is always available. Then, when AI fails, is unavailable, or encounters edge cases beyond its training, there’s no physician expertise to fall back on.
Historical Lesson: Every Technology Changed What Physicians Knew
This pattern is not new. Each wave of medical technology transformed physician capabilities:
Calculators and clinical scoring systems: Physicians no longer perform complex risk calculations mentally. They input variables into validated scoring systems. Mental arithmetic skills declined, but risk stratification improved.
Electronic health records (EHRs): Physicians rely on structured templates, autocomplete, and copy-forward functionality. Documentation became standardized but less personalized. Narrative skills atrophied. Many physicians struggle to write coherent clinical summaries without EHR templates.
Automated laboratory analyzers: Physicians stopped performing manual blood counts and chemical assays. Lab interpretation skills improved (access to more data), but understanding of underlying physiology declined. Few physicians today could estimate hemoglobin from examining a peripheral blood smear.
Advanced imaging (CT, MRI): Physical examination skills declined as imaging became routine for diagnosis. Studies show physicians miss physical findings that would have been obvious to previous generations, because they order imaging instead of examining patients thoroughly.
Each technology brought benefits and redistributed work. The effects on competence varied by task, specialty, training, and implementation.
AI may produce a broader shift because it can affect pattern recognition, diagnostic reasoning, documentation, and decision-making under uncertainty at the same time.
The Unsolved Tension: Competence vs. Dependency
There’s no simple resolution to this tension.
Refusing to use AI to preserve traditional skills would harm patients if AI genuinely improves outcomes. But uncritical adoption creates dangerous dependencies and erodes the expertise needed when technology fails.
Medical educators, licensing bodies, and specialty societies are grappling with fundamental questions:
- What clinical skills are truly foundational, requiring mastery regardless of AI availability?
- What tasks can we safely delegate to algorithms, accepting that human competency will decline?
- How do we maintain physician capability to recognize and override algorithmic errors?
- What backup competencies must physicians preserve for when AI systems fail?
- How do we train medical students to use AI effectively without becoming entirely dependent on it?
No consensus exists. Medical schools are experimenting with different approaches: some integrate AI early in training, others restrict its use until students demonstrate independent competency. Some specialties (radiology, pathology) are redesigning curricula around AI-augmented practice. Others (surgery, primary care) emphasize skills that remain distinctly human.
The historical lesson is clear: Technology always changes physician capabilities. The question is not whether AI will transform medical expertise, but how we manage that transformation to preserve the irreplaceable elements of clinical judgment while embracing genuine improvements.
Physicians who understand this historical pattern, and who treat skill preservation as a measurable implementation objective, will be better positioned to maintain critical competencies while integrating AI thoughtfully.
The long-term effects of AI on physician expertise remain uncertain. A safer response is to define the capabilities clinicians must retain, test performance with and without assistance, monitor override behavior and error recovery, and design periodic independent practice where the consequences justify it.