AI in Medical Education: Curriculum, Competency, and the Future Physician
A 2025 scoping review of AI in undergraduate medical education (UME) identified 310 publications on the topic, with 52% appearing in the eight months following the November 2022 release of ChatGPT (Simoni et al., BMC Medical Education, 2025). Medical schools worldwide are scrambling to develop curricula for a technology that did not exist when most faculty completed their training. The evidence base for what to teach, when to teach it, and how to assess competency remains nascent.
A 2026 PRISMA-ScR review now provides a more granular inventory for undergraduate medical education: 54 studies from 22 countries yielded 564 competency-relevant statements, organized into seven domains, 37 competencies, and 170 learning objectives. The recurring emphasis was ethics, legal oversight, professionalism, clinical applications, critical appraisal of AI outputs, research and innovation, and AI theory or foundations (Hunt et al., 2026).
After reading this chapter, medical educators will be able to:
- Understand the five-domain framework for AI applications in medical education from the Macy Foundation report
- Evaluate integration models for AI across undergraduate, graduate, and continuing medical education
- Recognize the tension between AI-assisted learning and foundational skill development
- Apply the FACETS framework when reporting AI educational interventions and distinguish reporting completeness from evidence of educational effectiveness
- Identify the 23 AI competencies with strong Delphi consensus for physician training (from 27 candidate items)
- Develop faculty development strategies for AI education
- Compare international approaches to AI medical education (UK, EU, Asia-Pacific)
- Assess learners on AI-specific competencies using validated frameworks
Introduction
The integration of AI into clinical practice creates an educational imperative: physicians must learn to work effectively with AI tools they did not encounter during training. This challenge spans the educational continuum from undergraduate medical education (UME) through graduate medical education (GME) to continuing medical education (CME) for practicing physicians.
The evidence base for AI medical education remains limited. A comprehensive scoping review of AI in undergraduate medical education identified 310 publications, but methodological quality varied substantially and most studies preceded the large language model revolution of 2022-2023 (Simoni et al., BMC Medical Education, 2025). Educational frameworks are emerging, but consensus on core competencies, optimal integration timing, and assessment methods remains elusive.
Three fundamental tensions shape AI medical education:
Efficiency vs. skill development: AI tools improve diagnostic efficiency, but training with constant AI assistance may prevent development of independent clinical reasoning
Curricular space: Medical curricula are already overcrowded. Adding AI competencies requires either displacing existing content or integrating AI into existing courses
Faculty readiness: Most medical educators completed training before clinical AI deployment. They cannot teach what they have not learned
Part 1: The Educational Imperative
Why AI Education Cannot Wait
Physicians entering practice today encounter AI in multiple clinical contexts:
- Clinical decision support: Sepsis prediction, deterioration alerts, drug interaction warnings
- Diagnostic imaging: CAD systems in radiology, pathology, ophthalmology
- Documentation: Ambient AI scribes, auto-generated notes, coding suggestions
- Information retrieval: LLM-based literature search, clinical question answering
Without formal training, physicians develop ad hoc responses to AI ranging from uncritical acceptance to blanket rejection. Neither extreme serves patients. When clinical AI is misunderstood or inadequately implemented, new error modes emerge: alert fatigue, automation complacency, and over-reliance on inaccurate predictions. The Epic Sepsis Model, deployed across U.S. hospitals, illustrates these failure modes at scale (Wong et al., JAMA Internal Medicine, 2021).
The Competency Gap
A 2025 report from the Josiah Macy Jr. Foundation identified five domains of medical education where AI can be applied (Boscardin et al., 2025):
- Admissions: AI-assisted screening and evaluation of applicants
- Classroom-Based Learning: AI tutoring, content generation, and personalized learning
- Workplace-Based Learning: AI clinical decision support during training
- Assessment and Feedback: Automated evaluation and certification processes
- Program Evaluation and Research: AI-augmented educational outcomes analysis
Most current medical education addresses none of these domains systematically. Surveys consistently show that medical students and residents feel unprepared for AI-integrated practice, citing lack of formal training as the primary barrier.
The Publication Surge
Interest in AI medical education has grown rapidly. The scoping review by Simoni et al. documented a dramatic publication increase following ChatGPT’s release in November 2022: 52% of all AI medical education publications appeared in the eight months after ChatGPT’s public release (Simoni et al., 2025). This surge reflects both genuine educational need and publication opportunism, making critical appraisal of the literature essential.
Structured AI training can produce measurable gains. A pre-post study of 326 physicians across 39 countries found that a 1.5-hour training intervention was followed by higher scores when GPT-4 assistance was permitted: average clinical-competence scores changed from 56.9% to 77.6%, and pass rates from 6.4% to 58.6% (Qunaibi et al., 2026). Without a randomized no-training AI comparator, the study does not establish that training mattered more than the tool or isolate the training’s causal effect. A meta-analysis of 11 RCTs involving 786 medical students found no significant difference in theoretical knowledge between AI and traditional teaching, but reported higher practical-skill scores and satisfaction with AI teaching amid heterogeneous interventions and outcomes (Li et al., 2025). Educational effectiveness remains intervention-, learner-, task-, comparator-, and outcome-specific.
Part 2: Undergraduate Medical Education (UME)
Current Landscape
Medical schools face mounting pressure to incorporate AI into curricula. The chapter review did not identify a dedicated LCME AI competency mandate, but current standards and interpretations must be checked directly for each accreditation cycle. Professional organizations have begun filling the educational gap: AMA policy H-295.857 supports AI learning objectives, curricular toolkits, specialty modules, research, faculty development, and CME across the continuum (AMA PolicyFinder, 2025); the AAMC published versioned Principles for Responsible AI Use in Medical Education, with Version 2.0 developed in collaboration with AMEE, IAMSE, and APMEN (AAMC, 2025).
Existing Integration Models:
| Model | Description | Advantages | Challenges |
|---|---|---|---|
| Standalone Course | Dedicated AI curriculum (elective or required) | Comprehensive coverage, faculty specialization | Siloed knowledge, curricular space competition |
| Integrated Threads | AI woven into existing courses (physiology, clinical skills) | Contextual learning, efficient use of time | Requires faculty-wide training, inconsistent depth |
| Clinical Immersion | AI exposure during clerkships with supervised use | Real-world application, immediate relevance | Dependent on clinical site AI deployment, variable exposure |
| Simulation-Based | AI scenarios in simulation center | Safe practice environment, standardized exposure | Resource-intensive, may lack authenticity |
Recommended Curricular Framework (UME)
Based on the Macy Foundation domains, a phased approach aligns AI education with medical school progression:
Pre-Clinical Years (Years 1-2):
- Foundational concepts: How machine learning works, training data, overfitting, validation vs. deployment performance
- Basic statistics for AI: Sensitivity, specificity, PPV, NPV, AUC-ROC interpretation in clinical AI contexts
- Ethical foundations: Algorithmic bias, data privacy, informed consent for AI-assisted care
- Critical appraisal: Evaluating AI studies using modified CONSORT-AI and SPIRIT-AI guidelines
Clinical Years (Years 3-4):
- Specialty-specific AI tools: Exposure to AI in radiology, pathology, cardiology during rotations
- Workflow integration: Observing how AI fits (or conflicts with) clinical workflows
- Human-AI teaming: Practicing appropriate reliance, recognizing when to override
- Documentation AI: Understanding ambient scribes, auto-generated notes, verification responsibilities
Integration Barriers
Faculty Preparedness:
The primary barrier to AI education in medical schools is faculty unfamiliarity. Surveys consistently indicate that medical educators have received little to no formal training in AI, and most do not feel prepared to teach these concepts. The Simoni et al. scoping review noted this gap as a recurring theme across the literature.
Curricular Crowding:
Medical school curricula already contain more content than students can master. Adding AI education without displacing other material risks superficial treatment. Integration into existing courses (anatomy faculty discussing AI in imaging, pathology faculty discussing computational pathology) spreads the burden but requires coordinated faculty development.
Assessment Challenges:
No single widely adopted, validated instrument assesses the full range of medical-student AI competencies. Existing instruments and local rubrics measure selected knowledge, attitudes, confidence, or performance. FACETS is a reporting framework for educational-AI studies, not an assessment taxonomy or competency test.
The LLM Challenge in Medical Education
Large language models create new educational challenges and opportunities:
Opportunities:
- Personalized tutoring and practice questions
- Clinical reasoning partners for case discussions
- Literature synthesis for learning
- Writing assistance for reports and applications
Threats:
- Academic integrity concerns (LLM-generated assignments)
- Over-reliance preventing deep learning
- Hallucinated medical information
- Reduced development of writing skills
Medical schools are developing LLM policies ranging from prohibition to required use with disclosure. The emerging consensus favors integration with transparency: students may use LLMs as learning tools but must disclose use and verify outputs.
LLM-Based Virtual Patients for Clinical Training
A 2026 systematic review of 39 studies examined LLM-powered virtual-patient systems for history-taking training (Li & Lutfi, 2026). Reported top-k accuracy, hallucination, and usability values came from heterogeneous systems, tasks, denominators, and evaluation methods. They should not be combined into a single performance profile or treated as evidence that virtual patients improve real clinical care.
Enhancement techniques that improved performance:
- Role-based prompts and few-shot learning
- Knowledge graph integration (improved top-k accuracy by 16%)
- Multiagent frameworks
- Multimodal inputs (speech, imaging)
Critical limitation: The review found that most systems focused on specific diseases and relatively few addressed multimorbidity. This limits transfer to the complex, comorbid patients who dominate practice (see Primary Care and Internal Medicine).
Implication for educators: LLM virtual patients offer promising tools for early history-taking training, but cannot replace exposure to multimorbid patients. Programs using virtual patients should supplement with cases featuring overlapping conditions, polypharmacy, and diagnostic uncertainty.
Part 3: Graduate Medical Education (GME)
The ACGME Landscape
The Accreditation Council for Graduate Medical Education (ACGME) has not issued formal AI competency requirements, but existing milestones implicitly encompass AI:
- Practice-Based Learning and Improvement: Requires use of technology for practice improvement
- Systems-Based Practice: Requires understanding of healthcare systems, including technology
- Medical Knowledge: Requires application of biomedical sciences, including computational tools
Current ACGME Common Program Requirements should be checked directly rather than inferred from consultation documents. In its November 2025 second stakeholder survey for the Common Program Requirements major revision, ACGME asked whether programs had training and policies for six AI uses and whether ACGME should create requirements. That survey documents consultation, not a final competency mandate or a forecast of final language (ACGME, 2025). A 2025 NEJM review proposed the DEFT-AI framework for supervising AI use in clinical training, identifying three failure modes: “never-skilling” (avoidance), “mis-skilling” (over-simplified learning from AI), and “deskilling” (competence erosion through overreliance). The framework distinguishes a “centaur” approach, in which humans delegate selected tasks while retaining oversight, from a “cyborg” approach involving closer human-AI integration (Abdulnour et al., 2025). These are conceptual approaches, not ACGME-prescribed training tracks.
Several specialty societies have begun developing AI-specific milestones, with radiology and pathology leading.
Specialty-Specific Integration
AI integration varies dramatically by specialty, reflecting differences in tool maturity and workflow impact:
| Specialty | AI Integration Status | Key Educational Priorities |
|---|---|---|
| Radiology | Mature (high volume of FDA-authorized AI-enabled devices; see FDA AI-Enabled Medical Device List) | CAD integration, deskilling prevention, AI triage |
| Pathology | Growing (computational pathology) | Digital slide analysis, quantitative assessment |
| Cardiology | Moderate (ECG AI, echo quantification) | Wearable AI, automated measurements |
| Emergency Medicine | Emerging (sepsis, deterioration) | Critical appraisal, workflow integration |
| Primary Care | Growing (documentation, screening) | LLM documentation, risk stratification |
| Surgery | Nascent (robotics, imaging guidance) | Intraoperative AI, outcome prediction |
The De-Skilling Problem
A critical concern in GME is cognitive deskilling: residents trained with constant AI assistance may have difficulty maintaining independent diagnostic skills when AI fails or is unavailable. Automation bias and inappropriate reliance are documented, but generalized longitudinal deskilling across residency training has not been established.
Evidence from Radiology:
Radiology research documents automation bias, including performance changes when readers receive incorrect computer advice. That evidence supports teaching independent interpretation and AI-failure recognition, but it does not establish that radiologists trained with CAD from early residency become broadly less competent. The evidence boundaries and practical safeguards are reviewed in the radiology chapter’s deskilling section.
Gastroenterology Evidence:
In a retrospective observational study at four Polish endoscopy centers, the adenoma detection rate during unassisted colonoscopy declined from 28.4% before routine AI exposure to 22.4% afterward, an absolute difference of 6.0 percentage points. The analysis included 1,443 unassisted examinations, 795 before and 648 after AI introduction (Budzyń et al., 2025). This result is a signal of possible behavioral deskilling, not proof that AI exposure caused the decline. The before-and-after observational design remains vulnerable to selection bias, confounding, temporal changes, and center-level differences.
Mitigating De-Skilling in GME
Training programs must balance AI proficiency with independent skill development:
Illustrative Staged Competency Model:
The sequence below is a program-design option, not a validated universal schedule or an accreditation requirement. Programs should set access according to the learner’s demonstrated competence, the clinical task, supervision, and the consequences of error rather than postgraduate year alone.
| Training Phase | AI Exposure | Rationale |
|---|---|---|
| Foundational phase | Limited or delayed access for selected assessment tasks | Establish observable unaided performance where the skill itself is an educational objective |
| Supervised integration | Interpret first, then compare with AI | Teach verification, disagreement analysis, and calibrated reliance |
| Practice integration | Workflow-specific use with periodic unaided assessment | Maintain independent capability while learning actual clinical systems |
Practical Strategies:
AI-free rotations: Dedicated blocks where trainees interpret without AI, especially for foundational rotations
Interpret-then-compare: Trainees form independent assessment before viewing AI output, documenting reasoning
AI failure case conferences: Periodic review of cases where AI failed, building recognition of AI limitations
Competency gating: Assessment of independent skills before unsupervised AI-assisted work when the local program determines that gating is appropriate
GME Program Director Responsibilities
Program directors face new responsibilities for AI education:
- Curriculum development: Ensuring AI content addresses specialty-specific needs
- Faculty development: Preparing attendings to teach AI concepts
- Assessment: Developing and implementing AI competency evaluation
- Workflow design: Structuring AI exposure to reduce deskilling risk
- Documentation: Tracking AI-related competency milestones
Part 4: Continuing Medical Education (CME)
The Practicing Physician Challenge
Practicing physicians face unique AI education challenges:
- Time constraints: Limited availability for additional education
- Variable baseline: Ranging from no AI experience to daily AI use
- Immediate application: Need practical skills, not theoretical foundations
- Regulatory requirements: Evolving malpractice and compliance considerations
CME Framework for AI
Essential CME Topics:
| Topic | Content | Urgency |
|---|---|---|
| AI Tools in Your Specialty | Overview of FDA-cleared and emerging tools | High |
| Critical Appraisal | Evaluating AI performance claims | High |
| Liability and Documentation | Legal requirements for AI use | High |
| Workflow Integration | Practical implementation strategies | Medium |
| Patient Communication | Explaining AI to patients | Medium |
| Emerging Technologies | Future AI capabilities | Low |
Delivery Models
Effective CME Formats:
- Case-based modules: AI decision points embedded in clinical cases
- Simulation exercises: Hands-on practice with AI tools in safe environment
- Conference workshops: Specialty-specific AI sessions at annual meetings
- Online self-paced: Accessible, flexible, but requires self-discipline
- Institutional training: Mandatory training for new AI tool deployment
The Documentation Imperative
For practicing physicians, documentation of AI-assisted care should follow the institution’s policy, the product’s intended use, the clinical action taken, and applicable law. No universal rule requires every AI output or probability to be copied into the medical record.
- When to document AI use: When the information is clinically material to continuity of care, explains a decision, or is required by local policy
- What to document: The relevant clinical assessment, action, and reasoning, with enough context for another clinician to understand the decision
- Override documentation: Record clinically meaningful disagreement when it changes care, without treating the record as a defensive transcript of every alert
Illustrative Documentation, Not a Legal Safe Harbor:
“AI decision support flagged high sepsis risk (92% probability). Clinical assessment: patient afebrile, hemodynamically stable, WBC 11.2, lactate 1.4. AI recommendation not followed; clinical presentation inconsistent with sepsis. Plan: continue monitoring, repeat labs in 6 hours.”
This example illustrates a concise clinical rationale. Appropriate documentation can support continuity, audit, and retrospective review, but it does not determine civil liability or protect a clinician merely because a template was completed. Liability remains fact-specific and jurisdiction-specific.
Part 5: Competency Assessment
The 23 AI Competencies
A 2022 Delphi consensus study identified 27 candidate competencies; 23 achieved strong consensus for physicians using AI in clinical practice, organized across knowledge, skills, and attitudes (Caliskan et al., 2022):
Knowledge Domain (8 competencies):
| Competency | Description |
|---|---|
| K1 | Understand basic AI/ML concepts (supervised, unsupervised, reinforcement learning) |
| K2 | Understand data requirements (training, validation, test sets) |
| K3 | Understand performance metrics (sensitivity, specificity, AUC, calibration) |
| K4 | Understand limitations (overfitting, distribution shift, adversarial attacks) |
| K5 | Understand ethical considerations (bias, fairness, transparency) |
| K6 | Understand regulatory frameworks (FDA, CE marking, liability) |
| K7 | Understand data privacy requirements (HIPAA, GDPR, consent) |
| K8 | Understand specialty-specific AI applications |
Skills Domain (9 competencies):
| Competency | Description |
|---|---|
| S1 | Interpret AI outputs in clinical context |
| S2 | Recognize AI failure modes and limitations |
| S3 | Integrate AI recommendations with clinical judgment |
| S4 | Communicate AI role to patients appropriately |
| S5 | Document AI use in medical records |
| S6 | Evaluate AI tools for clinical adoption |
| S7 | Participate in AI governance and oversight |
| S8 | Monitor AI performance post-deployment |
| S9 | Override AI appropriately when indicated |
Attitudes Domain (6 competencies):
| Competency | Description |
|---|---|
| A1 | Maintain appropriate skepticism toward AI claims |
| A2 | Embrace continuous learning as AI evolves |
| A3 | Prioritize patient safety over efficiency gains |
| A4 | Advocate for equitable AI that reduces disparities |
| A5 | Support transparent AI development and deployment |
| A6 | Accept responsibility for AI-assisted decisions |
The FACETS Reporting Framework
The BEME Guide No. 84 proposed FACETS as a structured way to report AI innovations in medical education (Gordon et al., 2024). FACETS improves description and comparability; it is not a validated effectiveness instrument or learner-competency assessment.
| FACETS element | Reporting question | Detail to make explicit |
|---|---|---|
| Form | What form does the AI application take? | Tutor, simulator, assessment aid, content generator, or another form |
| Use case | What educational task or problem is addressed? | The intended learner action and educational purpose |
| Context | Where and for whom is the intervention used? | Learner stage, profession, specialty, institution, and setting |
| Education form | How is teaching, learning, or assessment organized? | Course, workshop, workplace learning, simulation, or self-study |
| Technology | What system is actually being evaluated? | Model, data inputs, version, interface, and relevant technical function |
| SAMR | How does technology change the educational activity? | Substitution, augmentation, modification, or redefinition |
Assessment Methods
Current Assessment Tools:
| Method | What It Measures | Limitations |
|---|---|---|
| Multiple choice | Knowledge recall | Does not assess application |
| Case vignettes | Theoretical decision-making | Artificial context |
| Simulation | Performance in controlled environment | Resource-intensive, may lack authenticity |
| OSCE stations | Standardized clinical performance | Limited by scenario design |
| Workplace assessment | Real-world performance | Dependent on clinical AI availability |
| Portfolio | Reflective practice | Subjective evaluation |
Recommended Multi-Modal Assessment:
- Knowledge assessment: Written examination on AI concepts (K1-K8)
- Simulation assessment: AI-integrated scenarios testing appropriate use (S1-S5)
- Workplace observation: Attending evaluation of AI integration (S1-S9)
- Reflective portfolio: Documentation of AI learning and challenges (A1-A6)
Gaps in Assessment
No validated, widely-adopted assessment tool exists for AI competency in clinical practice. The Delphi competencies provide framework, but operationalization requires:
- Standardized case libraries with AI decision points
- Rubrics for evaluating appropriate AI use
- Benchmarks for competency levels by training stage
- Methods for assessing attitude and professional identity formation
Part 6: Faculty Development
The Faculty Gap
Medical educators cannot teach what they do not know. Surveys consistently show:
- Most medical faculty have received little to no formal AI training
- Most do not feel prepared to teach AI concepts
- Most faculty learned about AI through media coverage, not formal education
This gap threatens curricular implementation: even well-designed AI curricula fail if faculty cannot deliver them effectively.
Faculty Development Framework
Tier 1: AI Literacy (All Faculty)
Essential for all clinical educators:
- Basic AI concepts (what is machine learning, how do clinical decision support systems work)
- Limitations and failure modes
- Ethical and legal considerations
- How to discuss AI with learners
Tier 2: AI Integration (Course Directors, Clerkship Directors)
For faculty designing and implementing curricula:
- Curricular design for AI content
- Assessment of AI competencies
- Integration with existing courses
- Simulation and case development
Tier 3: AI Expertise (Specialty Champions)
For faculty leading institutional AI education:
- Deep technical knowledge
- Research in AI medical education
- Institutional governance participation
- External collaboration and advocacy
Delivery Models for Faculty Development
| Format | Advantages | Best For |
|---|---|---|
| Grand rounds | Reaches many faculty, low time commitment | Tier 1 awareness |
| Workshops | Hands-on practice, discussion | Tier 1-2 skill building |
| Online modules | Flexible, asynchronous | Tier 1 foundational knowledge |
| Fellowships | Deep expertise development | Tier 3 champions |
| Learning communities | Peer support, iterative improvement | All tiers, ongoing development |
| Industry partnerships | Access to tools and expertise | Tier 2-3 practical skills |
Institutional Strategies
Create incentives for AI education:
- Protected time for AI curriculum development
- Promotion credit for AI teaching innovation
- Funding for AI education research
- Recognition through teaching awards
Leverage external resources:
- Professional society educational materials
- Vendor training on specific tools
- Collaboration with computer science and engineering faculty
- External courses and certificates
Part 7: International Perspectives
United Kingdom
The UK has taken a coordinated approach to AI medical education through Health Education England (HEE) and the NHS Topol Review.
Topol Review (2019):
- Called for AI to be core to health professional education
- Recommended competency frameworks across professions
- Recommended preparing the workforce for a digital future through education, training, and protected time for development (NHS Topol Review)
Current implementation should be verified program by program:
- The Topol Digital Fellowship program provides participating health professionals with time, support, and training to lead digital-health work.
- Undergraduate modules, postgraduate requirements, and faculty-development offerings vary across institutions and should not be inferred from the 2019 review alone.
- A national strategy can establish direction without proving that every medical school or training program has implemented the same curriculum.
European Union
The EU AI Act creates a regulatory context for organizational AI literacy. Article 4 directs providers and deployers to take measures, to their best extent, to ensure sufficient AI literacy among staff and others operating AI systems on their behalf, with training calibrated to knowledge, experience, context, and affected persons (Regulation (EU) 2024/1689, Article 4).
Educational Implications:
- AI literacy measures should be proportionate to the system, user, setting, and affected population.
- High-risk medical-device classification does not itself create one universal medical-school curriculum.
- Product instructions, institutional governance, professional duties, and applicable national law must be reviewed together.
- CE marking and EU AI Act obligations should not be reduced to a generic claim that a particular course or credential is legally required.
Country-level verification:
Germany, France, the Netherlands, and Nordic countries all host digital-health and AI initiatives, but labels such as “emerging” and “advanced” conceal variation among schools, professions, and clinical sites. A reliable comparison should identify the named program, responsible institution, learner population, curriculum content, assessment method, and current source rather than rank countries from general reputation.
Asia-Pacific
Singapore: National AI and workforce-development initiatives create opportunities for health-professions education. Claims about curricular integration should be tied to a named institution, current syllabus, or official program record.
South Korea: Technology infrastructure and government-supported AI initiatives can support medical-education programs, but adoption and assessment vary by institution.
Japan: Medical-society and university initiatives should be described through current program sources. Demographic need alone does not demonstrate curricular effectiveness.
Australia/New Zealand:
- Royal Australian and New Zealand College of Radiologists (RANZCR) leading in radiology AI education
- Multi-society AI statements (see radiology chapter)
- Rural health focus for AI deployment
Low- and Middle-Income Countries (LMICs)
AI medical education in LMICs faces distinct challenges:
- Infrastructure limitations: Inconsistent technology access
- Faculty scarcity: Fewer trained educators, competing priorities
- Relevance questions: AI developed in high-income countries may not transfer
- Opportunity: Locally appropriate delivery models can widen access when they are designed for available infrastructure, language, faculty capacity, and clinical priorities
Promising approaches:
- Mobile-first educational platforms
- AI tools designed for LMIC contexts
- South-South collaboration and knowledge sharing
- Integration with existing telemedicine initiatives
Clinical Scenarios
The following cases are fictional teaching scenarios. Their learners, institutions, performance data, schedules, and outcomes are illustrative, not reports of observed people or programs. They are preserved to support discussion and should not be cited as clinical evidence.
Case: Third-year medical student on internal medicine clerkship uses ChatGPT to generate differential diagnoses for every patient. When asked to explain pathophysiology during rounds, the student struggles to articulate reasoning, relying on memorized outputs rather than understanding.
Attending observation: The student provides comprehensive differentials but cannot engage in Socratic discussion about mechanism, cannot prioritize diagnoses based on clinical likelihood, and becomes uncertain when asked follow-up questions that require synthesis.
Discussion:
What is happening?
One plausible interpretation is that the student is using the LLM as a cognitive crutch rather than a learning tool. Other explanations, including gaps in prior knowledge, supervision, language, or psychological safety, should be assessed before attributing the difficulty to AI use. The educational concern is that repeated outsourcing could reduce opportunities for deliberate practice in generating, prioritizing, and defending a differential diagnosis.
Why is this problematic?
Skill development: Clinical reasoning requires repeated practice. Outsourcing prevents the pattern recognition and knowledge organization that define clinical expertise.
Verification inability: Without understanding pathophysiology, the student cannot verify LLM outputs for accuracy or appropriateness.
Transfer failure: When LLM is unavailable (oral exams, bedside without device), the student lacks functional competency.
Professional identity: Medicine requires independent judgment. Dependence on AI early in training may prevent development of professional autonomy.
Educational intervention:
Explicit expectations: Clarify that LLMs may be used for learning (studying concepts) but not for clinical reasoning tasks that students are expected to develop
Process transparency: Require students to show their reasoning process, not just conclusions
LLM-free assessments: Evaluate clinical reasoning without AI access to establish baseline competency
Constructive use: Guide appropriate LLM use (explaining concepts, practice questions) vs. inappropriate use (generating clinical assessments)
Metacognitive discussion: Help student recognize the difference between having an answer (LLM output) and understanding the answer (clinical reasoning)
Case: Fourth-year radiology resident has read chest X-rays with CAD assistance throughout training. Program director notes that during independent reading sessions (CAD disabled for competency assessment), the resident misses significantly more findings than peers who trained with CAD-free rotations in PGY-2.
Synthetic performance data for discussion:
| Condition | Resident (CAD-trained) | Peers (Mixed training) |
|---|---|---|
| Independent sensitivity | 72% | 89% |
| CAD-assisted sensitivity | 94% | 95% |
| False positive rate (independent) | 18% | 8% |
Discussion:
What happened?
The pattern raises concern about automation complacency, but the table does not establish causation. Case mix, supervision, prior experience, test conditions, reader variability, and the validity of the assessment all require review. The educational task is to determine whether independent reasoning and pattern recognition were practiced and measured adequately.
Root cause analysis:
Foundational exposure: Determine whether the resident had enough supervised unaided interpretation to establish baseline skill
Workflow design: CAD output visible concurrent with images, creating anchoring before independent assessment
Assessment gap: No formal testing of independent skills until PGY-4
Lack of AI-failure education: Resident never systematically reviewed cases where CAD failed
Remediation approach:
CAD-free remediation period: A locally determined period of independent reading with expert feedback, based on demonstrated needs rather than a universal four-week prescription
Interpret-then-compare protocol: Form impression before viewing CAD, document reasoning
CAD failure case review: Systematic exposure to AI failure modes
Competency gating: Demonstrate independent proficiency before resuming CAD use
Ongoing assessment: Periodic CAD-free competency checks at intervals justified by the program and task risk
Program-level changes:
- Foundational phase: Preserve enough unaided interpretation to measure core skills
- Supervised integration: Use CAD as a second reader after an independent assessment when educationally appropriate
- All phases: Review AI failures and disagreements at a frequency supported by local case volume
Case: A fictional founding curriculum dean is planning a new medical school opening in 2027. The school leadership has committed to “AI-integrated education from day one.” The planning team must design the AI curriculum component.
Constraints:
- 4-year MD program, LCME accreditation required
- Faculty recruited from traditional medical schools (limited AI expertise)
- Budget for technology but not unlimited
- Regional clinical affiliates have varying AI deployment
Illustrative Curriculum Design Framework:
Phase 1: Pre-Clinical Years (Years 1-2)
Year 1 (Foundations):
- Module 1: Introduction to AI in Healthcare (8 hours)
- What is AI/ML, basic concepts
- Current applications overview
- Why physicians need AI literacy
- Integrated content:
- Biostatistics: AI performance metrics (sensitivity, specificity, AUC)
- Ethics: Algorithmic bias, consent, accountability
- Anatomy: AI in imaging (brief exposure)
- Assessment: Knowledge-based examination
Year 2 (Applications):
- Module 2: Clinical AI Tools (12 hours)
- Clinical decision support systems
- Imaging AI (radiology, pathology)
- Documentation AI
- Critical appraisal of AI studies
- Integrated content:
- Pharmacology: Drug interaction AI
- Pathophysiology: Risk prediction models
- Epidemiology: AI in public health
- Assessment: Case-based evaluation with AI decision points
Phase 2: Clinical Years (Years 3-4)
Year 3 (Clerkships):
AI exposure during rotations: Supervised use of clinical AI
Clerkship-specific objectives:
- Medicine: Sepsis prediction, deterioration algorithms
- Surgery: Risk stratification, imaging AI
- Pediatrics: Growth prediction, screening tools
- Ob/Gyn: Fetal monitoring AI, documentation
Longitudinal thread: Monthly AI case discussions across clerkships
Assessment: Workplace-based assessment of AI integration
Year 4 (Acting Internship/Electives):
Capstone project: AI evaluation or implementation project
Elective: Advanced AI in [specialty] (specialty-specific)
Preparation for residency: Documentation practices, liability awareness
Assessment: Portfolio demonstrating AI competency
Faculty Development Plan:
- Year -1 (before opening): Intensive faculty AI training
- Ongoing: Monthly faculty development sessions
- Champions: Recruit 2-3 faculty with AI expertise for leadership
Technology Requirements:
- Simulation center with AI-integrated scenarios
- Access to clinical AI tools for demonstration
- LLM policy and approved tools for educational use
Illustrative Success Measures:
- Student competency on assessment instruments with documented validity evidence for the intended use, where such instruments are available
- Faculty confidence in teaching AI
- Graduate preparedness surveys
- Alignment with current LCME requirements and institutional learning objectives, without implying that LCME has a dedicated AI-accreditation standard
The Path Forward
AI medical education remains in its early stages. Key priorities for the field:
Research Needs:
- Validated assessment tools for AI competency
- Longitudinal studies of AI-trained physicians’ performance
- Comparative effectiveness of educational models
- De-skilling prevention strategies
- Stakeholder prioritization of proposed UME AI competencies, because the 2026 taxonomy is broad and not yet an accreditation standard
Implementation Priorities:
- Faculty development at scale
- Integration with accreditation standards
- Specialty-specific competency frameworks
- CME for practicing physicians
Policy Advocacy:
- LCME guidance on AI competencies
- ACGME milestones for AI in specialty training
- Role-specific continuing education tied to the actual tools and workflows physicians use
- Funding for AI medical education research
What is the Macy Foundation 5-domain framework for AI in medical education?
The five domains are: admissions, classroom-based learning, workplace-based learning, assessment/feedback, and program evaluation/research. It guides comprehensive AI curriculum integration.
What are the 23 AI competencies for physicians?
A Delphi study identified 27 candidate competencies; 23 achieved strong consensus across knowledge, skills, and attitudes domains. They include AI evaluation, bias recognition, and clinical integration skills.
What is the FACETS framework?
FACETS is a reporting framework proposed in BEME Guide 84. It prompts authors to report the AI form, use case, context, education form, technology, and SAMR level. It is not a validated assessment instrument or an acronym for fidelity, acceptability, cost, effectiveness, transferability, and sustainability.
Will AI cause deskilling in medical trainees?
Automation bias and inappropriate reliance are documented, and one observational colonoscopy study found lower unassisted adenoma detection after routine AI exposure. Generalized longitudinal deskilling across medical training remains insufficiently established, so programs should measure both unaided and assisted performance.
How should AI be taught in medical school?
Integration across UME, GME, and CME stages is recommended. Focus on critical appraisal (when to trust AI), technical literacy (what AI can and cannot do), and ethical considerations.
AI education must preserve independent clinical reasoning, not replace it with AI dependence
Staged competency development: Build foundational skills before integrating AI tools, then measure both unaided and assisted performance
Faculty development is prerequisite: Invest in training the trainers before curricular change
Deskilling risk requires measurement: Automation bias is documented, while generalized longitudinal deskilling remains insufficiently established
Assessment remains underdeveloped: Multi-modal approaches are needed, and no single widely adopted instrument covers the full competency range
International frameworks provide models: Adapt to local context while learning from global experience
The curriculum is already crowded: Integration into existing courses more feasible than standalone additions
LLMs change access and policy questions: Policies for appropriate educational use, disclosure, privacy, and assessment are essential
CME should be versioned to practice: Practicing physicians need training tied to the actual tools, versions, and workflows they use
Equity matters: AI education should address disparities, not entrench them
Additional Resources
Key Publications:
- Boscardin CK, et al. (2025). Preparing Students for an AI-Enhanced Health Care Future. Academic Medicine. https://doi.org/10.1097/ACM.0000000000006107
- Hunt VM, et al. (2026). What the AI era doctor should know: a scoping review of proposed artificial intelligence competencies for medical education. npj Digital Medicine. https://doi.org/10.1038/s41746-026-02761-9
- Simoni AH, et al. (2025). AI in medical education: a scoping review. BMC Medical Education. https://doi.org/10.1186/s12909-025-08188-2
- Gordon M, et al. (2024). A scoping review of artificial intelligence in medical education: BEME Guide No. 84. Medical Teacher. https://doi.org/10.1080/0142159X.2024.2314198
- Caliskan SA, et al. (2022). AI competencies for physicians: a Delphi study. PLoS ONE. https://doi.org/10.1371/journal.pone.0271872
Professional Society Resources:
- AMA Digital Health Initiative: AI curriculum resources
- AAMC Core Entrustable Professional Activities (EPAs): AI integration guidance
- Specialty societies: See individual specialty chapters for society-specific resources
Cross-references:
- Diagnostic Imaging, Radiology, and Nuclear Medicine for deskilling evidence and cognitive bias discussion
- Healthcare Policy and AI Governance for regulatory context
- Physician AI Liability and Regulatory Compliance for legal frameworks