AI and Global Health Equity

Specialist access, digital infrastructure, research capacity, and disease burden are distributed unevenly within and across countries. AI can extend selected services, but performance can change with prevalence, language, devices, protocols, referral pathways, and population characteristics. In 2025, the International Telecommunication Union estimated that 2.2 billion people remained offline, most in low- and middle-income countries (ITU, 2025). Global-health AI succeeds only when the technical system, care pathway, infrastructure, governance, financing, and local workforce function together.

Learning Objectives

After reading this chapter, clinicians will be able to:

  • Understand the dual potential of AI: reducing vs. exacerbating health disparities
  • Evaluate AI applications in low- and middle-income countries (LMICs)
  • Recognize infrastructure, data, and resource constraints in global settings
  • Assess telemedicine and AI-enabled remote diagnostics for underserved populations
  • Identify algorithmic bias and its impact on health equity
  • Navigate ethical considerations for AI deployment in resource-limited settings
  • Advocate for equitable AI development and deployment globally

Clinical Context

AI has substantial potential for global health, including screening, acquisition guidance, translation, documentation, decision support, surveillance, and workforce education. Yet model availability does not resolve electricity, connectivity, maintenance, payment, referral capacity, data rights, language coverage, or workforce training. The equity question is empirical: who gains access, who bears errors and costs, and what happens after the pilot ends?

Key Applications

What Has Evidence Under Defined Conditions: - National-scale mobile maternal messaging and helpdesk infrastructure through South Africa’s MomConnect, a digital-health comparator rather than proof of AI efficacy - WHO-recommended computer-aided chest-radiography products for tuberculosis screening, subject to product, population, threshold, and workflow requirements - Prospective studies of autonomous retinopathy screening and image-acquisition guidance for specified populations and intended uses - Context-specific LLM benchmarks, simulated-vignette trials, and retrospective clinical-workflow safety evaluations

What Does Not Work (Yet): - Systems deployed without local acceptance testing or a functioning referral pathway - Cloud-dependent workflows where connectivity, latency, affordability, or downtime is unacceptable - Device-dependent diagnostics without an affordable, maintainable acquisition chain - Donor pilots without local ownership, financing, technical support, monitoring, and an exit plan - Data extraction partnerships with no local benefit-sharing (ethical violations, community mistrust)

What’s Uncertain: - Federated learning for privacy-preserving multi-country training (promising but untested at scale) - Edge AI for offline diagnostics (technically feasible, regulatory pathways unclear) - AI for neglected tropical diseases (limited research funding, small datasets)

Critical Insights

  • Inverse Care Law Applied to AI: AI benefits those who need it least (HIC populations) while underserving those who need it most (LMIC populations)
  • Adoption gap: A 50-country physician survey found high AI awareness and optimism but much lower clinical use, with formal training and institutional AI availability strongly associated with adoption. Global AI equity depends on workforce training and local implementation capacity, not model access alone.
  • Evidence gap: Research volume and funding do not consistently match disease burden, infrastructure constraints, or local priorities; a universal 90/10 estimate is not established
  • Data Colonialism Risk: Tech companies extract LMIC patient data for commercial AI, provide no local benefits
  • Algorithmic Bias Impact: Measurement bias, representation gaps, label choices, prevalence shifts, and workflow differences can produce unequal performance
  • Infrastructure Reality: ITU estimated 2.2 billion people offline in 2025, while the 2026 SDG 7 report estimated 655 million people without electricity; national and facility conditions vary (ITU, 2025; World Bank, 2026)

Clinical Bottom Line

Global health equity requires intentional design: AI must be built for the actual infrastructure and care pathway, with local partners and governance, and evaluated on relevant populations and endpoints. Without equity-focused design and measurement, AI can widen disparities. Physician role: Advocate for equitable development, demand local acceptance testing, support maintainable tools and open standards where appropriate, partner with LMIC institutions, and reject data extraction without a negotiated public or local benefit.

Essential Reading

  • Wahl et al. (2018). “Artificial intelligence (AI) and global health: how can AI contribute to health in resource-poor settings?” BMJ Global Health 3:e000798 (framework for equitable AI)
  • Schwalbe & Wahl (2020). “Artificial intelligence and the problem of knowledge collapse in global health.” The Lancet Global Health 8:e1444-e1445 (critique of data extraction)
  • Beede et al. (2020). “A human-centered evaluation of a deep learning system in diabetic retinopathy screening.” PACM HCI 4:1-30 (Google Health Thailand deployment study; illustrates how lab performance does not transfer to real-world clinical settings)
  • Tomson et al. (2022). “Global health equity in AI: a framework for algorithmic fairness in low-resource settings.” BMJ Global Health 7:e008822 (policy recommendations)
  • WHO (2021). “Ethics and governance of artificial intelligence for health.” (global governance framework)

Part 1: Hypothetical Thailand Malaria-Screening Failure

Taught us: Validation in Western populations does not guarantee LMIC performance. Infrastructure assumptions matter. Community trust is fragile.

Case Study Note

This is a constructed hypothetical, not an anonymized report of an actual Thai Ministry of Public Health pilot. The startup, partnership, dataset, publication, devices, clinics, performance values, costs, patient outcomes, community quotation, discontinuation decision, and later corporate pivot are fictional. The details illustrate transportability, infrastructure, acquisition, workflow, and monitoring questions. They must not be cited as measured evidence about Thailand, malaria AI, a company, or patient harm.

The Promise (2018-2019)

Background: Thailand-Myanmar border region endemic for malaria (primarily P. falciparum, P. vivax). Traditional diagnosis requires microscopy, but trained microscopists are scarce in remote border clinics serving migrant workers and refugees.

Technology: U.S.-based startup (anonymized here) developed smartphone microscopy AI for malaria detection: - Clip-on smartphone microscope lens ($50) - Blood smear imaging via smartphone camera - Cloud-based AI analyzes images, detects parasites - Results in 2-3 minutes (vs. 30-60 min for traditional microscopy)

Fictional laboratory validation claimed in the exercise: - Training dataset: 12,500 blood smears from CDC reference lab (Atlanta, U.S.) - Validation: 2,100 smears from U.S. hospital (imported malaria cases, mostly travelers) - Performance: Sensitivity 98.2%, Specificity 97.5% - Comparison: Expert microscopist sensitivity 95%, specificity 98%

The pitch: “AI outperforms human microscopists, works on $200 smartphones, democratizes malaria diagnosis for resource-limited settings.”

Deployment plan: Partnership with Thai Ministry of Public Health to deploy in 15 border clinics across 3 provinces (Tak, Kanchanaburi, Ranong). Target: Screen 50,000 patients over 18 months.

The Reality (2019-2020 Pilot)

Deployment challenges:

1. Dataset Mismatch - Training data: CDC reference lab smears (thin smears, optimal staining, high parasite density, primarily P. falciparum) - Field reality: Border clinic smears (thick smears for higher sensitivity, variable staining quality, low parasite density common, mixed P. falciparum + P. vivax infections) - Result: AI sensitivity degraded from 98% to 67% on field samples. 31 percentage point degradation

Specific failure modes: - Low-density infections (<100 parasites/μL): AI sensitivity 52% (expert microscopist 85%) - P. vivax detection: AI sensitivity 61% (trained primarily on P. falciparum) - Mixed infections: AI detected only 1 species in 73% of mixed cases - Poorly stained slides: AI rejected 18% as “insufficient quality” (vs. 3% rejection by expert)

2. Infrastructure Failures - AI design: Cloud-based processing (images uploaded, AI runs on remote servers, results returned) - Border clinic reality: Intermittent 3G connectivity (during monsoon season, clinics offline for days) - Wait times: When connectivity available, upload + processing: 8-15 minutes per patient (vs. promised 2-3 minutes) - Workflow breakdown: Clinics seeing 40-60 patients/day, AI could handle 10-15 patients/day during good connectivity

3. Smartphone Hardware Issues - Deployment devices: Mid-range Android phones ($200, as promised) - Camera quality: Variable image quality (12 MP cameras, older models struggled with low-light conditions in clinics without consistent electricity) - Battery life: Continuous use (imaging, uploading) drained batteries in 3-4 hours; clinics often lacked charging infrastructure or experienced power outages - Device failure: 40% of smartphones malfunctioned within 6 months (humid tropical environment, no protective cases, frequent drops)

4. Clinical Impact - False negatives: 33% of malaria cases missed by AI (using AI alone, not AI + microscopist backup) - Patients with missed diagnoses: Untreated malaria → severe disease, hospitalizations - Illustrative adverse outcomes: The scenario assumes progression to severe malaria after false-negative results. These events did not occur in a documented pilot. - Illustrative community response: The quoted reaction is fictional and represents a possible loss of trust.

Illustrative Scenario Numbers

Every value in this table is fictional. Its purpose is to show which laboratory, field, workflow, device, and equity measures a prospective pilot should compare.

Metric Lab Validation (U.S.) Field Performance (Thailand) Delta
AI sensitivity 98.2% 67% -31 percentage points
AI specificity 97.5% 89% -8.5 percentage points
P. vivax detection 97% 61% -36 percentage points
Low-density infection detection 96% 52% -44 percentage points
Processing time per patient 2-3 min 8-15 min (when online) 3-5x slower
Image rejection rate <3% 18% 6x higher
Device failure rate (6 months) <5% (assumed) 40% 8x higher

Project outcomes: - Target: 50,000 patients screened over 18 months - Actual: ~3,200 patients screened before pilot suspended (12 months) - Only 6% of target achieved

Illustrative economic assumptions: - Estimated pilot cost: $450,000 (devices, training, cloud infrastructure, monitoring) - Cost per successful diagnosis: $140 (vs. <$2 for traditional microscopy) - Thai Ministry of Public Health discontinued project after 12 months

The Lesson for Physicians

Why this failure matters:

1. Validation must match deployment population - AI trained on U.S. travelers failed on endemic malaria (mostly travelers returning from Africa, high P. falciparum density) failed on Southeast Asian endemic malaria (low density, P. vivax predominant) - Red flag: No validation in Thailand/Myanmar border populations before deployment

2. Infrastructure assumptions invisible in lab validation - Cloud-based AI assumed reliable internet (absent in 70% of target clinics) - Smartphone AI assumed HIC-level devices (mid-range phones inadequate for consistent imaging) - Red flag: No field testing of connectivity, device durability before large-scale deployment

3. Clinical context matters - U.S. imported malaria: High pretest probability (symptomatic travelers), thick smears, expert microscopy backup - Thai border clinics: Variable pretest probability (screening migrant workers), resource constraints, AI as replacement for microscopy (not adjunct) - Consequence: False negatives caused preventable severe disease, deaths

4. Stakeholder engagement critical - Technology developed in U.S., “deployed” in Thailand without co-design with local clinicians, microscopists - Community health workers not trained adequately, did not trust AI outputs - Result: Low adoption even when AI available

What should have been done differently:

Local validation before deployment: Pilot on 1,000+ Thai border region blood smears, measure performance on local malaria species, staining protocols Infrastructure assessment: Survey clinic connectivity, electricity, device charging before selecting cloud vs. edge AI Co-design with local stakeholders: Partner with Thai researchers, microscopists, CHWs from project inception Hybrid approach: AI assists microscopists (not replaces), human review of all AI-negative results in high-risk populations Ruggedization: Weatherproof device cases, solar chargers, offline-capable edge AI Phased deployment: Start with 1-2 clinics (silent mode → shadow mode → active mode), validate in real-world conditions before scaling

Real-world boundary: This fictional startup has no current status, and the chapter does not assert that Thailand adopted an AI-specific national requirement because of such a pilot. Real procurement should verify the current Thai regulatory pathway, device record, local evidence requirements, malaria diagnostic guidance, and institutional monitoring plan from primary sources.


Part 2: MomConnect as a National-Scale Digital-Health Comparator

What it teaches: A nationally owned, low-bandwidth digital-health service can reach scale when it is integrated into public care. MomConnect is not evidence that an AI chatbot caused reductions in maternal mortality, preterm birth, low birth weight, or cost. Its value here is as a non-AI or mixed digital-service comparator for delivery, ownership, language, and helpdesk design.

The Problem

South African maternal-health context: MomConnect launched in 2014 as a National Department of Health initiative to provide pregnancy and postpartum information and a helpdesk through mobile technology. The program addressed maternal and child health within the public system; it should not be reduced to a single historical mortality ratio or compared across countries without harmonized definitions.

Information and service gaps: Pregnant and postpartum users need timely, understandable guidance, appointment information, a way to ask questions, and a route for complaints or urgent escalation. Language, literacy, phone access, privacy, and continuity after delivery shape who benefits.

Technology choice: SMS reduced dependence on smartphones and mobile data. The current service also uses WhatsApp, but channel, language, cost, privacy, and accessibility differ. The design lesson is to choose the lowest-burden channel that can support the required clinical and communication function.

The Technology

MomConnect, led by the South African National Department of Health with implementation partners:

Design principles: - SMS-based, no smartphone required (works on 2G networks) - Free to users (government subsidizes SMS costs) - Opt-in (pregnant women register at first prenatal visit or via SMS short-code) - Automated chatbot and human-operated text helpdesk for questions and feedback; these functions should not be conflated with autonomous clinical diagnosis - SMS content in South Africa’s 11 official languages, while the current official technical page describes WhatsApp messages in English (South African National Department of Health)

How it works:

Phase 1: Registration and Profile - A pregnant user registers through the current supported channel or is assisted during care; codes and enrollment workflows can change and should be checked on the official program page - Chatbot asks: Due date? First pregnancy? HIV status known? Language preference? - Woman receives welcome message, information on nearest clinic

Phase 2: Weekly Health Messages - AI sends stage-appropriate health messages (nutrition, HIV testing, danger signs, childbirth preparation) - Messages tailored to gestational age (e.g., Week 20: “Your baby is growing. Remember to take your iron tablets daily. Next visit: [date]”)

Phase 3: Two-Way Communication - Woman can text questions: “I have headache and swelling” → AI assesses symptoms - Automated routing plus human helpdesk: - Urgent language should trigger conservative escalation under an evaluated protocol - Routine questions can receive approved information - Ambiguous or complex questions require human helpdesk or clinical review

Phase 4: Appointment Reminders - SMS reminders 2 days before prenatal visits, postnatal visits, infant immunizations - Whether reminders improve attendance must be evaluated rather than assumed

Phase 5: Feedback Loop - Women can report clinic experiences (long wait times, stock-outs, rude staff) - Data aggregated, sent to health system managers for quality improvement

The Evidence

National reach: The current National Department of Health page reports that almost 5 million mothers using public antenatal services have registered since 2014 across more than 95% of public health facilities (South African National Department of Health). This is an official cumulative program measure, not a count of active users, message exposure, or clinical benefit.

Early peer-reviewed delivery evidence: A descriptive analysis of system data from August 2014 through April 2017 reported 1,159,431 registrations. It also identified substantial attrition in the delivery chain: 26% of registration attempts in 2016 did not convert, fewer than one quarter of mobile health messages were received by pregnant users under the study’s exposure definition, and fewer than 6% of enrolled users were exposed to at least one message at 6–12 months postpartum (LeFevre et al., 2018). Enrollment, successful delivery, exposure, behavior, service use, and health outcomes are different endpoints.

Program-reported recent indicators: The official ten-year page reports survey and external-evaluation results including high antenatal-visit attendance, vaccination-card verification, and subgroup effects on breastfeeding knowledge or behavior and family-planning measures. These results should be read with their study designs, comparators, p values, and reports, not converted into a universal maternal-mortality effect (South African National Department of Health).

Evidence boundary: Peer-reviewed accounts credit government leadership, national integration, partners, and health-worker and user engagement for scale, while noting that summative evaluation could not cleanly link program exposure to maternal and child health outcomes (Barron et al., 2018). The earlier claims of a 17% mortality reduction, 18-percentage-point postnatal effect, 15:1 return, and $42 per disability-adjusted life-year were not supported by the sources reviewed and are not retained as measured results.

Why MomConnect Reached National Scale

Factor MomConnect Design Typical AI Pilot
Technology level SMS (works on basic phones, 2G) Smartphone app requiring 4G
Infrastructure assumptions Minimal (SMS works everywhere) Reliable internet, smartphones
Problem prioritization Locally identified (South African maternal mortality) Externally imposed (what’s interesting to researchers)
Stakeholder engagement Co-designed with SA Dept of Health, nurses, pregnant women Developed abroad, “deployed” locally
Sustainability model Government-funded, integrated into health system Donor-funded pilot, no long-term plan
Language/cultural fit 11 languages, culturally appropriate messages English-only or poor translations
Automation complexity Approved messaging, automated routing, and human helpdesk Complex clinical generation without an escalation pathway
Failure mode Graceful degradation (SMS delivery fails, retry) System crash, no offline mode

Scale and Sustainability

Related programs beyond South Africa: Other countries have implemented maternal messaging services, including Kilkari in India. These are separate programs with their own designs and evidence. They should not be counted as MomConnect expansion or used to create a combined user total without a direct program source.

Integration with health systems: - South Africa: MomConnect integrated with national health information system (appointment scheduling, immunization tracking, chronic disease management) - Data used for health system quality improvement (identifying clinics with high no-show rates, medication stock-outs)

Challenges remaining:

  1. Digital divide within countries: Phone ownership, control of a shared device, affordability, literacy, disability, and connectivity can exclude intended users
  2. Literacy: SMS requires reading ability (audio versions in development)
  3. Male partner engagement: Messages target pregnant women, miss fathers/male partners
  4. Misinformation: Some users receive conflicting advice from traditional healers, family
  5. Data privacy: Concerns about government access to sensitive health data (HIV status)

The Lesson for Physicians

Why MomConnect worked when high-tech AI pilots failed:

1. Appropriate technology for context - SMS, not smartphone apps, met users where they are - 2G network requirement (vs. 4G) ensured rural coverage

2. Locally prioritized problem - Maternal mortality was South African government’s top health priority - Solution co-designed with local stakeholders, not imposed externally

3. Bounded automation plus human escalation - Approved messages and routing are easier to constrain than autonomous diagnosis - SMS lowers device and bandwidth requirements but still depends on registration, delivery, phone access, and privacy - The helpdesk provides a route for complex questions and complaints

4. Sustainable business model - Government-funded from start, not pilot-dependent on external donors - Integrated into existing health system, not parallel vertical program

5. Equity-focused design - Free to users (no cost barrier) - Works on cheapest phones (no device barrier) - 11 languages (no language barrier) - Low-literacy accommodations (simple language, audio versions planned)

Questions for evaluating global health AI:

Does technology match infrastructure reality? (MomConnect: SMS works on 2G with basic phones) Was problem identified by local stakeholders? (MomConnect: SA govt priority) Is AI appropriate complexity? (MomConnect: Simple rules, not brittle LLMs) Is there sustainable funding? (MomConnect: Government-funded, integrated into health budget) Does it reduce or widen digital divide? (MomConnect: Reduces, accessible to poorest populations)


Part 3: Data Colonialism and Equitable Partnerships

Data colonialism: Extraction of LMIC health data without benefit-sharing by HIC institutions/companies for commercial AI development.

How Data Extraction Happens

Typical scenario:

  1. HIC tech company/research institution approaches LMIC hospital: “We’ll build AI for [disease], need your patient data for training”
  2. Data transfer: LMIC hospital provides de-identified patient records, imaging, genomics (often millions of records)
  3. AI development: HIC institution trains model, publishes research, files patents, commercializes product
  4. Deployment: AI sold back to LMIC hospitals at commercial rates (or only deployed in HIC markets)
  5. Benefit to data source: Zero (or minimal: co-authorship on 1-2 papers)

Hypothetical composite examples:

The following two cases are constructed negotiation exercises, not anonymized accounts of identifiable universities, hospitals, publications, companies, licensing fees, or grants.

Case 1: Tuberculosis chest X-ray AI - U.S. university partnered with 8 sub-Saharan African hospitals - Collected 250,000 chest X-rays + TB diagnoses - Developed AI, published in Nature Medicine, licensed to commercial vendor - African hospitals received: Co-authorship on paper - African hospitals did NOT receive: Access to AI tool (vendor charged $10K+ licensing fee), revenue share, capacity-building

Case 2: Cervical cancer screening - European consortium collected cervical images from 12 LMIC sites (Latin America, Africa, Asia) - Trained AI for HPV lesion detection - Secured €15M commercial funding for product development - LMIC sites received: Acknowledgment in paper footnotes - LMIC sites did NOT receive: Equity stake, free product access, training in AI development

Ethical Issues

1. Consent and lawful-use risks - Permission for clinical care does not automatically establish a lawful basis or ethically sufficient permission for secondary commercial AI development - Requirements depend on the jurisdiction, data, identifiability, protocol, ethics review, notices, agreements, and applicable research or privacy law - Analogy: Imagine donating blood for research, later discovering pharmaceutical company sold your cells for $1B (Henrietta Lacks case)

2. Benefit-Sharing Failure - Nagoya Protocol (biodiversity) established equitable benefit-sharing for genetic resources - Digital health data should follow same principles but currently does not - LMIC populations provide data, HIC institutions capture value

3. Capacity Extraction vs. Building - Data extraction without training local researchers → LMICs remain dependent on foreign expertise - Contrast: Capacity-building partnerships train local AI researchers, leave sustainable infrastructure

4. Exploitation in the AI Supply Chain

WHO’s 2025 guidance addresses lifecycle ethics and governance for large multimodal models (WHO, 2025). Labor conditions are a related procurement and research-ethics concern because data annotation, transcription, and content review can be outsourced across borders.

These workers frequently face:

  • Inadequate compensation: Piece-rate work and weak bargaining power can produce low or unpredictable pay; contracts and audits should establish the actual compensation and conditions
  • Psychological distress: Annotating traumatic medical images (injuries, pathology specimens) without mental health support
  • Poor working conditions: Long hours, repetitive tasks, minimal job security
  • No benefit-sharing: Workers contribute to billion-dollar AI products but receive only piece-rate payment

Equitable AI development must address not only patient data rights but also fair labor practices throughout the supply chain. Physicians evaluating AI vendors should ask: How was your training data annotated? What are the working conditions and compensation for data workers?

5. Data Sovereignty - Who owns patient data? Individuals? Hospitals? Governments? - Legal and institutional frameworks vary substantially and continue to change - A weak contract or unclear governance can permit extraction even where privacy or research rules exist

Solutions and Frameworks

1. Equitable Data Partnerships

Model: INDEPTH Network (International Network for the Demographic Evaluation of Populations and Their Health) - Network of health and demographic surveillance sites across Africa and Asia; current membership and governance should be verified from the network - Shared data governance principles: - Data hosted in country of origin (not extracted to HIC servers) - Local researchers lead analysis - External collaborators require approval from local ethics committees - Publications require local co-authorship (not just acknowledgment) - Commercial use requires benefit-sharing agreements

2. Benefit-Sharing Agreements

Essential elements: - Free access: AI tools developed from LMIC data available to source institutions at no cost - Commercial benefit: If commercialized, agreements can consider royalties, equity, milestones, subsidized access, or reinvestment, with terms negotiated rather than assumed from a standard percentage - Capacity-building: HIC partners train local researchers in AI development - Co-ownership: Shared intellectual property rights

Example: H3Africa (Human Heredity and Health in Africa) genomic research consortium - African researchers co-lead studies - Data access, governance, capacity development, and institutional leadership are addressed through consortium policies and project-specific agreements; storage and benefit terms should not be generalized from one sentence

Capacity-Building Frameworks for Sustainable AI

Many donor-supported pilots struggle to transition into durable local services, but no verified universal 80% failure rate or two-year cutoff applies. Sustainability must be designed and measured through ownership, financing, maintenance, workforce, supply chains, governance, and an exit plan.

RAD-AID Three-Pronged Strategy

RAD-AID International developed a framework specifically addressing the gap between AI promise and sustainable implementation in low-resource radiology settings (Mollura et al., 2020):

  1. Education: Train local radiologists, technologists, and referring physicians in AI-augmented interpretation. Build understanding of AI capabilities, limitations, and appropriate use cases before deployment.

  2. Infrastructure: Assess and address gaps in electricity, connectivity, PACS/RIS systems, and device maintenance capacity. AI cannot function where foundational infrastructure is absent.

  3. Phased AI integration: Deploy AI in progressive stages: silent mode (AI runs, outputs not shown) → shadow mode (AI outputs shown after human interpretation) → active mode (AI outputs shown before human interpretation). Each phase validates performance and builds user trust before increasing AI influence on clinical decisions.

Evidence boundary: The cited RAD-AID article describes a framework and implementation priorities; it does not support the earlier Guyana and Nigeria concordance table as a comparative clinical evaluation. Each site needs its own documented design, endpoint, and source.

Workforce Empowerment

The 2025 Johns Hopkins consensus workshop emphasized that AI should be implemented with, not on, LMIC workforces (Marey et al., 2025). Key principles:

  • Local ownership: LMIC clinicians and researchers should lead implementation, not serve as data sources for foreign projects
  • Skills transfer: Every AI deployment should include training that enables local teams to maintain, troubleshoot, and eventually improve systems
  • Career pathways: Create roles for “AI specialists” within LMIC health systems, with compensation and advancement opportunities

Technology Transfer Models:

Model Description Sustainability Examples
Hub-and-spoke Central academic center provides AI expertise to peripheral sites Depends on hub continuity, authority, and local transfer AMPATH Kenya, Partners in Health
Federated networks Multiple LMIC sites collaborate while limiting central data transfer Depends on compute, standards, governance, and shared authority INDEPTH Network, H3Africa
Open-source commons Software is available under an open license Avoids some license lock-in but still needs maintainers, infrastructure, support, and governance OpenMRS, DHIS2
South-South partnerships LMIC institutions collaborate directly Depends on funding, reciprocity, capacity, and institutional commitments Africa CDC and regional research networks

Lessons from Failures:

Common failure patterns when capacity-building is neglected:

  1. Vendor lock-in: Proprietary systems that LMIC teams cannot maintain after external support ends
  2. Brain drain: Local staff trained in AI leave for HIC opportunities or private sector
  3. Orphan technology: Devices without local repair capacity become e-waste within 2-3 years
  4. Mission drift: Projects pivot to HIC markets when LMIC revenue proves insufficient

Sustainable design requirements:

  • Open-source or open-standard technology (avoids vendor lock-in)
  • Local repair and maintenance capacity (trained technicians, spare parts supply chain)
  • Retention incentives for trained staff (competitive salaries, career growth)
  • Government integration (budget line items, not parallel donor-funded systems)

3. Data Sovereignty Regulations

India Digital Personal Data Protection Act, 2023: - The enacted law is the Digital Personal Data Protection Act, not the earlier bill - It does not create the blanket health-data localization rule stated in the prior version of this chapter - Cross-border restrictions, consent and notice, significant-data-fiduciary duties, exemptions, rules, and commencement status must be checked against the current official text and implementing measures (India Code) - The schedule contains obligation-specific monetary ceilings; it should not be summarized as a single ₹15 crore or 4% export penalty

African Union Data Policy Framework (2022): - Establishes continental data governance principles - Emphasizes data sovereignty, local value capture, capacity-building - Member states developing national data protection laws

4. Open Science Models

Global Alliance for Genomics and Health (GA4GH): - Open-source tools for federated data analysis (data stays in country, AI travels to data) - International standards for responsible data sharing - Emphasis on public good, not commercial extraction

Physician Responsibilities

When approached for LMIC data partnerships:

Red flags to reject: - No benefit-sharing beyond co-authorship - Data transfer to HIC without local storage - No capacity-building commitment - Commercial use without LMIC equity stake - Short-term extractive relationship

Green flags to support: - Data hosted locally (or federated learning without transfer) - Shared IP ownership - Free access to resulting AI for source institutions - Multi-year capacity-building commitment (training local researchers) - LMIC researchers in leadership roles (not just acknowledged)


Part 4: Algorithmic Bias and Global Health Equity

AI trained on non-representative data performs poorly on underrepresented populations, worsening health disparities.

Mechanisms of Bias

1. Training Data Bias - Most medical AI trained on HIC populations (predominantly white, North American/European) - Underrepresentation of LMIC populations, racial/ethnic minorities

2. Label Bias - Disease definitions, diagnostic criteria differ across populations - Example: Heart failure diagnostic thresholds optimized for Western populations may miss disease in Asian populations with different body size distributions

3. Measurement Bias - Medical devices calibrated for specific populations - Example: Pulse oximeters overestimate oxygen saturation in dark-skinned patients → AI using pulse ox data inherits this bias

4. Prevalence Bias - Disease prevalence differs dramatically between HIC and LMIC - Example: TB AI trained on U.S. data (TB prevalence <10 per 100,000) miscalibrated for India (TB prevalence 200+ per 100,000)

Measured and Illustrative Bias Examples

Illustrative Example 1: Skin Cancer Detection AI

The algorithm, sample size, table values, and mortality inference below are synthetic teaching data. They demonstrate how to inspect sensitivity, specificity, and representation by skin tone; they are not results from a cited melanoma model.

Algorithm: Deep learning model for melanoma detection, trained on 130,000 dermatology images

Performance by skin tone (Fitzpatrick scale):

Skin Tone Sensitivity Specificity Training Data %
Type I-II (light) 91% 89% 78%
Type III-IV (medium) 83% 84% 18%
Type V-VI (dark) 65% 76% 4%

Illustrative disparity: 26 percentage points lower sensitivity for dark skin (Type V–VI versus I–II). A real analysis would need uncertainty, case mix, reference-standard quality, spectrum, access, stage, and outcome data before attributing clinical harm.

Root cause: Training dataset 78% light skin tones, only 4% dark skin tones (despite dark skin being majority globally)

Illustrative Example 2: Sepsis Prediction Models

The subgroup table below is synthetic and should not be attributed to the Epic Sepsis Model or the 38,455-encounter external validation. The actual external validation reported poor overall sensitivity and positive predictive value and did not establish patient-outcome benefit (Wong et al., 2021).

Algorithm: Sepsis early warning AI (Epic Sepsis Model), trained on 405,000 U.S. hospital encounters

Performance by race/ethnicity (external validation, n=38,455):

Patient Group Sensitivity Specificity PPV
White 63% 95% 18%
Black 51% 96% 12%
Hispanic 48% 94% 10%
Asian 44% 97% 11%

Illustrative disparity: 19 percentage points lower sensitivity for Asian versus white patients. The synthetic table does not establish delayed treatment or higher mortality.

Root cause: Training data 67% white patients, only 6% Asian. Model learned disease patterns from white patients and missed Asian-specific presentations

Example 3: Pulse Oximetry Bias

Device: Pulse oximeters measure oxygen saturation (SpO₂) Problem: Overestimate SpO₂ in dark-skinned patients (melanin interferes with light absorption)

Hidden hypoxemia (arterial O₂ <88% despite pulse ox reading of 92-96%): - Black patients: 11.7% hidden hypoxemia - White patients: 3.6% rate of hidden hypoxemia - 3.2x higher risk in Black patients (Sjoding et al., 2020)

AI implications: - COVID-19 AI models using pulse ox data → biased predictions - Sepsis AI using pulse ox → underestimates severity in Black patients - Cascade effect: Biased device leads to biased training data, which creates biased AI and perpetuates disparities

Solutions to Algorithmic Bias

1. Diversify Training Datasets

Strategy: - Actively collect data from underrepresented populations - Oversample minority groups to balance representation - Multi-site training including LMIC hospitals

Illustrative dataset-redesign exercise: A development team could prospectively recruit sites and populations that address representation and device gaps, then compare subgroup performance with uncertainty. The previously stated “CheXpert Version 2” demographics and 12-percentage-point improvement were not supported by a source and should not be treated as an actual dataset release.

2. Fairness-Aware Machine Learning

Techniques: - Equalized odds: Constrain AI to achieve equal sensitivity/specificity across demographic groups - Demographic parity: Equal positive prediction rates across groups - Calibration: Predicted probabilities match actual outcomes for all groups

Trade-offs: - Perfect fairness across all metrics impossible (mathematical constraints) - Fairness constraints can change overall and subgroup performance in different directions; the trade-off must be measured for the chosen objective and operating point - Choose metrics and thresholds through an explicit account of who bears false positives, false negatives, delay, and exclusion

3. External Validation in Deployment Populations

Requirement: Validate AI in populations where it will be deployed, BEFORE deployment

WHO governance boundary: WHO’s 2021 guidance provides six principles and governance recommendations; it does not establish the previously stated universal three-site noninferiority rule or 10% subgroup threshold (WHO, 2021). Institutions should prespecify sites, subgroups, metrics, uncertainty, and action thresholds from intended use and consequences.

4. Bias Auditing and Monitoring

Continuous monitoring: - Track AI performance by demographic subgroups post-deployment - Alert when a prespecified clinically meaningful performance or access boundary is crossed - Investigate before retraining; recalibration, workflow change, use restriction, rollback, or retirement may be more appropriate

Regulatory mandates: - FDA submissions and quality systems require evidence appropriate to the device and intended use; there is no single public rule that every AI device must pass one universal bias test - The EU AI Act imposes risk-management, data-governance, documentation, monitoring, and other obligations for high-risk systems; publication and applicability depend on the provision, actor, and implementation timeline

LMIC Regulatory Pathways for AI Medical Devices

AI medical devices face fragmented regulatory landscapes across low- and middle-income countries, creating barriers to deployment even when technology is validated and effective. Unlike the FDA’s centralized system, most LMICs lack dedicated AI/ML regulatory frameworks, and medicolegal mechanisms for AI-based tools remain ambiguous or absent in many jurisdictions (Marey et al., 2025).

Regional Regulatory Bodies:

Country Named authority to verify Current questions for the exact product
South Africa SAHPRA Is the function a regulated medical device, what local authorization is required, and what postmarket duties apply?
Nigeria NAFDAC What classification and registration pathway applies, and is any foreign authorization relevant but insufficient?
Kenya Pharmacy and Poisons Board and other competent health authorities as applicable Which authority governs the software function, data flow, clinical study, and facility use?
India CDSCO How is the software classified under current medical-device rules, and what local evidence and registration are required?
Bangladesh Directorate General of Drug Administration What current device, software, import, and facility rules apply?
Pakistan Drug Regulatory Authority of Pakistan What current classification, licensing, vigilance, and clinical-evidence requirements apply?
Indonesia Ministry of Health and competent device authorities What pathway applies to the software function, local representative, data processing, and facility deployment?
Thailand Thai FDA What current classification, registration, cybersecurity, change, and postmarket requirements apply?

This table is a routing aid, not a claim that a country has no framework, automatically accepts foreign authorization, or assigns one review duration. Each answer must be verified from the current primary national source before procurement or research.

Key Regulatory Challenges:

  1. Regulatory capacity and fit: Review capacity, classification, evidence, and timelines differ across jurisdictions and products; a U.S. review duration cannot be used as a benchmark for another regulator.

  2. Reliance and recognition pathways: Some jurisdictions may consider foreign authorization or certification, but local legal requirements and population, device, protocol, and workflow fit still need verification.

  3. Post-market surveillance gaps: Even where pre-market review exists, systematic post-market monitoring is rare. Performance degradation after deployment often goes undetected.

  4. Medicolegal ambiguity: When AI contributes to adverse outcomes, liability frameworks are unclear. Who is responsible: the vendor, the deploying institution, or the clinician? Many jurisdictions have not addressed this question.

Alternative Pathways:

WHO prequalification is product-category specific and is not a general substitute pathway for every AI medical device. Procurement teams should verify whether a product category is eligible, what WHO recommendation or listing exists, and what national authorization remains required.

Implications for Deployment:

Physicians deploying AI in LMICs should:

  • Document regulatory status: Note whether device has local approval, foreign approval only, or no formal approval
  • Obtain institutional ethics approval: When regulatory pathways unclear, ethics committee review provides governance layer
  • Establish accountability protocols: Written roles, escalation authority, audit access, insurance, indemnity, and incident duties, recognizing that a contract does not decide all statutory or tort liability
  • Report adverse events: Even without formal surveillance systems, document and report AI-related adverse outcomes to build evidence base

Part 5: Infrastructure and the Digital Divide

AI deployment assumes infrastructure that is often absent in LMIC settings.

Infrastructure Realities

Electricity: - The 2026 Tracking SDG 7 report estimated that 655 million people remained without electricity and emphasized that progress in sub-Saharan Africa had slowed (World Bank, 2026) - A national access percentage does not establish power quality, backup capacity, voltage stability, or uptime at a particular health facility

Internet connectivity: - ITU estimated that 2.2 billion people remained offline in 2025, despite mobile broadband coverage exceeding 96% globally (ITU, 2025) - Coverage is not meaningful access. Affordability, device ownership, speed, reliability, data caps, latency, and digital skills determine whether a clinical workflow functions

Devices: - Ownership, control of a shared device, operating-system support, camera and sensor quality, storage, repair, charging, replacement, and data cost vary within countries - Procurement must test the actual lowest-supported device and the full acquisition workflow rather than assuming that network coverage implies a compatible phone

Digital literacy and accessibility: - Reading ability, language, disability, trust, prior digital experience, and availability of assisted access affect use - Voice, icon, SMS, USSD, human helpdesk, and community-health-worker pathways should be evaluated with intended users rather than treated as universal solutions

Design Principles for Low-Resource Settings

1. Offline-First Design (Edge AI)

Rationale: Cloud AI requires internet, edge AI runs on-device

Product verification questions: - Does the exact authorized version run fully on the acquisition device, a local server, or a vendor cloud? - Which functions remain available offline, how long can data queue, and what happens after synchronization failure? - What hardware, operating system, calibration, cybersecurity, and update dependencies apply?

Trade-offs: - Edge AI requires more powerful devices (higher cost) - Model updates harder (vs. cloud models updated centrally) - Local and cloud implementations may differ in model, compression, latency, privacy, and update process. The performance difference must be measured for the exact versions.

2. Low-Power, Solar-Compatible

Design features: - Energy-efficient algorithms (reduce computational load) - Solar charging capability - Battery and charging requirements matched to the intended shift, acquisition volume, and outage pattern

Implementation example: Community-health-worker platforms such as CommCare illustrate offline data collection and synchronization patterns. Current device price, battery life, deployment geography, and solar-charging support should be verified for the specific implementation rather than treated as product-wide constants.

3. Robust to Poor Data Quality

LMIC data challenges: - Lower-resolution imaging (older equipment) - Incomplete EHR data (paper records common, partial digitization) - Variable data quality (inconsistent protocols across sites)

AI robustness techniques: - Transfer learning: Pre-train on HIC data, fine-tune on limited LMIC data - Data augmentation: Simulate poor quality (blur, noise) during training - Uncertainty quantification: AI flags low-confidence predictions for human review

4. Simplicity and Usability

User-centered design: - Training burden measured through task completion, errors, retention, escalation, and support needs - Intuitive interfaces (icons, minimal text for low-literacy users) - Voice-based interaction for illiterate users - Local language support

Example: Medic Mobile (CHW app for maternal/child health): - Icon-based navigation (pictures, not text-heavy) - SMS workflows (for feature phones) - Audio prompts in local languages - Deployment evidence, language support, literacy requirements, and current geographic reach should be verified from the implementing program

Bridging the Digital Divide

Infrastructure Investments: - Expand electricity access (grid extension, mini-grids, solar home systems) - Subsidize internet connectivity for health facilities (government programs, partnerships with telecom companies)

Device Affordability: - Low-cost smartphones ($50-100 range) designed for developing markets - Shared devices for community health workers (government-funded)

Digital Literacy Programs: - Training for patients, CHWs in basic digital skills - Integration into primary/secondary education curricula

Timeline boundary: SDG and connectivity targets express policy goals, not reliable forecasts of when a particular facility will have adequate electricity or internet. Design for measured current conditions and define the infrastructure change required for expansion.

Updated Barriers Framework (2025)

The 2025 Johns Hopkins workshop on AI in global health radiology identified five interlocking barriers that explain why AI often fails to deliver promised benefits in LMIC settings (Marey et al., 2025):

  1. Infrastructure: Unreliable electricity, limited internet connectivity, inadequate imaging equipment, and absence of digital health records

  2. Data: Scarcity of labeled LMIC datasets, poor data quality, lack of standardization across sites, and data sovereignty concerns

  3. Workforce: Shortage of radiologists and AI-literate health professionals, limited training opportunities, and brain drain to HICs

  4. Regulatory: Absence of AI-specific frameworks, medicolegal ambiguity, and reliance on inappropriate foreign standards

  5. Financing: Dependence on short-term donor funding, lack of sustainable business models, and inability to demonstrate ROI to health ministries

Critical Insight:

These barriers are interlocking, not independent. Addressing infrastructure without workforce development leaves systems unmaintained. Building workforce capacity without sustainable financing leads to brain drain. Regulatory clarity without data standards creates compliance theater.

The workshop consensus emphasized that technology is not value-neutral: AI can either reinforce existing inequities or help overcome them, depending on how implementation addresses all five barriers simultaneously.

Implication for Physicians:

When evaluating AI for LMIC deployment, assess all five dimensions. A technically excellent AI system will fail if deployed into a context where three of five barriers remain unaddressed. Success requires coordinated intervention across infrastructure, data, workforce, regulatory, and financing domains.

User Sentiment Paradox: Optimism Despite Barriers

Anthropic reported themes from 80,508 voluntary interviews with users across 159 countries, including regional differences in optimism and stated aspirations (Anthropic, 2026). This is a vendor-run sample of people already using an AI service, not a representative population survey and not evidence of clinical acceptance. It can generate questions about regional aspirations, but it cannot establish that communities facing the largest healthcare gaps are uniformly more receptive to medical AI.


Part 6: Large Language Models in Global Health

Large language models (LLMs) represent a distinct paradigm from the narrow diagnostic AI covered in previous sections. Unlike task-specific systems (radiology AI detecting pneumonia, pathology AI grading cancer), LLMs are general-purpose language systems that can assist clinical documentation, answer medical questions, synthesize literature, and support clinical reasoning through natural language interaction.

This versatility makes LLMs particularly relevant for global health: they can potentially address multiple healthcare gaps simultaneously, augment overburdened workforces, and scale to diverse settings without task-specific retraining. However, this same flexibility creates unique risks, hallucinations (confident but false information), privacy concerns, and deployment challenges that differ from narrow AI systems.

Open-Weight Models: Democratizing Access or Premature Deployment?

The computational efficiency breakthrough:

Open-weight LLMs such as DeepSeek, Llama, and Mistral can permit local hosting and adaptation, but “open-weight” does not guarantee low computational cost, open training data, clinical validity, security, or maintainability. Comparative performance and cost depend on model size, quantization, hardware, task, serving volume, and evaluation design (Ong et al., 2026).

DeepSeek in Chinese hospitals:

Chen and colleagues described reported adoption of DeepSeek across Chinese hospitals and potential applications in decision support, communication, and administration (Chen et al., 2025). Hospital counts in a narrative or news-derived account do not establish active clinical use, patient benefit, or regulatory status. In one ophthalmology benchmark of 300 cases across 10 subspecialties, an estimated model-query cost was 6.71% of a comparator. That result is task- and pricing-specific, not the total cost of ownership or proof of equivalent clinical care.

The “too fast, too soon” warning:

Chinese medical researchers have raised substantial concerns about rapid deployment without adequate clinical validation. A JAMA research perspective led by Zeng and Wong warns that DeepSeek’s tendency to generate “plausible but factually incorrect outputs” could lead to “substantial clinical risk” when deployed at scale without rigorous prospective validation (Zeng et al., JAMA, 2025).

The dual-edged reality:

Advantage Risk
Potential serving-cost control: Local operators can choose model and infrastructure Safety monitoring gaps: Guardrails, evaluation, and monitoring vary across both open-weight and proprietary systems
Local deployment: Runs on institutional servers, enhancing data sovereignty Integration complexity: Requires technical expertise for deployment and maintenance
No vendor lock-in: Avoids dependency on foreign commercial platforms Version control challenges: No centralized updates; each deployment potentially different
Customization potential: Can be fine-tuned on local data and languages Hallucination risks: Same fundamental limitations as proprietary LLMs, but with less safety testing

Model provenance and data-sovereignty concern:

Institutions should evaluate the model supplier, weights, license, training and evaluation documentation, hosting location, telemetry, update channel, legal jurisdiction, subcontractors, and ability to audit or disconnect the service. This analysis applies to vendors from every country. Local hosting can reduce some data-transfer risks but does not by itself resolve model provenance, software-supply-chain, cybersecurity, licensing, or governance concerns.

Clinical implication:

Open-weight LLMs offer an opportunity to control hosting and adaptation, but deployment requires rigorous local validation, safety monitoring, and clinical oversight regardless of licensing model. Reported hospital adoption should not be converted into an unsupported 300-hospital clinical-deployment claim or a comparison with unspecified regulatory requirements. For institutions prioritizing data sovereignty, local deployment may reduce transfer risk, while provenance and the full software supply chain still require review.

LLM-Enhanced Global Health Applications

DeepDR-LLM: Hybrid AI for Diabetes Care

A multimodal system combining image-based deep learning with language models demonstrates how LLMs can augment primary care capacity in resource-limited settings.

System design:

DeepDR-LLM integrates two components (Li et al., Nature Medicine, 2024):

  1. DeepDR-Transformer: Image-based screening for diabetic retinopathy
  2. LLM module: Personalized diabetes management recommendations for primary care physicians

Training data:

The system was trained on 371,763 real-world management recommendations from 267,730 participants, providing context-specific guidance adapted to Chinese primary care settings.

Prospective validation results:

In a prospective study comparing patients under unassisted primary care physicians (n=397) versus those with PCP + DeepDR-LLM support (n=372):

  • Medication adherence: Patients with newly diagnosed diabetes in the PCP+DeepDR-LLM arm showed significantly better self-management behaviors throughout follow-up (p<0.05)
  • Diabetic retinopathy referrals: For patients with referable DR, those in the PCP+DeepDR-LLM arm were more likely to adhere to referrals (p<0.01)
  • Diagnostic accuracy: Average PCP accuracy for identifying referable DR increased from 81.0% unassisted to 92.3% with DeepDR-Transformer assistance

Key insight:

Hybrid systems can constrain which component performs each task, but greater reliability than an LLM alone must be demonstrated for the exact workflow. A task-specific image model does not automatically prevent unsafe language-model guidance.

MomConnect Enhanced with LLMs

Ong and colleagues describe language-model support for triaging MomConnect enquiries as an emerging application (Ong et al., 2026). The official service includes an automated chatbot and a human-operated helpdesk, but the chapter should not imply that every SMS user is served by an LLM or that an LLM caused the program’s historical reach (South African National Department of Health).

Implementation approach:

  • Base system remains SMS: Preserves accessibility for users with basic feature phones
  • LLM layer for triage: Natural language processing identifies urgency signals in patient messages
  • Human escalation pathway: Urgent cases flagged for immediate nurse review
  • Maintains simplicity for users: No change to patient experience; complexity absorbed by backend systems

Why this hybrid approach is plausible:

The design can use language processing for message routing while preserving a human escalation pathway. Whether reviewers consistently catch errors, whether urgent cases are identified, and whether the pathway improves response time or outcomes must be measured.

Transformer-Based Malaria Detection

Transformer architectures similar to those underlying LLMs have been applied to smartphone-based malaria detection from blood smears, providing scalable alternatives to conventional computer vision approaches (Liu et al., Patterns, 2023).

Technical innovation:

The AIDMAN system uses transformer models optimized for mobile deployment, achieving 98% accuracy in controlled settings on microscopy images captured via smartphone cameras with clip-on lenses.

Deployment challenge:

As with the Thailand malaria AI failure documented in Part 1, performance in real-world border clinics has been more variable. The lesson remains: laboratory validation must be followed by prospective field testing before scale-up.

AfriMed-QA: Addressing the Evaluation Gap

The problem:

Most medical AI benchmarks (USMLE, MedQA) are developed from Western medical education contexts, using disease patterns, treatment options, and healthcare infrastructures common in high-income countries. LLMs that perform well on these benchmarks may fail when applied to African or other LMIC healthcare contexts.

The solution:

AfriMed-QA is the first large-scale pan-African, multi-specialty medical question-answer dataset designed to evaluate LLM performance in contexts relevant to African healthcare (Olatunji et al., ACL 2025).

Dataset composition:

  • ~15,000 questions spanning 32 clinical specialties
  • Contributors: 621 medical professionals from over 60 medical schools across 16 African countries
  • Question types: Expert multiple-choice questions (4,000+), short-answer questions (1,200+), and consumer health queries (10,000)
  • Recognition: Awarded Best Social Impact Paper at ACL 2025

Why this matters:

When 30 different LLMs (large, small, open-weight, proprietary, biomedical-specific, and general-purpose) were evaluated using AfriMed-QA, performance patterns differed substantially from Western benchmarks. Models that excelled on USMLE showed weaker performance on Africa-specific questions, revealing gaps in knowledge about:

  • Endemic infectious diseases (malaria, schistosomiasis, trypanosomiasis)
  • Resource-adapted treatment protocols
  • Traditional medicine interactions
  • Local drug formularies and availability
  • Cultural and linguistic considerations in patient communication

Clinical application:

Before deploying any LLM in African healthcare settings, validation against AfriMed-QA or similar context-specific benchmarks is essential. Performance on Western medical exams does not guarantee performance in LMIC contexts.

Access:

The dataset is publicly available at afrimedqa.com and through the GitHub repository, enabling local researchers to evaluate and fine-tune LLMs for their specific contexts.

Workflow Safety Evidence from African Primary Care

Global physician adoption is structurally constrained:

A 2026 cross-sectional survey of 1,049 physicians across 50 countries and territories found that most respondents reported at least fundamental AI understanding (86.5%) and believed AI would improve clinical practice (80.2%), but only 27.8% had used AI in practice and only 17.7% had received formal training. Formal training was associated with more than threefold higher odds of AI use, while working in an institution with AI technologies was associated with more than eightfold higher odds (Bold et al., 2026).

Clinical interpretation:

The global AI access problem is not limited to open-weight models or cloud costs. If physicians lack formal training and institutions lack deployable systems, high awareness does not translate into safe use. Equity-oriented AI implementation should pair model access with training, procurement support, local validation capacity, and governance infrastructure.

A safety evaluation from African primary care (2026):

A retrospective evaluation of an EMR-embedded LLM clinical decision support system across 16 outpatient clinics in Kenya (July–September 2024, n=1,469 patient encounters) provides the most granular real-world evidence to date for LLM deployment in African primary healthcare (Agweyu et al., Nature Health, 2026).

Key findings:

  • Clinical guidance aligned with Kenyan Ministry of Health guidelines in 99% of encounters
  • Actively harmful recommendations appeared in 7.8% of encounters (115 cases), including inappropriate antibiotic prescribing and incorrect referral guidance
  • In 62% of encounters where LLM output appeared in the documentation, clinicians did not modify that text before finalizing the record. This is compatible with reliance but does not by itself prove automation bias.
  • Harmful output was reflected in final documentation in 67 encounters, 5% of the evaluated sample, a direct safety signal in this retrospective review

Clinical interpretation:

This study directly complicates an optimistic framing of LLMs in LMIC settings. High scores on broad guideline-alignment items coexisted with actively harmful recommendations in 115 of 1,469 encounters. Low editing and harmful text in final documentation warrant investigation, but the retrospective design does not establish how the system changed patient outcomes relative to no LLM.

The lesson is not that LLMs should not be deployed in LMIC primary care, but that deployment without explicit safety monitoring and clinician override training is premature. The AfriMed-QA benchmark evaluates knowledge; this study measures safety in real clinical workflows. Both are necessary before scale-up.

A randomized simulated-vignette trial with physicians in Pakistan:

Sixty licensed physicians in Pakistan were randomized to GPT-4o access plus conventional resources or conventional resources alone following a 20-hour AI-literacy curriculum, and 58 completed the study. Physicians with LLM access scored 27.5 percentage points higher on diagnostic reasoning across clinical vignettes (71.4% versus 42.6%) (Qazi et al., 2026). This is evidence for a trained, simulated-vignette workflow, not patient diagnosis or outcomes. Unstructured access without training may not reproduce the result.

A multi-country randomized trial of GPT-4o assistance on standardized clinical vignettes (N=249 physicians in Indonesia, Kenya, and the Netherlands) found larger absolute performance gains in Kenya (+18 percentage points) than in Indonesia (+10.7 percentage points) or the Netherlands (+7.2 percentage points), with a reduction in the Kenya–Netherlands performance gap under GPT-4o assistance (Rounding et al., 2026). This is vignette performance under controlled conditions: the control arm lacked traditional resources such as internet and guidelines, harms were not assessed, and the result is not patient-outcome evidence.

Environmental Justice and LLM Sustainability

The hidden cost of computational intensity:

Training and deploying LLMs consume vast computational resources, with environmental impacts that disproportionately affect LMICs even when the technology is developed and deployed primarily in high-income countries.

Energy and emissions profile:

  • Training: Energy and emissions depend on hardware, data center, grid, model, training regime, and accounting boundary
  • Inference costs: While individual queries consume relatively little energy, cumulative daily use across healthcare applications (documentation, triage, decision support) scales dramatically
  • Water consumption: Data centers require substantial water for cooling; water scarcity is more acute in many LMIC regions
  • Hardware lifecycle: Rare-earth mining for GPUs and electronic waste disposal create environmental burdens concentrated in resource-extraction regions

The equity dimension:

LMICs contribute minimally to AI development emissions but bear disproportionate climate impacts. As healthcare systems in high-income countries adopt LLM-based documentation, clinical decision support, and administrative tools at scale, the cumulative carbon footprint grows while benefits accrue primarily to well-resourced settings.

Sustainable deployment strategies:

  1. Compare edge and cloud: Local hosting can reduce some data transfer and sovereignty risks but may use less efficient hardware; measure total energy, utilization, cooling, maintenance, and grid intensity
  2. Model efficiency: Smaller, task-optimized models rather than general-purpose large models where appropriate
  3. Renewable energy: Data centers powered by solar, wind, or hydroelectric rather than fossil fuels
  4. Shared infrastructure: Regional services may improve utilization, but jurisdiction, resilience, network, access, and concentration risks require evaluation

Policy implication:

As LLMs are evaluated for global health applications, environmental sustainability should be an explicit criterion, alongside clinical effectiveness and cost-effectiveness. The appropriate system is the least resource-intensive one that meets the verified clinical and operational objective.

Global Governance and Coordination Initiatives

WHO Global Initiative on AI for Health (GI-AI4H):

Launched in July 2023 by WHO, the International Telecommunication Union (ITU), and the World Intellectual Property Organization (WIPO), GI-AI4H provides an institutional framework for coordinating responsible AI development and deployment globally (WHO/ITU/WIPO, 2023).

Strategic focus areas:

  1. Standards development: International standards and normative guidance for AI evaluation, ethics, clinical validation, and benchmarking
  2. Knowledge transfer: Facilitating data sharing, collaboration, and best practice dissemination among stakeholders worldwide
  3. Health system strengthening: Prioritizing low- and middle-income countries through a scaling program initially targeting 12-18 countries with relevant AI use cases

Lancet Global Health Commission on AI and HIV:

This Commission synthesizes evidence on AI’s economic and health impacts across different settings, with explicit focus on guiding responsible AI model development and creating actionable guidance for stakeholders in regulation and adoption (Lancet Global Health, ongoing).

Gates Foundation AI equity initiatives:

Philanthropic funding targeting AI equity, language inclusivity, and ensuring equitable access to AI benefits in all settings. Projects focus on developing language technologies for under-resourced languages and supporting local capacity building (Gates Foundation, 2024-2026).

Implication for physicians:

These frameworks provide actionable guidance for institutions deploying LLMs in global health contexts. Rather than navigating regulatory ambiguity alone, leveraging WHO standards, Lancet Commission recommendations, and philanthropic partnership opportunities can accelerate responsible implementation.

Practical Guidance for LLM Deployment in LMICs

Pre-deployment checklist:

Before deploying any LLM system in resource-limited settings:

  1. Validate on context-specific benchmarks (AfriMed-QA for Africa; similar datasets for Asia, Latin America where available)
  2. Assess infrastructure requirements (internet connectivity, computational capacity, electricity reliability)
  3. Evaluate language support (does LLM support local languages and dialects, or only English?)
  4. Verify safety mechanisms (how does system handle uncertainty, avoid hallucinations, escalate to humans?)
  5. Establish clinical oversight (who reviews LLM outputs before clinical action?)
  6. Plan for sustainability (who maintains system when external funding ends? what is long-term cost structure?)
  7. Address data sovereignty (where is patient data processed and stored? does this comply with local regulations?)
  8. Measure environmental impact (what are energy/water requirements? can renewable energy support deployment?)

Red flags warranting rejection:

  • LLM vendor cannot demonstrate performance on LMIC-specific benchmarks
  • System requires always-on internet connectivity in settings with unreliable access
  • No mechanism for clinical oversight of LLM outputs
  • Deployment plan lacks sustainable financing beyond donor pilot period
  • Patient data will be processed on foreign servers without clear data protection agreements
  • Vendor dismisses environmental impact questions or lacks sustainability data

Green flags supporting adoption:

  • Prospective validation in deployment setting (not just Western benchmarks)
  • Offline/edge deployment capability for low-connectivity environments
  • Human-in-the-loop design with clear escalation pathways
  • Open-weight model with local customization and fine-tuning potential
  • Integration with existing health information systems and workflows
  • Government or institutional ownership with budget commitment
  • Explicit environmental sustainability assessment

The Path Forward: Co-Development Not Deployment

The extractive model (what to avoid):

  1. HIC institution develops LLM on Western data
  2. Pilots system in LMIC with external funding
  3. Publishes papers on “global health AI”
  4. Funding ends, system abandoned
  5. No local capacity remains

The co-development model (what to pursue):

  1. Joint problem definition: LMIC and HIC partners identify priority use cases together
  2. Shared data governance: Training data includes LMIC contexts, with local ownership and benefit-sharing
  3. Capacity building from inception: Local researchers co-lead development, not just deployment
  4. Prospective validation: Local clinical validation before scale-up
  5. Sustainable financing: Government integration and budget commitment, not donor dependency
  6. Knowledge transfer: Local teams can maintain, improve, and adapt systems independently

Physician role:

Physicians in high-income countries evaluating LLM vendors or research partnerships should demand evidence of co-development, equitable partnerships, and sustainable implementation rather than extractive “deploy and abandon” models. Physicians in LMICs should advocate for local ownership, capacity building, and benefit-sharing rather than accepting passive recipient roles.


Check Your Understanding

All three exercises are hypothetical. The organizations, products, authorizations, datasets, patient mix, performance results, harms, publications, financing, contracts, utilization values, legal positions, and remedies are fictional unless an external source is explicitly cited. They are prompts for due diligence, not reports of events in the Democratic Republic of Congo, Kenya, or Bangladesh.

Scenario 1: Deploying Western AI in LMIC Without Validation

You’re an internist working with Doctors Without Borders (MSF) in rural Democratic Republic of Congo (DRC). Hospital receives donated AI chest X-ray system (FDA-cleared in U.S.) for pneumonia detection.

AI specifications: - Trained on 200,000 U.S. chest X-rays (95% from academic medical centers) - Validation: Sensitivity 92%, specificity 88% on U.S. test set - FDA 510(k) cleared (2022)

Your context: - DRC hospital serves population with high HIV prevalence (8%), TB (350 per 100,000), malnutrition (18% of adults underweight) - X-ray machine: 15-year-old analog system (vs. digital systems in U.S. training data) - No on-site radiologist (you interpret X-rays with basic training)

You deploy AI as primary pneumonia screening tool (all patients with respiratory symptoms get CXR + AI interpretation).

3 months later: Retrospective chart review by visiting radiologist identifies problems: - AI sensitivity for TB: 58% (missed 42% of culture-proven TB) - AI sensitivity for Pneumocystis jirovecii pneumonia (PCP, in HIV+ patients): 34% - AI false positive rate: 35% (vs. 12% in U.S. validation)

Question 1: Why did AI fail in DRC?

Training data mismatch: 1. Disease prevalence: U.S. training data had <1% TB, <0.1% PCP; DRC has 30%+ TB, 5%+ PCP in respiratory presentations 2. Patient characteristics: U.S. training data mostly non-HIV, normal BMI; DRC population 8% HIV+, 18% underweight (atypical radiographic findings) 3. Image quality: AI trained on digital X-rays; DRC analog system produces lower resolution, different artifact patterns 4. Co-morbidities: AI learned “pneumonia” from U.S. bacterial pneumonia; missed opportunistic infections (PCP, TB) uncommon in training data

Question 2: What harm occurred?

Clinical impact: - 42% of TB missed, leading to delayed diagnosis, transmission to contacts, and TB mortality - 66% of PCP missed, resulting in HIV+ patients with untreated PCP and respiratory failure (PCP mortality 30-50% if untreated) - 35% false positives, leading to unnecessary antibiotics, costs, and patient anxiety

Trust impact: - Hospital staff lost confidence in AI: “The American computer doesn’t work here” - Some stopped using AI, others over-relied on negative results

Question 3: Are you liable?

Liability cannot be determined from the vignette. Relevant legal and ethical issues include:

Standard of care: Even in resource-limited settings, physicians must provide care meeting local standards. - Local validation, intended use, reasonable care under the available conditions, institutional duties, and causation would require fact- and jurisdiction-specific analysis

Informed consent: Did you inform patients that AI was not validated in DRC population?

FDA authorization is jurisdiction- and intended-use specific: It does not create local authorization or guarantee transportability to another population, device, protocol, or care pathway

Plaintiff (patient or family) argument: - Physician deployed unvalidated AI - AI missed TB/PCP, patient died from delayed treatment - Standard of care requires validation in deployment population OR not using AI for high-stakes decisions

Defense argument: - Resource-limited setting, no radiologist available - AI was FDA-cleared, represented “best available tool” - Physician acted in good faith with limited resources

Lesson: 1. FDA clearance ≠ universal validity. Validate in deployment population 2. High-risk populations (HIV, malnutrition, endemic diseases) require local validation 3. If validation impossible, use AI as adjunct (not replacement) for clinical judgment 4. Document limitations: “AI chest X-ray interpretation not validated in this population; clinical judgment takes precedence”

Scenario 2: Data Partnership with Benefit-Sharing Failure

You’re chief of medicine at teaching hospital in Nairobi, Kenya. U.S. university approaches with proposal:

Proposal: “We’re developing AI for early sepsis detection. We need diverse African data to make our model generalizable. Can you provide 50,000 patient records (ICU admissions, vitals, labs, outcomes)? In return, we’ll co-author you on publications and acknowledge your hospital.”

You agree. Hospital IT department exports 50,000 de-identified records, sends to U.S. university.

18 months later: - U.S. team publishes 3 papers in Nature Medicine, JAMA, Critical Care Medicine - Your hospital listed in acknowledgments (not co-authorship as promised) - AI commercialized: U.S. startup licenses technology, raises $25M Series A funding - Startup offers to sell AI back to your hospital: $50,000 annual licensing fee

Your hospital administration: “We provided the data for free, now they want us to pay $50K/year for our own data? This is exploitation.”

Question 1: What went wrong?

Classic data extraction: 1. Verbal agreement only: No written contract specifying co-authorship, IP rights, benefit-sharing 2. Data transfer without equity stake: Hospital gave away valuable asset (50K patient records) for vague promises 3. No benefit-sharing clause: U.S. team commercialized AI, Kenyan hospital received nothing 4. Insufficient oversight: Hospital did not involve legal team, tech transfer office before data sharing

Question 2: Do you have legal recourse?

Weak legal position: - No written contract → hard to enforce verbal promises - Data was “de-identified” → hospital may not have ownership claim - Kenya data protection laws (2019) were new, untested in court - U.S. university likely claims IP ownership (researchers developed AI)

Possible arguments: - Breach of verbal contract (co-authorship promised, not delivered) - Unjust enrichment (U.S. team profited from Kenyan data without compensation) - Data sovereignty violation (Kenya Data Protection Act requires data localization for processing)

Outcome: The exercise cannot predict litigation or settlement. Counsel would need the actual agreements, data rights, representations, ethics approvals, privacy law, intellectual-property facts, and jurisdiction.

Question 3: What should you have done differently?

Before sharing data:

Written data-sharing agreement specifying: - Co-authorship: Kenyan researchers as co-first/co-corresponding authors on all publications - IP ownership: Shared IP rights (hospital owns data, university owns algorithm, jointly own AI) - Benefit-sharing: If commercialized, negotiate the form and amount of equity, royalties, milestones, subsidized access, reinvestment, or public benefit rather than importing arbitrary percentages - Free access: Resulting AI provided to hospital at no cost - Capacity-building: U.S. team trains 3-5 Kenyan researchers in AI development - Data governance: Data hosted in Kenya (or federated learning, no data export)

Institutional approvals: - Legal team review - Ethics committee approval - Tech transfer office involvement (protect IP) - Data protection officer review (Kenya Data Protection Act compliance)

Pilot first: Start with 1,000 records, evaluate partnership quality before sharing 50,000

Lesson: Data is valuable. Demand equitable partnerships, not exploitation. LMIC hospitals should negotiate from position of strength (you have data they need, so demand fair value).

Scenario 3: AI Exacerbating Health Disparities

You’re public health official in Ministry of Health, Bangladesh. Government considering national rollout of AI-powered telemedicine platform for rural primary care clinics.

Platform features: - Smartphone app (patients video-call physician + AI clinical decision support) - AI provides differential diagnosis, treatment recommendations based on symptoms + patient history - Targets rural areas with physician shortages (1 physician per 50,000 people)

Deployment plan: - Distribute subsidized smartphones ($100 each) to 100,000 rural households - Train 500 community health workers to assist patients using app - Budget: $15 million over 3 years

12 months post-launch, evaluation shows:

Utilization by wealth quintile:

Wealth Quintile % Using Telemedicine % Using Traditional Clinics Sum of Channel-Use Percentages, Not Unique Access
Richest 20% 62% 45% 107% (using both)
Middle 60% 28% 38% 66%
Poorest 20% 9% 22% 31%

Digital divide drivers: - Poorest 20% often lack literacy (41% illiterate) and struggle with smartphone app despite CHW assistance - Smartphone distribution focused on households with electricity (excluded 38% of poorest quintile) - Data costs (mobile internet): $2-5/month equals 5-10% of poorest households’ income, making it unaffordable despite subsidized smartphones - Language: App in Bengali only, excluded 12% of population speaking minority languages (Chittagonian, Sylheti, tribal languages)

Health equity impact: - Richest quintile: Healthcare access increased 7% (using telemedicine in addition to clinics) - Poorest quintile: Healthcare access decreased 9% (traditional clinics closed due to “telemedicine coverage”, but poorest cannot access telemedicine)

Illustrative result: The fictional program widened the modeled access gap by 16 percentage points. This is not a measured Bangladesh result, and the channel-use percentages cannot be added to estimate unique people with access when users overlap.

Question 1: Why did telemedicine worsen equity?

Inverse care law (Julian Tudor Hart, 1971): “The availability of good medical care tends to vary inversely with the need for it in the population served.”

Applied to AI: - Telemedicine benefited those with smartphones, literacy, internet, electricity (already better-off) - Excluded those lacking these resources (poorest, who need healthcare most) - Clinic closures harmed poorest (who depended on in-person care) while richest benefited from telemedicine

Question 2: What should have been done differently?

Equity-focused design:

Universal access pre-requisites: - Ensure electricity, internet, smartphones, literacy BEFORE telemedicine rollout - OR: Design for low-tech (SMS, voice calls, not video; feature phones, not smartphones)

Hybrid model: - Telemedicine supplements clinics (not replaces) - Maintain in-person care for those who cannot access digital tools

Targeted support for poorest: - Free data plans (not just subsidized phones) - Voice-based apps (for illiterate users) - Minority language support - CHW home visits for those unable to use technology

Equity monitoring: - Track utilization by wealth, literacy, language from Month 1 - If disparities emerge, pause deployment and redesign

Question 3: How to fix the program now?

Immediate actions:

  1. Reopen closed clinics in poorest areas (do not rely solely on telemedicine)
  2. Free data plans for poorest quintile (government subsidy to mobile operators)
  3. Voice-based app version for illiterate users
  4. Multilingual support (Chittagonian, Sylheti, tribal languages)
  5. CHW-assisted telemedicine: CHWs help poorest households access platform (human-in-the-loop)

Lesson: AI can widen disparities if designed for already-privileged populations. Equity requires intentional design for most marginalized, not just “average” users. Monitor equity impacts from start and correct course quickly.


Questions About AI and Global Health Equity

Does Western-trained AI work in low-income countries?

Performance cannot be inferred from the country where a system was developed. It can change with disease prevalence, devices, protocols, language, image quality, referral pathways, and population characteristics, so local acceptance testing and monitoring are required.

What AI tools work for tuberculosis screening?

Computer-aided detection products can support tuberculosis screening from digital chest radiographs, but performance, thresholds, intended use, and workflow differ by product and setting. WHO recommendations and each current product record should be checked before procurement.

Can smartphone AI diagnose diseases?

Smartphone-connected imaging can support selected acquisition and classification tasks, but a laboratory or retrospective accuracy result does not establish field performance, clinical utility, affordability, device durability, or regulatory status.

Why do AI health pilots fail in LMICs?

Pilots can fail when they do not plan for local ownership, financing, maintenance, workforce training, connectivity, power, supply chains, monitoring, and decommissioning. There is no universal failure percentage or requirement that every system be cloud-independent.

What is data colonialism in health AI?

Data colonialism occurs when health data is extracted from LMICs for AI development without local benefit. Ethical AI requires equitable data partnerships and capacity building.

Key Takeaways

  1. Validation is contextual: Performance can change across populations, devices, protocols, prevalence, language, and workflows. Require local acceptance testing and monitoring before consequential use.

  2. Low-burden technology can scale: MomConnect reached national scale through public-system integration, messaging, and a helpdesk. Its scale does not establish a 15:1 return or AI-caused health outcomes. Match technology to infrastructure and care pathways.

  3. Data Extraction ≠ Partnership: Demand benefit-sharing (equity, royalties, free access) when sharing LMIC data for AI development. Reject exploitation.

  4. Algorithmic bias is measurable: Representation, labels, devices, prevalence, and workflow can create unequal performance. Diversify evidence and conduct subgroup and access audits with uncertainty.

  5. Infrastructure matters: Current official estimates identify hundreds of millions without electricity and billions offline, while facility-level reliability and affordability vary. Design for measured conditions, not global averages.

  6. Equity Requires Intention: AI follows “inverse care law” and benefits privileged populations unless explicitly designed for marginalized groups. Monitor equity impacts from Day 1.

  7. Capacity-Building > Extraction: Train local AI researchers, leave sustainable infrastructure. Extractive partnerships perpetuate dependency.

  8. Physician Advocacy Role: Demand equitable AI development, reject unvalidated deployments, support open-source tools, partner with LMIC institutions as equals.