Pediatrics and Neonatology
Children are not small adults, and pediatric AI must account for this fundamental reality. A normal heart rate for a neonate would trigger tachycardia alerts in adults. Growth trajectories, vital sign ranges, and medication dosing vary continuously with age. Most medical AI has been developed in adults, leaving pediatric populations systematically underrepresented and creating both safety concerns and opportunities for child-specific applications.
After reading this chapter, readers will be able to:
- Evaluate AI systems for neonatal intensive care monitoring and prediction
- Understand pediatric imaging AI and developmental considerations
- Assess AI tools for growth monitoring, developmental screening, and chronic disease management
- Navigate ethical challenges of AI in pediatric medicine
- Identify failure modes specific to pediatric populations
- Recognize equity concerns in pediatric AI (algorithmic bias against children)
- Apply evidence-based frameworks for pediatric AI adoption
Introduction
Pediatric AI faces challenges that adult medicine does not: smaller training datasets, rapidly changing developmental physiology, and higher stakes when errors affect an entire lifespan. Adult reference ranges, dosing algorithms, and imaging parameters do not transfer directly to children.
Several pediatric computational interventions have demonstrated clinical value, but they do not all use machine learning. The Kaiser Permanente Neonatal Early-Onset Sepsis Calculator is a conventional multivariable risk model. HeRO is conventional heart-rate-variability monitoring. Closed-loop oxygen control and automated insulin delivery are control systems whose algorithms should be described from their labeling rather than automatically classified as AI. Canvas Dx is a machine-learning diagnosis aid that received FDA De Novo authorization in 2021.
Regulatory tracking shows how far pediatrics still has to go. Of 876 FDA-authorized AI/ML devices (through March 2024), only 149 (17.0%) were labeled for pediatric use, and only 28 of those (18.8%) explicitly reported validation datasets that included children (Brewster et al., JAMA Pediatrics, 2025).
A scoping review of LLMs using pediatric clinical text found only retrospective observational studies and no evaluation within real-time pediatric workflows. Age-specific analyses and reporting were often incomplete. These findings concern clinical-text research through July 2025, not all pediatric AI or subsequent deployments (Huang et al., 2026).
AI Applications in Neonatology (NICU)
1. Neonatal Sepsis Prediction
Early-Onset Sepsis (EOS) Risk Calculators:
Traditional approach: Categorical maternal-risk assessment or serial clinical examinations
Risk model: Kaiser Permanente Neonatal Early-Onset Sepsis Calculator (Kuzniewicz et al., 2017). The calculator is a conventional multivariable statistical model, not a machine-learning system.
Evidence:
- In an observational implementation study of 204,485 infants born at 35 weeks or later, empiric antibiotic use during the first 24 hours fell from 5.0% to 2.6% and blood-culture use from 14.5% to 4.9%, with no reported increase in adverse outcomes or readmissions (Kuzniewicz et al., 2017)
- The AAP clinical report identifies multivariate risk assessment with the EOS Calculator as one acceptable strategy alongside categorical assessment and serial clinical examinations (Puopolo et al., 2018)
- A systematic review and meta-analysis found reduced empiric antibiotic use with calculator-guided management; evidence on safety was limited. This review pools implementation cohorts and is not a separate external-validation cohort (Achten et al., 2019)
- The underlying risk model was derived from 350 culture-confirmed infections among 608,014 live births and showed good discrimination (c statistic 0.800) (Puopolo et al., 2011)
How it works:
- Integrates maternal risk factors (GBS status, intrapartum antibiotics, temperature, ROM duration)
- Infant clinical signs (activity, respiratory status, temperature)
- Provides personalized infection risk estimate
- Guides empiric antibiotic decision-making
Model version and limitations:
The developer group updated the calculator using a contemporary birth cohort of infants born at 35 weeks or later, re-estimating model coefficients and clinical-status likelihood ratios. This is a model update, not independent external validation; implementation should identify the calculator version and the population evaluated (Kuzniewicz et al., 2024).
- Accurate maternal-risk and infant clinical-status data are necessary for interpretable estimates
- Reduced antibiotic exposure in one implementation does not establish the same effect in every setting
- A sick-appearing infant requires clinical assessment and management independent of a reassuring numerical estimate
A systematic review of retrospective and prospective cohorts found that the calculator could delay or omit initial antibiotic recommendations in some culture-positive cases compared with NICE guidance. Its outcome was a difference in treatment recommendations, not an estimate of excess mortality; it should be interpreted alongside prospective workflow evidence and serial clinical assessment (Pettinger et al., 2020).
2025 Cluster RCT Update: An open-label, two-arm cluster randomized trial across 10 hospitals in the Netherlands enrolled 1,830 neonates born at 34 weeks or later with at least one EOS risk factor. EOS Calculator-guided management reduced antibiotic starts for suspected EOS from 26.6% to 7.2% (absolute risk reduction 19.0%, 95% CI 11.3–26.7). Fewer newborns met the trial’s predefined harm criteria (7.0% versus 14.6%; relative risk 0.48, 95% CI 0.36–0.63), while median antibiotic duration among treated infants was longer in the calculator arm (5.5 versus 2.1 days) (van der Weijden et al., 2025). This randomized result supports one risk-stratified EOS workflow; it is not class-wide evidence for pediatric AI.
Late-Onset Sepsis (LOS) Prediction in Preterm Infants:
Challenge: Preterm infants are at high risk for sepsis, and most very low birth weight infants receive empiric antibiotics in the absence of a culture-confirmed infection (Puopolo et al., 2018), but clinical signs of late-onset sepsis are nonspecific
Computational approaches:
- Continuous heart-rate-variability monitoring
- Statistical or machine-learning analysis of vital-sign and laboratory trajectories
- Multimodal research models combining physiologic and clinical data
Evidence:
- HeRO (Heart Rate Observation) monitor (Moorman et al., 2011):
- Analyzes heart rate characteristics (reduced variability, decelerations)
- Randomized trial (N=3003 VLBW infants) showed 22% relative reduction in mortality (Moorman et al., 2011)
- FDA 510(k)-cleared under K021230 for heart-rate-variability analysis and reporting under licensed supervision in neonatal or pediatric intensive care (FDA K021230)
- The FDA indication states that HeRO does not interpret the measurements or provide a diagnosis, and its public device record does not establish a machine-learning model
- Published in The Journal of Pediatrics (Moorman et al., 2011)
Subsequent implementation studies:
- Variable mortality benefit in real-world settings (Kumar et al., 2020)
- Depends on care team responses to alerts
- Heart rate characteristic trajectories varied widely between infants, and trajectory patterns predicted 7-day and 30-day mortality beyond individual HRC values in 2,989 very low birth weight infants across nine NICUs (Zimmet et al., 2020)
Limitations:
- A rising heart-rate-characteristics index is nonspecific and can create additional evaluations or treatment when response protocols are poorly calibrated
- Does not identify pathogen (empiric antibiotics still required)
- Requires continuous cardiorespiratory monitoring infrastructure
- Training required for appropriate alert interpretation
Implementation challenges:
- Integration with existing NICU monitors
- Nurse and physician education on alert response
- Protocols needed to avoid reflexive antibiotics for every alert
HeRO has FDA 510(k) clearance for HRV analysis and reporting, plus randomized evidence for the complete monitor-display-and-clinician-response intervention (FDA K021230; Moorman et al., 2011). The trial does not establish that the monitoring score alone causes benefit, and the public FDA record does not establish HeRO as machine learning. Implementation still requires explicit evaluation and treatment protocols to limit alert fatigue and reflexive antibiotic use (Kumar et al., 2020; Zimmet et al., 2020).
2. Neonatal Respiratory Support
Automated Oxygen Titration for Preterm Infants:
Clinical problem: Preterm infants require narrow oxygen saturation targets (88-95%) to minimize retinopathy of prematurity (ROP) risk and bronchopulmonary dysplasia (BPD) while preventing hypoxia (Stenson et al., 2013).
Manual titration limitations:
- Frequent SpO2 fluctuations
- Nurse workload (adjustments every 15-30 minutes)
- Time within the intended SpO2 range was only 32% in manual mode in a crossover study of 32 ventilated preterm infants (Claure et al., 2011)
Automated-control intervention: Closed-loop oxygen controllers adjust inspired oxygen from pulse-oximetry feedback. This intervention is clinically relevant to AI-enabled care, but closed-loop control is not necessarily machine learning.
Evidence:
- Multiple randomized crossover studies show that automated systems can increase time in the intended saturation range (Claure & Bancalari, 2015):
- Time in range was 62% with automated control versus 54%–58% with manual control in an 80-infant crossover study. Time with SpO2 below 80% was consistently reduced, while hyperoxemia fell only when the lower saturation range was targeted (van Kaam et al., 2015)
- Cochrane review (18 studies, 457 infants) found automated delivery probably increases time in the target SpO2 range (mean difference 13.5 percentage points, moderate-certainty evidence), but whether this translates into clinical benefit remains unclear (Stafford et al., 2023)
Regulatory status: Product-specific authorization and labeling must be verified for the exact controller and ventilator configuration being procured. Evidence for one system should not be transferred to another.
Long-term outcomes:
- No difference in neurodevelopment at 2 years (Lal et al., 2015)
- Reduced severe ROP in some studies (Zapata et al., 2014)
Limitations:
- Requires reliable pulse oximetry (motion artifacts problematic)
- Does not replace clinical assessment for escalation/de-escalation of support
- Alarms still require nurse response
- Procurement, integration, training, and maintenance costs vary by system and institution
Randomized crossover evidence and a Cochrane review support improved time within the target saturation range (Claure & Bancalari, 2015; Stafford et al., 2023). Whether the control strategy improves longer-term patient outcomes remains uncertain, so procurement should specify both process and clinical endpoints.
3. Neonatal Parenteral Nutrition
Total parenteral nutrition (TPN) is essential in premature and critically ill neonates, but orders are complex, vary by clinician, and create safety and workload burdens in NICUs.
TPN2.0: A 2025 Nature Medicine study trained a data-driven neonatal TPN recommendation system on 79,790 orders from 5,913 Stanford patients and externally validated it on 63,273 orders from 3,417 patients at a second hospital. The model identified 15 standardized TPN formulas with strong agreement against expert practice (Pearson R=0.94), and physicians rated the recommendations higher than current practice in a blinded study of 192 orders (Phongpreecha et al., 2025).
Clinical interpretation: This is a promising physician-in-the-loop standardization tool, not proof that AI-guided TPN improves neonatal outcomes. Associations between disagreement with TPN2.0 and morbidity were observational. Before deployment, NICUs need local formulary review, pharmacy integration, nutrition-team oversight, and prospective monitoring for electrolyte, growth, and complication outcomes.
4. Retinopathy of Prematurity (ROP) Screening
AI-Assisted ROP Detection:
Clinical problem: ROP affects 14,000+ preterm infants annually in US (Hellström et al., 2013). Requires serial dilated retinal exams by ophthalmologists. Severe ROP requires urgent treatment to prevent blindness.
Traditional screening: Modified from AAP guidelines (Fierson et al., 2018) - Infants <1500g birthweight or ≤30 weeks gestation - First exam at 31 weeks postmenstrual age or 4 weeks chronologic age - Serial exams until retina mature - Ophthalmologist-intensive process
AI solution: Automated ROP detection from retinal images
Evidence:
- i-ROP system (Brown et al., 2018):
- Identifies plus disease (severe ROP) with 93% sensitivity, 94% specificity on an independent 100-image test set
- Trained on 5,511 retinal photographs collected from 8 academic institutions in the i-ROP cohort
- Published in JAMA Ophthalmology (Brown et al., 2018)
- Matches expert consensus better than individual ophthalmologists
- Automated detection of ROP stage (Chen et al., 2021):
- A model trained on North American images reached AUROC 0.99 and 94% sensitivity on North American test images, but sensitivity fell to 52% on Nepali images; training on the combined dataset raised sensitivity to 98% and 82% respectively
- Performance depends on how closely the deployment population and camera match the training data
- Published in Ophthalmology Retina (Chen et al., 2021)
- Deep learning models (Redd et al., 2019):
- Evaluated on 4,861 examinations from 870 infants at seven centers
- Detect type 1 ROP with AUROC 0.96; at a vascular severity score threshold of 3, sensitivity was 94% and specificity 79%, with 13% positive predictive value and 99.7% negative predictive value
Current status:
- The studies cited here do not establish an FDA-authorized autonomous ROP diagnosis; current status must be verified for the exact product in the FDA record
- Research and telemedicine workflows retain ophthalmologist confirmation
- Telemedicine applications for under-resourced NICUs (Greenwald et al., 2020)
Limitations:
- Image quality critical (hazy media, poor dilation reduce accuracy)
- Peripheral retina visualization challenging
- Does not eliminate need for ophthalmologist expertise
- Rare cases may be missed (sensitivity not 100%)
- Most studies from academic centers with high-quality imaging
Equity implications:
- Could improve access to ROP screening in rural/under-resourced areas
- Telemedicine + AI may reduce disparities in ophthalmologist availability
- But requires imaging infrastructure and technical support
Future direction: Autonomous ROP screening remains a regulatory and clinical-evidence question. Research performance cannot predict whether a product will receive FDA authorization or improve patient outcomes after deployment.
The studies support AI-assisted reader and triage research (Brown et al., 2018; Chen et al., 2021). They do not establish autonomous diagnosis, broad transportability, or improved visual outcomes. Ophthalmologist confirmation and an explicit pathway for ungradable images remain necessary.
Infantile fundus abnormality screening
Closed-set retinal models can be overconfident when infant disease presentations or image quality fall outside the training set. A 2026 npj Digital Medicine multicenter study trained an uncertainty-inspired open-set YOLOv8 system (UIOS-Y) on 319,998 fundus images from 15,647 infants at 39 centers (Zhao et al., 2026). Internal AUC was 0.995; excluding high-uncertainty predictions (UIOS-Y+θ) raised AUC to 0.997 and reserved ambiguous cases for clinician review. Five independent external datasets held AUCs of 0.969–0.982, and the system detected out-of-distribution images and outperformed resident physicians and general-purpose chatbots (ChatGPT and Gemini; all p < 0.001). This is retrospective multi-class screening with uncertainty triage, not autonomous ROP diagnosis and not an FDA-authorized device.
5. Neonatal Neuroimaging and Brain Injury Prediction
Hypoxic-Ischemic Encephalopathy (HIE) Severity Assessment:
Clinical problem: HIE affects 1-2/1000 term births (Kurinczuk et al., 2010). Therapeutic hypothermia improves outcomes if initiated <6 hours after birth. Severity assessment guides cooling decisions and prognostication.
Traditional assessment: Clinical exam (Sarnat staging) + aEEG or EEG
AI approaches:
- MRI pattern grading for injury prediction (Martinez-Biarge et al., 2011):
- Severity grading of basal ganglia-thalamic injury on early MRI in 175 term infants with HIE, by expert visual assessment rather than a deep learning model
- Predicts motor outcome and death at 2 years
- Predictive accuracy of severe basal ganglia-thalamic lesions for severe motor impairment was 0.89 (95% CI 0.83-0.96)
- Published in Neurology (Martinez-Biarge et al., 2011)
- EEG pattern recognition (Pavel et al., 2020):
- Automated seizure detection in neonates
- In a randomised trial, algorithm support raised the proportion of seizure hours correctly identified from 45% to 66%, without improving recognition of which neonates had seizures
- Evaluated with continuous conventional EEG (cEEG), not amplitude-integrated EEG
- Early EEG grade at 6 hours stratified 5-year outcome: intact survival was 75% with mild, 46% with moderate, 43% with major EEG abnormalities and 0% with inactive EEG, versus 97% in comparison infants (Murray et al., 2016)
- Multi-modal prediction models (Wusthoff et al., 2022):
- Combine clinical data, MRI, EEG, biomarkers
- Predict outcomes at 18-24 months with AUC 0.88-0.92
- Published in Pediatric Research (Wusthoff et al., 2022)
Limitations:
- MRI typically performed day 4-7 (after acute decisions made)
- aEEG expertise limited outside major centers
- Prediction models need prospective validation in diverse populations
- Group-level discrimination does not establish reliable prognosis for an individual infant
- Long-term outcomes (school age, adolescence) less well predicted
Ethical considerations:
- Outcome predictions influence decisions about withdrawal of life-sustaining treatment
- False predictions have devastating consequences (both directions)
- Must not be sole basis for prognostic discussions
- Family values and goals central to decision-making
- Cultural attitudes toward disability and life-sustaining treatment vary
The research is promising (Martinez-Biarge et al., 2011; Wusthoff et al., 2022), but a prediction score should not be the sole basis for prognosis or withdrawal-of-treatment decisions. Study-level discrimination does not remove uncertainty for an individual infant. Serial clinical assessment, multimodal evidence, family values, and specialist judgment remain essential.
AI Applications in General Pediatrics:
6. Growth and Development Monitoring
Automated Growth Chart Analysis:
Application:
- WHO/CDC growth chart plotting from EHR weight/height data
- Identification of abnormal growth patterns (failure to thrive, obesity, growth deceleration)
- Alerts for crossing percentiles
Evidence:
- Automated screening of EHR growth data identified implausible measurements with 97% sensitivity and 90% specificity against physician review, across 2,000,595 measurements from 280,610 children (Daymont et al., 2017)
- Clean growth data is a prerequisite for detecting true growth abnormalities such as Turner syndrome (short stature), celiac disease (growth deceleration), and growth hormone deficiency
- Published in JAMIA (Daymont et al., 2017)
Implementation:
- Built into most modern EHR systems (Epic, Cerner)
- Requires accurate measurement documentation
- False positives with measurement errors (incorrect length/height)
Limitations:
- Depends on accurate anthropometric measurements
- Growth chart reference populations may not represent all ethnic groups
- Does not replace clinical judgment (constitutional growth delay vs. pathology)
Automated detection of implausible measurements is a low-risk data-quality application (Daymont et al., 2017). Whether and how it should be deployed depends on the EHR, measurement workflow, alert burden, and response pathway.
Developmental Screening AI:
Traditional screening: AAP recommends standardized developmental screening at 9, 18, 30 months (Lipkin et al., 2020) - Ages and Stages Questionnaires (ASQ) - Parents’ Evaluation of Developmental Status (PEDS) - Modified Checklist for Autism in Toddlers (M-CHAT)
AI-enhanced tools:
- Automated analysis of screening questionnaires
- Video analysis of infant motor development
- Speech/language delay detection from parent-recorded videos
Evidence:
- Cognoa Canvas Dx (AI-based autism diagnosis aid) (FDA De Novo DEN200069):
- Analyzes parent questionnaires + home videos via mobile app
- Aids diagnosis of ASD in patients 18 through 72 months who are at risk for developmental delay
- In the 425-subject analysis population, the device returned a positive or negative result for 135 children and returned no result for 290 (68%). Among children with a diagnostic output, positive predictive value was 81%, negative predictive value 98%, sensitivity 98%, and specificity 79% (FDA DEN200069)
- FDA De Novo authorization was granted in June 2021; this was authorization of a prescription diagnosis aid, not an autonomous screening test
- Earlier validation: Kanne et al., 2018 in Autism Research
Limitations:
- Cannot replace clinical diagnosis by developmental pediatrician
- Cultural and linguistic bias in screening tools (Zuckerman et al., 2014)
- Video quality and parent compliance variable
- Overdiagnosis risk (low PPV in low-prevalence populations)
- Delays in accessing diagnostic services after positive screen
Ethical concerns:
- Stigma of early autism labeling
- Parental anxiety from false positives
- Access to diagnostic services after positive screen variable (6-12 month waits common)
- Insurance discrimination concerns
The Cognoa ASD Diagnosis Aid received FDA De Novo authorization under DEN200069 in June 2021. Its successor, Canvas Dx, is indicated as a prescription aid for ASD diagnosis in patients 18 through 72 months who are already at risk for developmental delay. It is not a screening tool or stand-alone diagnosis (FDA K243558). A diagnostic output must be interpreted with history, observation, and other clinical evidence. A no-result output, which was common in the pivotal De Novo study, requires an alternative evaluation pathway.
7. Pediatric Emergency Department AI
Pediatric Sepsis Early Warning Systems:
Challenge: Pediatric sepsis is associated with substantial morbidity and mortality, while early recognition remains difficult because signs can be nonspecific (Weiss et al., 2020).
Traditional tools: Pediatric Early Warning Scores (PEWS), pediatric SIRS criteria
AI-enhanced systems:
- Continuous monitoring of vitals, labs, clinical documentation
- Age-adjusted warning criteria (pediatric SIRS not sensitive (Goldstein et al., 2005))
Evidence:
- Bedside PEWS severity score (Parshuram et al., 2018):
- EPOCH cluster randomized trial across 21 hospitals in 7 countries, covering 144,539 patient discharges
- All-cause hospital mortality did not differ (1.93 vs 1.56 per 1000 discharges; adjusted odds ratio 1.01, P=.96), and the authors concluded the findings do not support using the system to reduce mortality
- Significant clinical deterioration events were lower with BedsidePEWS (0.50 vs 0.84 per 1000 patient-days; adjusted rate ratio 0.77, P=.03)
- BedsidePEWS is a manually scored severity tool, not a machine learning system
- Published in JAMA (Parshuram et al., 2018)
- Adult sepsis evidence does not establish pediatric performance (Giannini et al., 2019):
- A random forest model deployed for adult non-ICU admissions at a Philadelphia health system reached sensitivity 26%, specificity 98%, and positive predictive value 29%
- Alerts produced a small increase in lactate testing and intravenous fluids but no significant difference in mortality, discharge disposition, or ICU transfer
- This model was neither developed nor validated in children
- Published in Critical Care Medicine (Giannini et al., 2019)
- Pediatric-specific sepsis ML models (Masino et al., 2019):
- Trained on NICU EHR data at a single children’s hospital
- Predict sepsis at least 4 hours before clinical recognition
- AUC 0.80-0.82 for culture-positive sepsis, and 0.85-0.87 when clinically positive evaluations were included
- Published in PLoS ONE (Masino et al., 2019)
Critical limitation:
- Most sepsis AI trained on adults, inadequate pediatric validation
- Age-appropriate vital sign thresholds essential
- Parental recognition of illness often precedes algorithmic detection
- Alert fatigue major implementation challenge
Pediatric Early Warning Scores and pediatric machine-learning studies answer different questions (Parshuram et al., 2018; Masino et al., 2019). Adult sepsis-model performance cannot be assumed to transfer to children. Pediatric deployment requires age-specific development or validation, prospective workflow evidence, and locally measured alert consequences.
Fracture Detection AI in Pediatric Imaging:
Application: AI analysis of pediatric radiographs for fracture detection
Evidence:
Rayan and colleagues developed a pediatric elbow-fracture model from 21,456 radiographs and evaluated it on an independent set; this research result should not be transferred to unrelated commercial products (Rayan et al., 2019)
Potential uses include triage and second-reader support in busy emergency departments
Whether a system reduces missed subtle fractures must be measured prospectively in the intended workflow
Published in Radiology: AI (Rayan et al., 2019)
Wrist fracture detection (Kim & MacKinnon, 2018):
- Transfer learning from a network pre-trained on non-medical images, applied to lateral wrist radiographs
- AUC 0.954 on a 100-image test set; at the cut-off maximising both, sensitivity was 90% and specificity 88%
- A proof-of-concept classification study; effects on missed injuries were not measured
- Published in Clinical Radiology (Kim & MacKinnon, 2018)
Limitations:
- Growth plates mimic fracture lines (AI false positives)
- Child abuse screening requires clinical correlation (algorithmic detection insufficient)
- Does not replace radiologist interpretation
- Performance varies by fracture location and subtlety
Clinical and medicolegal considerations:
- Missed fractures in child abuse cases have severe consequences
- AI should assist, not replace, careful skeletal survey interpretation
- Documentation should accurately record the clinician’s interpretation, relevant discordance, and resulting management; AI use does not itself create liability protection
These studies support continued evaluation of pediatric fracture triage and second-reader workflows (Rayan et al., 2019; Kim & MacKinnon, 2018). They do not establish class-wide commercial performance or improved patient outcomes. Skeletal surveys and suspected abuse require specialist interpretation and clinical correlation.
8. Pediatric Chronic Disease Management
Type 1 Diabetes AI Applications:
Artificial Pancreas Systems (Hybrid Closed-Loop Insulin Delivery):
Systems:
- Commercial automated insulin-delivery systems include Medtronic, Tandem, Insulet, and Beta Bionics products. Age indications and compatible components change with product-specific labeling and should be verified in the current FDA record.
- FDA cleared the insulin-only iLet ACE Pump and iLet Dosing Decision Software in 2023 for people aged 6 years and older with type 1 diabetes. The system is initialized with body weight and replaces conventional carbohydrate counting with meal announcements; it does not eliminate meal input (FDA, 2023).
- Automated insulin delivery is an adaptive control application. It should not be counted as machine-learning evidence unless the product’s technical record establishes that category.
Evidence:
- Pediatric RCTs (Breton et al., 2020):
- Time-in-range improved from 53% to 67% with Control-IQ versus 51% to 55% with a sensor-augmented pump (adjusted difference 11 percentage points)
- Overnight time-in-range was 80% with closed-loop control versus 54% with the sensor-augmented pump
- Hypoglycemia was low in both groups (time below 70 mg/dL 1.6% vs 1.8%), and the HbA1c difference of -0.4 percentage points (95% CI -0.9 to 0.1, P=0.08) did not reach significance
- Published in NEJM (Breton et al., 2020)
- Real-world outcomes (Pinsker et al., 2020):
- Similar benefits in routine clinical use
- Quality of life improvements for children and parents (Cobry et al., 2021)
- Reduced diabetes distress and parental fear of hypoglycemia
- Published in Diabetes Technology & Therapeutics (Pinsker et al., 2020)
- Very young children (Wadwa et al., 2023):
- PEDAP randomized 102 children aged 2 to under 6 years to closed-loop control or standard care over 13 weeks
- Time-in-range rose from 56.7% to 69.3% with closed-loop control versus 54.9% to 55.9% with standard care (adjusted difference 12.4 percentage points, 95% CI 9.5-15.3)
- Published in the New England Journal of Medicine (Wadwa et al., 2023)
Limitations:
- Requires continuous glucose monitor (CGM) + insulin pump (technology burden)
- User and caregiver training is essential, with time requirements varying by system and clinical program
- Patient costs and insurance coverage vary by device, components, plan, and jurisdiction
- Alarm burden varies by device configuration, glucose patterns, and user settings
- Does not eliminate need for carbohydrate counting and diabetes self-management
- System failures require backup conventional insulin regimen
Equity concerns:
- Access limited by insurance, SES, health literacy
- Disparities in technology use by race/ethnicity (Agarwal et al., 2021)
- Published in Diabetes Technology & Therapeutics (Agarwal et al., 2021)
Randomized evidence supports improved glycemic time in range for specific pediatric automated insulin-delivery systems (Breton et al., 2020; Wadwa et al., 2023). The result is product- and population-specific. Training burden, device wear, alarm tolerance, backup insulin planning, and affordability remain part of shared decision-making. Access disparities require measurement rather than an assumption that authorization produces equitable uptake.
AI-Enhanced Asthma Management:
Applications:
- Inhaler adherence monitoring (smart inhalers with Bluetooth)
- Exacerbation prediction from symptom tracking apps
- Environmental trigger identification (pollen, air quality, allergens)
Evidence:
- Smart inhalers (Chan et al., 2015):
- Median adherence to inhaled corticosteroids was 84% with audiovisual reminders versus 30% without, in a randomised trial of 220 children aged 6-15 years
- Audiovisual dose reminders rather than inhaler technique feedback; days absent from school did not differ between groups
- Published in Lancet Respiratory Medicine (Chan et al., 2015)
- Exacerbation prediction (Finkelstein & Jeong, 2017):
- ML models predict asthma exacerbations 3-7 days in advance
- Accuracy modest (AUC 0.70-0.75)
- Published in Ann NY Acad Sci (Finkelstein & Jeong, 2017)
- Pediatric-specific validation limited:
- Most studies in adults
- Adherence improvement not consistently translated to outcome improvement (ED visits, hospitalizations)
Smart inhalers (Chan et al., 2015) are promising tools for motivated families struggling with medication adherence. But we need pediatric-specific RCTs showing they actually reduce exacerbations and hospitalizations, not just improve adherence metrics.
9. Pediatric Oncology AI
Pediatric Cancer Diagnosis and Risk Stratification:
Applications:
- Neuroblastoma risk stratification from genomics
- Leukemia subtype classification from blast morphology
- Brain tumor segmentation and classification from MRI
Evidence:
- Neuroblastoma genomic classifiers (Cohn et al., 2009):
- Integrate genomic data to refine risk stratification
- Improve prediction of treatment response
- Published in Journal of Clinical Oncology (Cohn et al., 2009)
- ALL subtype classification (Arber et al., 2016):
- The 2016 WHO revision defines leukemia subtypes by combined morphologic, immunophenotypic, and genetic criteria, the reference standard any classifier must reproduce
- This is a consensus classification document; it does not evaluate an AI classifier
- Published in Blood (Arber et al., 2016)
- Pediatric brain tumor classification (Tampu et al., 2025):
- MRI-based deep learning models classify tumor types
- Reported performance varies across tumor type, imaging sequence, reference standard, and external validation setting; the cited review does not establish a universal 80%–90% comparison with pathologists
- Published in Neuro-Oncology Advances (Tampu et al., 2025)
Critical limitations:
- Pediatric cancer rare (limited training data)
- Genomic classifiers expensive, not universally available
- Clinical validation in prospective pediatric trials lacking
- Most studies retrospective, single-institution
- Integration with established risk stratification systems (COG protocols) incomplete
Ethical concerns:
- Prognostic predictions influence treatment intensity decisions (more vs. less chemotherapy)
- False reassurance (underestimating risk) or false alarm (overestimating risk) both problematic
- Family involvement in research consent complex (parental permission + child assent)
This research supports careful external validation (Cohn et al., 2009; Tampu et al., 2025), but it does not establish routine clinical utility. Pediatric cancers are rare, datasets are limited, and many studies are retrospective or single-institution. Treatment-changing use requires multi-institutional prospective evaluation and compatibility with established Children’s Oncology Group protocols.
10. Pediatric Mental and Behavioral Health AI
Suicide Risk Prediction:
Application: ML models analyzing EHR data to identify children/adolescents at high suicide risk
Evidence:
- Suicide attempt prediction (Walsh et al., 2017):
- AUC 0.84 at 7 days before an attempt, falling to 0.80 at 720 days, with recall of 0.95 across all prediction windows
- Reported precision was 0.74-0.79, but this came from a case-control sample; positive predictive value would be far lower at the true population base rate
- Published in Clinical Psychological Science (Walsh et al., 2017)
- Adolescent-specific models (Su et al., 2020):
- AUC 0.81-0.86 across prediction windows from 0 to 365 days, in 41,721 patients aged 10-18 years
- At 90% specificity the models detected 53-62% of patients who went on to attempt suicide
- Published in Translational Psychiatry (Su et al., 2020)
Implementation challenges:
- What to do with high-risk predictions? (Resource-intensive interventions)
- False positives cause family distress and labeling concerns
- True positives may not be preventable with current interventions
- Liability if identified patient not contacted and dies by suicide
Ethical concerns:
- Screening vs. surveillance (are we identifying risk to help or monitor?)
- Adolescent privacy and confidentiality (HIPAA allows parental access to minor records, but teens may not disclose SI if parents informed)
- Parental notification requirements (varies by state)
- Potential for discrimination (insurance, employment, education)
Discrimination in a retrospective case-control dataset does not establish population positive predictive value or clinical benefit (Walsh et al., 2017). Deployment requires a defined intervention pathway, locally estimated alert burden, age-appropriate privacy and family-communication policy, and prospective evaluation against validated clinical screening. A model should not be represented as endorsed by a professional society without a direct, current policy source.
ADHD Diagnosis Support:
Tools:
- AI analysis of continuous performance tests (CPTs)
- Classroom behavior observation algorithms
- Parent/teacher rating scale analysis
Evidence:
- Objective measures correlate with ADHD diagnosis but do not replace clinical assessment (Emser et al., 2018)
- Regulatory status must be checked for the exact product and intended use; no cited source in this section establishes an authorized autonomous ADHD diagnosis
- DSM-5 criteria remain gold standard (requires clinical judgment, developmental history, functional impairment assessment)
- Published in Behavioral and Brain Functions (Emser et al., 2018)
Limitations:
- ADHD heterogeneous (inattentive, hyperactive, combined types)
- Comorbidities common (anxiety, depression, learning disabilities)
- Cultural and contextual factors influence symptom expression
- No biomarker or objective test diagnostic
AI tools analyzing continuous performance tests or behavior ratings may support clinical assessment (Emser et al., 2018), but they do not replace comprehensive ADHD evaluation. Developmental history, school performance, family assessment, impairment across settings, and comorbidity screening remain necessary.
Equity and Bias Concerns in Pediatric AI:
Training Data Bias:
- Most medical AI trained on adult populations
- Pediatric data scarce, often from academic medical centers
- Underrepresentation of minority children, rural children, low-income children (Rajkomar et al., 2018)
Examples of Documented Bias:
1. Pulse Oximetry and Skin Pigmentation:
- Sjoding and colleagues analyzed paired measurements from hospitalized adults, not children, and found occult hypoxemia more often in Black than White patients (Sjoding et al., 2020)
- The adult study establishes a sensor-equity concern but does not quantify pediatric error or prove a fixed 2%–3% overestimate in children
- Any pediatric model that uses pulse oximetry can inherit measurement error from its input, so pediatric validation should include paired arterial measurements and performance by a justified skin-pigmentation measure
2. Neonatal Sepsis Calculators:
- Validation studies predominantly white populations
- Performance in diverse populations uncertain
- Social determinants of health not incorporated (maternal prenatal care access, housing stability)
3. Developmental Screening Tools:
- Cultural and linguistic bias in questionnaires (Zuckerman et al., 2014)
- Video analysis trained on majority populations
- Autism screening tools show racial disparities in referral (Constantino et al., 2020)
- Black and Hispanic children diagnosed later, at higher severity (Constantino et al., 2020)
- Published in Pediatrics (Constantino et al., 2020)
4. Growth References:
- WHO standards and CDC growth references differ in source populations, construction, and intended interpretation
- A growth reference does not diagnose disease; trajectory, measurement quality, feeding, gestational age, and clinical context matter
- Breastfeeding and formula-feeding growth trajectories can differ (Dewey et al., 1992)
5. Asthma Prediction Models:
- Many trained on insured, suburban populations
- Underperform in urban, low-income settings
- Miss environmental triggers specific to disadvantaged neighborhoods (mold, pests, pollution)
Consequences:
- Delayed diagnosis in minority children
- Overdiagnosis or underdiagnosis based on race/ethnicity
- Widening of existing health disparities (Obermeyer et al., 2019)
- Erosion of trust in pediatric care systems among minority families
Mitigation Strategies:
- Require diverse pediatric training datasets (by race, ethnicity, SES, geography)
- Validate algorithms across demographic subgroups
- Report performance stratified by demographics (mandate transparency)
- Engage community stakeholders in AI development
- Continuous monitoring for bias after deployment
- Independent equity audits before and after implementation
Ethical Frameworks for Pediatric AI: The proposed ACCEPT-AI framework addresses age, communication, consent and assent, equity, data protection, and technical transparency in research using pediatric data. It recommends checking for age-related bias during dataset curation, training, and testing, including after deployment. It was published as a proposal without an E-Delphi consensus process, not as a clinical practice guideline (Muralidharan et al., 2023).
1. Best Interest Standard:
- AI must serve child’s best interest, not just efficiency or cost reduction
- Long-term consequences matter (children have decades ahead)
- Parents and children should participate in AI deployment decisions
- A 2025 Pediatrics perspective offers a trustworthy-pediatric-AI framework aligned with National Academy of Medicine principles; it is not a formal AAP clinical practice guideline (Johnson et al., 2025)
2. Permission, Assent, and Communication:
- Routine clinical use of software does not automatically create a separate legal consent requirement; requirements depend on jurisdiction, intervention, research status, institutional policy, and material risk
- Parents or guardians generally authorize pediatric care, while developmentally appropriate assent and adolescent involvement remain ethically important
- When a meaningful non-AI alternative exists, families should understand the role of the system and available choices
- Explanations should be understandable to parents and, when appropriate, children
3. Privacy and Confidentiality:
- Children’s health data requires special protection (Toh et al., 2019)
- Longitudinal records follow children into adulthood
- Data sharing for AI training must have strict safeguards
- Adolescent confidentiality particularly sensitive (reproductive health, mental health, substance use)
- COPPA applies to covered online services directed to children under 13, or services with actual knowledge that they collect personal information from such children; it is not a universal medical-device privacy law
4. Equity and Justice:
- AI must not worsen existing disparities in pediatric care (Obermeyer et al., 2019)
- Access to beneficial AI should not depend on insurance status
- Validation in diverse populations mandatory before deployment
- Attention to digital divide (not all families have smartphones, reliable internet)
5. Avoid Premature Deployment:
- Higher bar for pediatric AI evidence than adult AI
- Vulnerable population justifies extra caution (precautionary principle)
- Pilot studies in pediatric populations essential before broad deployment
- Long-term safety monitoring required
6. Transparency:
- Families should know when AI influences their child’s care
- Explainable AI particularly important for parental trust
- Physicians must be able to explain AI recommendations in plain language
- Black-box algorithms ethically problematic in pediatrics
Clinical Practice Guidelines for Pediatric AI:
This is an authorial adoption framework informed by the evidence and ethical sources cited in this chapter. It is not an American Academy of Pediatrics clinical practice guideline.
Before Adopting Pediatric AI:
- Demand pediatric-specific validation:
- Adult validation insufficient
- Stratify performance by age groups (<1 year, 1-5, 6-12, 13-18)
- Include diverse populations (race, ethnicity, SES, geography)
- Published prospective studies, not just retrospective accuracy (Taylor et al., 2019)
- Assess benefit-risk for children:
- Does this improve outcomes or just efficiency?
- What are failure modes and consequences?
- Are there safer alternatives?
- Is the benefit worth the risk? (especially for vulnerable neonates)
- Evaluate equity implications:
- Will this widen or narrow disparities?
- Is training data representative?
- Can all families access this technology? (SES, insurance, language, health literacy)
- Published equity analysis required
- Consider family preferences:
- Some families prefer human-only care (religious, cultural, personal reasons)
- Cultural attitudes toward technology vary
- Offer alternatives when possible
- Respect parental autonomy
- Ensure child-appropriate interfaces:
- Language and visuals appropriate for developmental stage
- Avoid frightening or confusing children
- Involve child life specialists in design
- Gamification should not trivialize medical care
Safe Implementation:
- Staged rollout: Begin in the population and workflow supported by the strongest local and external evidence; do not extrapolate across age groups without validation
- Risk-proportionate monitoring: Set review frequency from severity, exposure, detectability, deployment scale, and expected drift rather than a universal monthly or quarterly schedule
- Incident reporting: Capture adverse events and near misses internally, and report externally when applicable law, regulation, authorization, or institutional policy requires it
- Family feedback: Systematically collect parent and adolescent experiences
- Physician oversight: AI should support, not replace, pediatrician judgment
- Continuous validation: Monitor real-world performance across demographic subgroups
Red Flags (Avoid These Systems):
- No pediatric validation (only adult data)
- Claims to diagnose complex conditions autonomously (autism, ADHD, mental health)
- Lack of age-stratified performance data
- No mechanism for parents to review AI inputs/outputs
- Vendor resistance to equity audits
- Black-box models without explanation capability
- No FDA clearance when clearance required
Future Directions in Pediatric AI:
Demonstrated or Clinically Deployed:
- Multivariate EOS risk assessment, conventional HRV monitoring, closed-loop oxygen control, and automated insulin delivery have clinical evidence, but represent different computational categories
- AI-assisted ROP and fracture interpretation has retrospective or reader-study evidence, with transportability and workflow effects varying by product and setting
- Pediatric diagnosis aids such as Canvas Dx have narrow product-specific FDA indications
Theoretical or Under Evaluation:
- AI-assisted developmental assessment integrated into well-child workflows
- Rare-disease support combining clinical and genomic data
- Mental-health risk models with an actionable, resource-matched intervention pathway
- Wearable monitoring for selected chronic conditions
Beyond Current Evidence:
- Autonomous replacement of pediatric primary care
- Fully automated diagnosis of complex developmental or behavioral conditions
- School-based surveillance without settled evidence, privacy, governance, and consent structures
- A single pediatric model that transfers safely across developmental stages without age-specific validation
Key Research Gaps:
Validation Studies:
- Prospective RCTs of AI interventions in children
- Multi-site validation across diverse populations
- Long-term outcome studies (does AI improve health trajectories to adulthood?)
- Cost-effectiveness analyses from healthcare system and family perspectives
Equity Research:
- Performance of AI across racial/ethnic groups (stratified reporting mandatory)
- Impact on health disparities (helpful or harmful?)
- Access barriers to beneficial AI technologies
- Community-based participatory research in AI development
Implementation Science:
- Best practices for integrating AI into pediatric workflows
- Training needs for pediatricians, pediatric nurses, pediatric specialists
- Family acceptance and preferences across cultures
- Strategies to minimize alert fatigue in pediatric settings
Weiser and colleagues evaluated eight large language models on 304 board-style pediatric cardiology multiple-choice questions, with and without single-textbook retrieval-augmented generation (NadasGPT); the top models reached 98.4% accuracy without RAG, and single-textbook RAG did not significantly change performance for any model except gpt-5.1 on fellow-level questions (Weiser et al., 2026). MCQ accuracy is not educational effectiveness and is not bedside clinical decision support.
Ethics Research:
- How to obtain meaningful consent/assent for AI use (developmental stage considerations)
- When is AI use in children justified? (ethical frameworks)
- Balancing innovation with precautionary principle
- Long-term consequences of childhood health data collection
Safety Research:
- Adverse event surveillance for pediatric AI
- Failure mode analysis specific to children
- Human factors research (how do pediatricians interact with AI?)
Conclusion
Pediatric computational interventions span distinct evidence categories. The HeRO trial evaluated an HRV display plus clinician response (Moorman et al., 2011). ROP studies evaluated image classification and reader support, not prevented blindness (Brown et al., 2018). Pediatric automated insulin-delivery trials measured glycemic outcomes for specific systems (Breton et al., 2020). Children’s unique vulnerabilities demand age-specific evidence, equity evaluation, and careful consideration of long-term consequences.
Pediatricians should embrace AI tools with robust evidence while advocating for children in AI development, demanding diverse representation in training data, and insisting on pediatric-specific validation before deployment.
The principle remains constant: First, do no harm, especially to children who cannot fully advocate for themselves.
Artificial intelligence must serve the best interests of all children, not just those well-represented in training datasets. A 2025 Pediatrics perspective calls for trustworthy pediatric AI aligned with National Academy of Medicine principles (Johnson et al., 2025).
Check Your Understanding
The three cases below are explicitly hypothetical teaching exercises. Institutions, products, patient counts, model outputs, costs, outcomes, and legal arguments are illustrative unless a source is linked in the same sentence. They are not reports of actual events, regulatory findings, or legal holdings. The purpose is to practice base-rate reasoning, subgroup auditing, and safe escalation without presenting invented details as fact.
Scenario 1: Hypothetical Suicide-Risk Model in Adolescent Medicine
An academic children’s hospital is considering a fictional suicide-risk prediction model for adolescent emergency and inpatient encounters.
AI system: Analyzes EHR data (diagnoses, medications, prior visits, social history) and flags patients at high risk for suicide attempt within 30 days.
Illustrative first-month performance (adolescent patients aged 12–17):
- Patients flagged as high-risk: 487 out of 1,200 adolescent encounters (41%)
- Actual suicide attempts within 30 days: 3 patients
- True positives: AI correctly identified 3/3 (100% sensitivity)
- False positives: 484 patients flagged but no suicide attempt
- Positive predictive value: 0.6% (3/487)
Clinical workflow impact:
- All flagged patients require: Psychiatric consult, safety plan, social work assessment, close follow-up
- Psychiatric service overwhelmed: 487 consults vs. usual 120/month
- Wait time for psych consult: 8 hours → 24+ hours
- Parents of flagged children upset: “Why does AI think my child will hurt themselves?”
Illustrative week 3 event:
- 15-year-old girl presents to ED with asthma exacerbation
- AI flags as high suicide risk (prior depression diagnosis 2 years ago, now in remission)
- Psychiatric consult delayed 26 hours due to backlog
- During wait, girl becomes agitated, family frustrated
- Girl not suicidal, discharged home after 30-hour ED stay
- Family files complaint about unnecessary psychiatric hold
Answer 1: What is the problem with this AI implementation?
Unacceptably low positive predictive value (0.6%):
- 99.4% of flagged patients are false positives
- For every 1 true suicidal patient, AI flags 161 non-suicidal patients
Why such low PPV?:
- Low base rate: Suicide attempts rare (3/1,200 = 0.25% prevalence)
- High sensitivity optimization: AI designed to catch all true cases (100% sensitivity)
- Result: At 0.25% prevalence, even 90% specificity yields PPV <2%
System overload:
- Psych service cannot handle 4× increase in consult volume
- Delays in care for truly high-risk patients
- Resource diversion from patients who need help
Clinical harm:
- False positive patients subjected to unnecessary psychiatric evaluation
- Stigma, family distress, prolonged ED stays
- Delayed care for asthma (primary presenting complaint)
Answer 2: Why should this not be presented as a real Vanderbilt deployment?
The numbers in this exercise are fictional. They should not be attributed to Vanderbilt or another institution without a connected source establishing the cohort, model, endpoint, and deployment result.
Transferable failure mode:
- Rare outcomes constrain positive predictive value even when discrimination appears strong
- A threshold that maximizes sensitivity can create an unmanageable number of false-positive alerts
- Case-control precision cannot be transferred to a real population without accounting for prevalence
- Deployment claims require prospective evidence from the actual workflow
The hypothetical pediatric implementation:
- Same problem: Optimized for sensitivity at cost of specificity
- Same PPV (~0.6%)
- Same clinical consequence: System overwhelm, false positive burden
Why low PPV is worse in pediatrics:
- Parents involved in all decisions (family distress amplified)
- Psychiatric resources more limited in pediatrics
- Stigma potentially greater for children/adolescents
- Longer-term implications of psychiatric labeling in childhood
Answer 3: What are the liability implications?
Potential risk questions for the hospital:
If AI misses a suicide (false negative):
- Plaintiff argument: Hospital deployed suicide prediction tool but failed to flag patient → negligence
- Defense: AI had 100% sensitivity in pilot; this case was unpredictable
No liability conclusion follows from this hypothetical sensitivity estimate. Duty, causation, damages, and the standard of care are fact- and jurisdiction-specific.
If false positive causes harm:
- Plaintiff argument (asthma patient delayed care):
- AI incorrectly flagged patient as suicidal
- Triggered unnecessary 26-hour psychiatric hold
- Delayed asthma treatment
- Family trauma, stigma
- Defense: Hospital acting in abundance of caution, suicide risk assessment standard of care
- Unresolved question: A court or regulator would examine the actual facts, applicable law, necessity of the hold, and reasonableness of the workflow
Governance concern: Failure to monitor and adjust the system:
- Hospital deploys system with 99.4% false positive rate
- Continues deployment despite evidence of harm (psych service overload, care delays)
- Does not recalibrate or pause system
- Plaintiff: Hospital knew system was causing harm but continued anyway
Answer 4: How should suicide risk AI be implemented in pediatrics?
Pre-implementation considerations:
- Model the consequences of a low base rate
- At an illustrative prevalence of 0.25%, positive predictive value depends jointly on sensitivity, specificity, and the deployed threshold
- The organization should calculate expected alerts, true positives, false positives, and missed cases before deployment
- Resource capacity check
- Can psychiatric service handle 2-4× increase in consults?
- If no, system will fail (alert fatigue, delays, staff burnout)
- Tiered risk stratification:
- Define response tiers from local validation, clinical urgency, staffing, and the harms of false positives and false negatives
- Do not import fixed positive-predictive-value thresholds from a hypothetical case
- Reserve intensive interventions for a tier with a clinically justified response pathway
Implementation safeguards:
- Clinical override:
- Pediatrician reviews AI flag, determines if psychiatric consult truly needed
- AI is screening tool, not mandate
- Family communication:
- Explain AI screening to families: “Routine screening tool flagged some concerns. We’ll ask some questions to better assess.”
- Avoid alarming language: “AI thinks your child is suicidal”
- Continuous monitoring:
- Track positive predictive value and alert volume at a risk-proportionate interval
- Use prespecified safety and capacity triggers for pausing, recalibrating, or retiring the system
- Alternative approach: Universal brief screening
- Instead of AI, use validated brief tools (ASQ, C-SSRS) for all adolescents
- More clinically actionable, less false positive burden
Documentation:
- “AI suicide risk algorithm flagged patient. Clinical assessment: No current suicidal ideation, intent, or plan. Patient cooperative, family supportive. Low acute risk. Outpatient f/u arranged.”
Lesson: Suicide-risk prediction in pediatrics faces a fundamental base-rate challenge. A model can discriminate in retrospective data yet create an unsustainable false-positive burden in practice. Implementation requires a tiered response, clinical oversight, capacity modeling, privacy safeguards, and prospective monitoring. Validated brief clinical screening remains an important comparator.
Scenario 2: Hypothetical ROP Subgroup-Performance Audit
A Level III NICU is evaluating a fictional AI-assisted ROP screening system. No named commercial or research product should be inferred from the illustrative data below.
AI system: Analyzes retinal images, classifies as:
- Plus disease present → Urgent ophthalmology referral
- Pre-plus disease → Close monitoring
- No plus disease → Routine screening
Your NICU demographics:
- 60% Black/Hispanic infants
- 30% White infants
- 10% Asian infants
- Serves predominantly low-income community
Illustrative month 3 performance review:
| Race/Ethnicity | AI Sensitivity (Plus Disease) | AI Specificity | Ophthalmology-Confirmed Plus Disease |
|---|---|---|---|
| White infants | 95% (19/20 cases detected) | 88% | 20 cases |
| Black infants | 78% (25/32 cases detected) | 85% | 32 cases |
| Hispanic infants | 72% (18/25 cases detected) | 83% | 25 cases |
Illustrative missed cases:
- 7 Black infants with plus disease misclassified as “no disease” by AI
- 7 Hispanic infants with plus disease misclassified as “no disease” by AI
- All received delayed treatment (2-3 weeks later than optimal)
- 3 infants progressed to Stage 3+ ROP requiring laser therapy (might have been prevented with earlier treatment)
Answer 1: What could explain the apparent subgroup disparity?
The table alone does not identify a cause. Race and ethnicity are social categories, not image mechanisms. The review should test several explanations before attributing the difference to biology, pigmentation, or training composition:
- Sampling and uncertainty:
- The subgroup numerators are small and the table omits confidence intervals
- Referral patterns, disease prevalence, and missing follow-up may differ across groups
- Image acquisition and site effects:
- Camera type, operator, dilation, illumination, focus, compression, and ungradable-image handling can alter performance
- These operational variables may be correlated with site, insurance, language, or recorded race and ethnicity
- Training and labeling representation:
- Training data may underrepresent the deployment population, cameras, disease spectrum, or clinical sites
- Reference labels may vary across experts, and subgroup label quality should be audited
- A justified pigmentation measure can be analyzed separately from race and ethnicity when image contrast is a plausible concern
Similar to dermatology AI bias:
- Image-based evaluations should use a justified skin-tone measurement approach and report subgroup performance, uncertainty, image capture conditions, and external validation. Fitzpatrick type and race labels alone are insufficient. See Dermatology for the specialty-specific evidence and adoption framework.
Answer 2: What are the liability implications of using biased AI?
Hospital and clinician risk considerations for missed ROP cases:
Plaintiff argument (parents of Black infant with missed ROP):
- Hospital deployed AI system that performed worse on Black infants (78% vs. 95% sensitivity)
- This is discriminatory medicine, a different standard of care based on race
- Our child was harmed (delayed treatment, worse outcome) because of racial bias in AI
- Hospital knew or should have known AI had lower sensitivity for Black/Hispanic infants
Legal and civil-rights context:
- Federal and state civil-rights obligations can apply to healthcare programs and activities
- A subgroup disparity warrants immediate investigation and mitigation, but this hypothetical table does not establish a legal violation
- Legal exposure depends on facts including the applicable statute, intent or effect standard, knowledge, causation, safeguards, and resulting harm
Plaintiff damages:
- Child now requires laser therapy (might have been prevented with earlier detection)
- Potential vision impairment
- Lifelong consequences of preventable ROP progression
Defense arguments:
- AI overall performance acceptable (average 82% sensitivity)
- Physician still reviewed images (AI was decision support, not final decision)
- ROP difficult to detect even for human experts
Factors a reviewer would examine:
- A claim could be strengthened if:
- Hospital aware of racial performance disparity but continued deployment
- No additional safeguards for higher-risk populations
- Multiple missed cases demonstrating pattern
Regulatory implications:
- Product submissions and postmarket obligations are device- and authorization-specific
- Subgroup performance, training-data representativeness, and clinically relevant failure modes should be reviewed in the exact FDA record
- Local equity monitoring complements, but does not replace, manufacturer and regulatory evidence
Answer 3: How should ROP screening AI be implemented equitably?
Pre-deployment validation:
- Stratified performance testing:
- Test AI on representative sample of your NICU population
- Report sensitivity/specificity by race/ethnicity
- Prespecify uncertainty-aware safety triggers for pausing or narrowing deployment; no universal 10-percentage-point rule applies
- Vendor accountability:
- Demand race-stratified validation data from vendor
- Ask: “What is sensitivity for Black, Hispanic, Asian infants specifically?”
- If evidence is inadequate for the intended population, restrict or defer deployment until the uncertainty is resolved
- Match evidence to the deployment population:
- A model validated in a narrow population, camera, or site requires local and external evidence before broader use
Implementation safeguards:
- Safety-net discordant or uncertain cases:
- Use a clinically justified pathway for ungradable images, low-confidence results, and disagreement with examination findings
- Do not assign different care solely from race or ethnicity without evidence that the policy improves safety and equity
- Hybrid approach:
- AI can triage images while an ophthalmologist reviews cases defined by clinical risk, image quality, uncertainty, or discordance
- The review policy should be evaluated for both missed disease and unequal burden
- Human expert involvement:
- Neonatologist or ophthalmologist reviews AI classifications
- Clinical override when AI classification conflicts with exam findings
Monitoring:
- Ongoing surveillance:
- Track missed ROP cases by race
- If pattern emerges (more misses in Black/Hispanic infants) → investigate AI bias
- Risk-proportionate audits:
- Compare AI performance to ophthalmology gold standard by race
- Report to hospital equity committee
Advocacy:
- Demand better AI:
- Work with AI vendors to improve performance on diverse populations
- Insist on diverse training datasets
- Research funding:
- Support research creating diverse ROP datasets
- Partner with vendors to improve AI for underrepresented populations
Documentation:
- “AI ROP screening classified as [result]. Reviewed images personally. [Agree/Disagree] with AI assessment. Ophthalmology referral [made/deferred] based on clinical judgment.”
Lesson: An observed subgroup performance gap is a safety signal, not proof of its cause. Require representative validation, uncertainty estimates, image-quality analysis, monitoring, human oversight, and a documented mitigation path. Do not turn race into a biological mechanism or a fixed referral rule without evidence.
Scenario 3: Hypothetical Neonatal Jaundice App in a Rural Setting
A rural health clinic is evaluating a fictional smartphone bilirubin-estimation app. The nearest children’s hospital is 120 miles away in this teaching scenario.
Clinical challenge: The clinic sees an illustrative 15–20 newborns per month. A transcutaneous bilirubinometer is unavailable, and serum bilirubin testing requires a fictional 24–48-hour reference-laboratory turnaround.
Fictional AI tool: A smartphone app called BiliPhoto estimates bilirubin from an infant photograph. - Illustrative marketing claims: “93% accuracy, replaces transcutaneous bilirubin measurement, FDA-cleared” - Illustrative cost comparison: $50 per month versus $15,000 for a transcutaneous bilirubinometer
The claims and prices are deliberately unverified examples of what a procurement team must investigate. No FDA authorization or commercial availability is implied.
Your usage:
- Use BiliPhoto for initial screening
- If BiliPhoto suggests bilirubin above an illustrative threshold of 15 mg/dL, send serum bilirubin and consider referral
Case 1 - Success:
- 4-day-old term infant, looks jaundiced
- BiliPhoto estimate: 17.2 mg/dL
- Serum bilirubin (sent immediately): 16.8 mg/dL
- Transfer to children’s hospital for phototherapy
- Good outcome, kernicterus prevented
Case 2 - Near miss:
- 3-day-old term infant, appears mildly jaundiced
- BiliPhoto estimate: 11.2 mg/dL (below the fictional workflow threshold)
- Parents reassured, discharged home
- Infant returns 2 days later (day 5) lethargic, poor feeding
- Emergency serum bilirubin: 24.8 mg/dL (critical)
- Emergency transfer, exchange transfusion required
- Infant survives but develops kernicterus (bilirubin encephalopathy)
Investigation: Why might BiliPhoto have underestimated? - Infant has darker skin tone (Hispanic) - Room lighting was fluorescent (not natural light as recommended) - The fictional vendor’s validation population and lighting conditions were not adequately verified before use
Answer 1: What was the error in using BiliPhoto?
Over-reliance on an unverified tool:
- Unverified performance across skin pigmentation: The clinic did not establish accuracy across its patient population
- Skin pigmentation, camera processing, and illumination are plausible effect modifiers that require measurement, not assumptions based on race or ethnicity
- Training and validation composition must be obtained from the manufacturer rather than inferred
- Lighting conditions: Smartphone camera + room lighting ≠ controlled medical device
- The fictional app’s acquisition protocol requires controlled lighting, which the clinic did not follow
- Color temperature, shadows, reflections affect accuracy
- Acquisition variation: Smartphone positioning, focus, distance, and device processing can affect a color-based estimate
- Not standardized like medical device
- No local validation: The clinic did not compare BiliPhoto with total serum bilirubin across the intended population and conditions of use
Even an FDA-authorized device must be used within its exact indication and labeling:
- Authorization does not establish performance in every population, camera, lighting condition, or workflow
- The regulatory pathway, intended use, contraindications, validation population, and uncertainty must be verified from the exact FDA record
Answer 2: What liability questions would the hypothetical event raise?
No categorical liability answer is possible. Key questions include:
Standard of care:
- What is standard approach for neonatal jaundice screening in resource-limited settings?
- What do the applicable guideline, available resources, device labeling, consultation options, and local standard require?
- Did the clinic rely on an unvalidated output instead of obtaining or arranging a clinically indicated bilirubin measurement?
Plaintiff argument:
- You used smartphone app instead of validated medical device (transcutaneous bilirubinometer)
- App underestimated bilirubin due to known limitations (skin tone, lighting)
- Failed to obtain confirmatory serum bilirubin despite moderately elevated visual jaundice
- Infant developed preventable kernicterus due to missed diagnosis
Defense argument:
- The fictional app was used according to the information available to the clinician
- The clinic documented resource constraints, arranged follow-up, and responded to the estimated risk
- The reasonableness of those steps would require expert and jurisdiction-specific analysis
Potential expert questions:
- Was the app authorized for this intended use, population, device, and lighting condition?
- Did the clinician follow the AAP hyperbilirubinemia guideline, which bases treatment thresholds on age in hours, gestational age, total serum bilirubin, and neurotoxicity risk factors rather than a fixed app threshold?
- Was follow-up timely and feasible given the child’s risk and distance from definitive care?
Outcome boundary:
- The result would depend on jurisdiction, facts, causation, expert testimony, documentation, and the clinical options that were realistically available
- This teaching case does not predict a verdict or settlement
Answer 3: How should smartphone-based medical AI be used safely in resource-limited settings?
Validation before deployment:
- Local validation study:
- Compare BiliPhoto estimates with total serum bilirubin under a prespecified protocol
- Use a sample size justified by the endpoint, required precision, and important subgroups rather than an arbitrary 50–100-patient rule
- Measure skin pigmentation with a justified method and record camera, lighting, operator, age in hours, gestational age, and relevant risk factors
- Determine calibration, error distribution, and clinically important misclassification in the intended setting
- Lighting standardization:
- Take all photos in same location (near window, natural light)
- Avoid fluorescent/LED lighting
- Use white background, standard distance
- Operator training:
- All staff using BiliPhoto trained on proper technique
- Inter-rater reliability testing
Clinical integration:
- Use only within verified evidence and labeling:
- A photographic estimate cannot be assumed to replace transcutaneous or total serum bilirubin measurement
- Confirmatory testing and follow-up should follow age-in-hours, gestational-age, and risk-specific guidance rather than the fictional fixed thresholds in this scenario (Kemper et al., 2022)
- Do not rely on AI when clinical exam conflicts:
- If visual assessment, feeding history, examination, or risk factors conflict with a reassuring app output, obtain an appropriate bilirubin measurement or urgent consultation
- Clinical judgment includes recognizing that visual assessment alone is also insufficient for precise bilirubin quantification
- Risk-specific escalation:
- Gestational age, age in hours, feeding, weight trajectory, hemolysis risk, G6PD deficiency, sepsis, and clinical instability should inform measurement and escalation under applicable guidance
- Race or ethnicity should not substitute for a measured clinical risk factor
Documentation:
- “BiliPhoto estimate: 11.2 mg/dL in this hypothetical case. The estimate conflicts with clinical concern and has not been validated for this patient or acquisition condition. Total serum bilirubin and risk-specific follow-up arranged.”
Advocacy for access:
- Telemedicine: Photo sent to neonatologist for assessment
- Point-of-care serum bilirubin: Advocate for access to rapid testing
- Equipment grants: Apply for funding for transcutaneous bilirubinometer
Lesson: Smartphone-based estimation may improve access, but color-based measurement can be sensitive to acquisition conditions, device processing, and population shift. A product must have verified regulatory status, intended-use evidence, and local workflow validation before clinical use. Resource constraints should be documented and addressed through safer measurement, consultation, referral, or access pathways, not converted into evidence that an unverified app is adequate.