Obstetrics and Gynecology
AI in obstetrics and gynecology operates in a setting where maternal autonomy, fetal assessment, reproductive decisions, and persistent inequities intersect. In final 2024 U.S. data, the maternal mortality rate was 44.8 deaths per 100,000 live births among non-Hispanic Black women and 14.2 among non-Hispanic White women (CDC/NCHS, 2026). An algorithm that performs well on average can still worsen care if its errors, thresholds, or downstream actions differ across patient groups.
After reading this chapter, you will be able to:
- Evaluate AI systems for fetal monitoring and pregnancy risk prediction
- Understand AI applications in prenatal screening and ultrasound interpretation
- Assess AI tools for cervical and breast cancer screening
- Navigate ethical challenges specific to maternal-fetal medicine
- Identify failure modes in obstetric AI (high-stakes, two-patient scenarios)
- Recognize equity concerns in women’s health AI
- Apply evidence-based frameworks for OBGYN AI adoption
Introduction
OBGYN AI faces challenges that differ from many other specialties. A recommendation may affect maternal health, fetal status, delivery timing, reproductive choices, and future fertility at the same time. Maternal mortality disparities remain stark: final 2024 U.S. data reported 44.8 maternal deaths per 100,000 live births among non-Hispanic Black women, more than three times the 14.2 rate among non-Hispanic White women (CDC/NCHS, 2026). The high-liability environment adds complexity, but FDA status or an algorithm output does not determine the legal standard of care.
AI applications in this field must navigate these realities while addressing genuine clinical needs: fetal monitoring interpretation with inter-observer agreement of only κ=0.30-0.50, preterm birth prediction where effective interventions remain limited, and screening tools requiring nuanced counseling about positive predictive values.
Obstetric AI Applications
Fetal Monitoring and Interpretation
Electronic Fetal Monitoring (EFM) Interpretation:
Clinical problem: Cardiotocography (CTG) monitoring is universal in US labor management, but interpretation is highly subjective. Inter-observer agreement κ=0.30-0.50 (poor) (Blackwell et al., 2011).
AI approaches:
- Automated CTG interpretation
- Fetal heart rate pattern recognition
- Prediction of fetal acidemia/hypoxia
Evidence:
- Multiple ML systems developed (Comert et al., 2018; Zhao et al., 2019)
- Reported sensitivity on the open-access CTU-UHB database ranges from 63.5% for a time-frequency support vector machine (Comert et al., 2018) to 98.2% for a convolutional neural network (Zhao et al., 2019)
- Corresponding specificity ranges from 65.9% to 94.9% on that same database
- Published in Computers in Biology and Medicine (Comert et al., 2018)
Randomized outcome evidence:
The INFANT trial randomized 47,062 women receiving continuous intrapartum electronic fetal monitoring to computerized CTG decision support or no decision support. The composite poor neonatal outcome occurred in 172 infants (0.7%) in the decision-support group and 171 (0.7%) in the control group (adjusted risk ratio, 1.01; 95% CI, 0.82–1.25), with no significant developmental difference at two years (INFANT Collaborative Group, 2017). This is direct evidence that adding one computerized interpretation system did not improve maternal or neonatal outcomes; it is not a class-wide verdict on every later model.
Major limitations and implementation questions:
- High false positive rates leading to increased cesarean sections
- A large RCT of one computerized decision-support system showed no improvement in the prespecified neonatal outcome
- Risk of automation bias (providers over-relying on AI categorizations)
Antenatal comparator evidence: A Cochrane review of antenatal CTG found no clear overall evidence of improved perinatal outcome. In two small studies involving 469 women, computerized antenatal CTG was associated with lower perinatal mortality, while cesarean rates and Apgar scores did not differ (Grivell et al., 2015). This antenatal question is distinct from intrapartum computerized interpretation and should not be used as product-specific AI evidence.
Retrospective classification accuracy, antenatal CTG evidence, and the INFANT intrapartum randomized trial answer different questions. Procurement decisions should require the exact product’s intended use, external validation, alert burden, clinician response workflow, subgroup performance, and comparative maternal and neonatal outcomes.
Preterm Birth Prediction
Clinical problem: Preterm birth affects 10% of US pregnancies (Martin et al., 2022) and is the leading cause of neonatal morbidity/mortality.
Traditional screening:
- Cervical length ultrasound
- Fetal fibronectin testing
- Clinical history
AI enhancement:
- Integrate EHR data (demographics, medical history, labs, medications)
- Predict spontaneous preterm birth <37, <34, <28 weeks
Evidence:
- In 112,963 nulliparous women, the best model reached AUC 0.60 using first-trimester data and 0.65 using second-trimester data, rising to 0.80 only when complications arising during the pregnancy were included (Arabi Belaghi et al., 2021)
- Better than individual risk factors alone but modest improvement
- Published in PLOS ONE (Arabi Belaghi et al., 2021)
- A systematic review of this literature found that data handling, feature selection, and evaluation metrics are frequently unjustified, threatening internal and external validity (Sharifi-Heris et al., 2022)
Limitations:
- Positive predictive value depends on the exact outcome, horizon, population, and threshold and should be reported with uncertainty
- Unclear how to act on predictions (progesterone, cerclage only effective in specific subgroups)
- Social determinants of health (stress, racism, housing instability) not captured in EHR
- Racial disparities in preterm birth (Black women 1.5x rate) (March of Dimes, 2015) not explained by medical factors
The research is promising (Arabi Belaghi et al., 2021), but clinical utility remains uncertain. A risk model does not by itself specify which intervention should follow. Models limited to EHR variables also cannot be assumed to represent structural racism, chronic stress, housing conditions, environmental exposure, or access barriers.
Delivery Date Prediction from Ultrasound Imaging
Clinical problem: Traditional delivery date estimation relies on last menstrual period and ultrasound-based gestational age dating, both of which carry uncertainty, particularly in pregnancies with irregular cycles, late prenatal care, or discordant dating criteria. Inaccurate delivery date estimates can lead to inappropriate timing of interventions including induction and cesarean delivery.
Delivery Date AI (Ultrasound AI): FDA granted De Novo authorization on February 11, 2026 (FDA, DEN250007). The prescription post-processing software analyzes pregnancy ultrasound images to provide a predicted delivery date for patients aged 18 years or older with a singleton pregnancy from 14 0/7 through 36 6/7 weeks who lack a reliable estimated delivery date because of an unreliable last menstrual period or no first-trimester ultrasound. The device does not identify at-risk pregnancies or specific complications.
Evidence:
- The PAIR (Perinatal Artificial Intelligence in Ultrasound) Study evaluated 5,714 patients in collaboration with the University of Kentucky, reporting an R-squared value of 0.92 for predicting days to delivery (vendor-sponsored study; Ultrasound AI, PAIR Study, Journal of Maternal-Fetal & Neonatal Medicine, 2025)
Clinical considerations:
- Intended as an aid to clinical judgment in the labeled population, alongside standard methods for assessing gestational age
- Compatible with most existing ultrasound machines with minimal workflow disruption
- Independent prospective validation beyond the pivotal study has not yet been published
- Performance across diverse populations (race, BMI, gestational complications) requires further characterization
- The R-squared metric describes correlation with actual delivery timing; clinical impact on obstetric decision-making and outcomes remains to be demonstrated in prospective trials
Prenatal Genetic Screening and Ultrasound
Cell-Free DNA Screening (NIPT) Enhanced Reporting:
Application: Computational analysis of sequencing data for aneuploidy screening. The clinical evidence for cfDNA screening should not be described as proof of a distinct machine-learning intervention unless the evaluated algorithm is specified.
- Trisomy 21, 18, 13 detection
- Sex chromosome abnormalities
- Microdeletion syndromes (emerging)
Evidence:
- NIPT sensitivity >99% for Trisomy 21 (Norton et al., 2015)
- Published in NEJM (Norton et al., 2015)
- Research continues on algorithmic signal processing and reporting, but product-specific claims require separate validation (Lee et al., 2022)
Important caveats:
- NIPT is screening, not diagnostic (amniocentesis/CVS for confirmation)
- False positives occur (especially low-prevalence conditions)
- Incidental findings (maternal malignancy) require counseling infrastructure
- Access and coverage vary by payer and jurisdiction and should be checked at the time of care
cfDNA is the most sensitive and specific screening test for common fetal aneuploidies, but it is not equivalent to diagnostic testing (ACOG, 2026). A positive result should be followed by genetic counseling, detailed anatomic evaluation, and an opportunity for CVS or amniocentesis. Positive predictive value depends on the condition, prior probability, population, and test.
Automated Fetal Ultrasound Analysis:
Applications:
- Nuchal translucency measurement
- Fetal biometry (head circumference, abdominal circumference, femur length)
- Anatomic survey automated views and measurements
- Placental localization
Evidence:
- Automated gestational-age estimation from 3D fetal brain ultrasound correlated closely with true gestational age (r=0.98, accurate within 6.10 days) and outperformed standard clinical dating by 4.57 days in the third trimester (Namburete et al., 2015)
- Standard scan-plane detection reached an average F1 score of 0.80 in real-time classification and 90.1% accuracy for retrospective frame retrieval, with 77.8% accuracy on anatomical localization (Baumgartner et al., 2017)
- Published in Medical Image Analysis (Namburete et al., 2015)
Limitations:
- Image quality dependent (maternal habitus, fetal position)
- Rare anomalies may be missed
- Does not replace skilled sonographer/perinatologist interpretation
- Liability concerns: who is responsible for missed anomalies?
Automated gestational-age estimation (Namburete et al., 2015) and standard scan-plane detection (Baumgartner et al., 2017) improve standardization and efficiency. But this cannot replace human expertise for comprehensive fetal assessment, and the liability question remains unresolved when AI misses a critical anomaly.
Newly FDA-Cleared Fetal Ultrasound AI (2025-2026):
Two systems received clearance that expand clinical capabilities beyond biometry:
Sonio Suspect (510(k) K243614, cleared February 21, 2025): a concurrent reading aid for interpreting physicians that automatically identifies and characterizes specified abnormal fetal ultrasound findings on detected views. The labeling states that patient-management decisions should not be made solely from its analysis.
Butterfly Gestational Age Tool (510(k) K252148, cleared in 2026): guides trained professionals through fundal-height measurement and blind ultrasound sweeps using Butterfly iQ+ or iQ3, then estimates gestational age for a singleton intrauterine pregnancy presumed to be 16–37 weeks. Its labeling states that the output is not intended for prenatal management or delivery planning.
The FDA records define different populations, inputs, outputs, and limitations; evidence for one tool cannot support the other. Clearance is not proof of improved prenatal-detection, management, or neonatal outcomes beyond the authorized use.
Maternal Risk Prediction
Preeclampsia Prediction Models:
Traditional screening:
- First trimester: maternal factors + PAPP-A + PlGF
- Fetal Medicine Foundation algorithm
AI enhancement:
- ML models integrating clinical + biochemical + ultrasound data
- Predict early-onset (<34 weeks) vs. late-onset preeclampsia
Evidence:
- Combined first-trimester screening by maternal factors, mean arterial pressure, uterine artery pulsatility index and placental growth factor detected 90% of early preeclampsia and 75% of preterm preeclampsia at a 10% screen-positive rate in 61,174 pregnancies (Tan et al., 2018)
- Better discrimination than maternal factors alone, though this is a Bayesian competing-risks model rather than machine learning
- Published in Ultrasound in Obstetrics & Gynecology (Tan et al., 2018)
- Prospective validation ongoing (ASPRE trial derivatives)
Clinical utility:
- In the ASPRE randomized trial, aspirin 150 mg reduced preterm preeclampsia among women identified as high risk by a defined first-trimester screening pathway (Rolnik et al., 2017)
- This treatment result does not establish that a machine-learning model identified the patients who benefited, and it should not be transferred to a different screening algorithm or aspirin regimen
- Current U.S. ACOG and SMFM guidance recommends aspirin 81 mg for defined clinical high-risk factors and combinations of moderate-risk factors (ACOG and SMFM, 2021)
This is promising. Combined biomarker screening achieves better discrimination than maternal factors alone (Tan et al., 2018), and an effective intervention exists for appropriately selected high-risk patients. The screening method, treatment regimen, population, and endpoint must remain connected when interpreting the evidence. Prospective comparative evaluation and cost-effectiveness analysis are still needed before an AI model can be credited with improving outcomes.
Postpartum Hemorrhage (PPH) Prediction:
Challenge: PPH is the leading cause of maternal mortality globally and difficult to predict.
AI approaches:
- Real-time prediction during labor
- Integrate vitals, labs, medications, obstetric factors
Evidence:
- In 152,279 births, models built from risk factors available at labor admission predicted PPH with C statistics of 0.93 for extreme gradient boosting and 0.87 for logistic regression, holding across temporal and site validation (Venkatesh et al., 2020)
- Modest improvement over the statistical models built from the same clinical risk factors
- Published in Obstetrics & Gynecology (Venkatesh et al., 2020)
Limitations:
- Many PPH cases occur in low-risk women (unpredictable)
- Interventions (uterotonics, surgical preparedness) already standard
- Alert fatigue if predictions not actionable
Modest improvement over the logistic model built from the same clinical risk factors (Venkatesh et al., 2020), but clinical benefit uncertain. The fundamental challenge: many hemorrhages occur in women with no identifiable risk factors, and prophylactic uterotonics are already routine. Research ongoing, but actionability remains the issue.
Digital CBT for Perinatal Depression
Bijen et al. pooled 16 RCTs (7,577 participants) of digital CBT versus control for perinatal depressive symptoms (Bijen et al., 2026). Post-intervention EPDS scores were lower with digital CBT (pooled MD -2.03, 95% CI -2.73 to -1.34), with substantial heterogeneity. In meta-regression, postpartum delivery looked more effective than antenatal delivery (b = -1.99, 95% CI -3.90 to -0.08, p = 0.04), but that contrast became p = 0.051 after excluding one high-risk-of-bias study. This is RCT meta-analytic evidence for digital CBT as a perinatal intervention, not a validated obstetric AI device, and the timing advantage is not robust.
Gynecologic AI Applications
Cervical Cancer Screening
AI-Assisted Cytology and HPV Testing:
Traditional screening:
- Pap smear cytology
- HPV DNA testing
- Co-testing strategies (ASCCP guidelines)
AI enhancement:
- Automated cytology interpretation
- HPV genotyping risk stratification
- Colposcopy image analysis
Evidence:
Automated Pap cytology: Detected 92.6% of CIN2 and 96.1% of CIN3+ in referred women (Bao et al., 2020)
Sensitivity equivalent to skilled cytologists with 26% higher specificity, and 12% higher sensitivity than primary-hospital cytology doctors
Published in Gynecologic Oncology (Bao et al., 2020)
AI evaluation of cervical images: Automated visual evaluation of cervigrams identified cumulative precancer and cancer with an AUC of 0.91 (95% CI 0.89-0.93), versus 0.69 for expert cervigram interpretation and 0.71 for conventional cytology (Hu et al., 2019)
Published in JNCI (Hu et al., 2019)
Meta-analysis (77 studies) confirms AI superior to experienced colposcopists: OR 1.75 (95% CI 1.33-2.31; p<0.0001), with 81% vs. 74% substantial agreement with histopathology (Liu et al., eClinicalMedicine, 2024)
AI-based VIA and point-of-care screening in low-resource settings:
The largest prospective evaluation of AI cervical screening enrolled 24,447 women across 57 facilities in Malawi, Rwanda, Senegal, Zambia, and Zimbabwe (African AVE Collaborative, Lancet Glob Health, 2026):
- AI ran on Samsung A21s smartphones (no specialist equipment required)
- Automated visual evaluation (AVE) sensitivity: 60.1% (95% CI 55.5-64.5%) for CIN2+; specificity 81.9%
- Sensitivity exceeded standard VIA with moderate loss in specificity; funded by Unitaid
For cytology screening where laboratory infrastructure is absent, a $300 compact microscope paired with an Att-Transformer AI (trained on 3,510 slides, validated across 4 independent datasets) achieves accurate precancer detection at approximately 1/100th the cost of conventional systems (Bai et al., Nat Commun, 2025).
A 2026 multicentre randomised crossover trial (n=1,920) tested AI assistance in routine cervical cytopathology. Deep learning improved non-expert cytopathologist sensitivity from 71.3% to 85.7% (difference 14.3%, 95% CI 7.6% to 21.1%; p<0.001), exceeding the prespecified 5% superiority margin, with comparable specificity (86.5% vs 85.1%). Reading time dropped from 175 to 31 seconds per slide (Xue et al., npj Digital Medicine, 2026). RCT-level evidence for AI-assisted diagnostics remains uncommon; most published cervical AI studies are retrospective.
AI-assisted cervical screening has been evaluated in automated cytology, colposcopic image interpretation, and smartphone-based visual assessment. The 2026 five-country study was a prospective observational diagnostic-accuracy study, not a randomized screening-outcomes trial. AVE increased sensitivity relative to VIA with lower specificity, and only 18,086 of 24,447 eligible participants had confirmed final status. The result supports task-specific diagnostic evaluation in government facilities; it does not establish mortality reduction or universal performance across devices and programs.
Breast and Ovarian Cancer Risk Prediction
Breast Cancer Risk Models Enhanced with AI:
Traditional models: Gail, Tyrer-Cuzick, BRCAPRO
AI enhancement:
- Integrate mammographic density, SNPs, family history, reproductive factors
- Polygenic risk scores
Evidence:
- A mammography-based deep-learning model was compared with established risk models in a defined screening cohort (Yala et al., 2019)
- Published in Radiology (Yala et al., 2019)
- Identifies women who may benefit from enhanced screening (MRI, tomosynthesis)
Mammography-based and multimodal risk models may improve stratification in some cohorts, but a risk-model performance result is not itself a screening recommendation. Screening intensity should follow applicable breast-imaging guidance, with model-specific calibration and subgroup validation. Detailed breast-imaging evidence is covered in the Radiology chapter.
Ovarian Cancer Early Detection:
Challenge: Ovarian cancer is often diagnosed at late stage. Screening strategies (CA-125, ultrasound) have high false positive rates.
AI approaches:
- Multimarker panels analyzed with ML
- Ultrasound-based ovarian mass characterization
Evidence:
- O-RADS is a structured ultrasound risk-stratification system, not itself an AI model (Cao et al., 2021)
- AI studies may attempt to reproduce or augment structured assessment, but comparative management and surgical outcomes require separate evaluation
- Published in Gynecologic Oncology (Cao et al., 2021)
Limitation: No screening strategy (including AI-enhanced) has been shown to reduce ovarian cancer mortality in average-risk women. USPSTF recommends against screening (USPSTF, 2018).
Structured imaging assessment and investigational AI may help characterize an already identified adnexal mass. That use must not be conflated with population screening. The USPSTF recommendation against ovarian-cancer screening in asymptomatic average-risk women is not changed by adding AI to an unproven screening pathway. (USPSTF, 2018)
Surgical Planning
Endometriosis Detection and Mapping:
AI applications:
- MRI-based endometriosis detection
- Surgical planning for deep infiltrating endometriosis
- Predicting surgical complexity
Evidence:
- A systematic review of AI applied to gynecologic ultrasound covered endometriosis among other benign disorders and found that most included studies carried a high risk of bias in subject selection or index-test design, frequently lacking external validation (Moro et al., 2025)
- Reported accuracy figures come largely from single-center retrospective series and should not be read as deployment-ready performance
- Potential value is in helping surgeons plan approach and counsel patients about surgical complexity, which remains unproven prospectively
Myomectomy/Hysterectomy Planning:
- AI analysis of fibroid size, location, vascularity
- Predicts surgical approach (laparoscopic vs. abdominal)
- Estimates blood loss risk
Evidence: Limited feasibility studies. Systematic review of AI applied to gynecologic ultrasound found most studies at high risk of bias and short of external validation, so surgical-planning use remains investigational rather than established (Moro et al., 2025). Clinical judgment remains central.
2025 evidence update: A systematic review of AI applied to ultrasound imaging for benign gynecologic disorders included 59 studies across PCOS, infertility, benign ovarian pathology, endometrial and myometrial disease, pelvic floor disorders, and endometriosis (Moro et al., 2025). Most studies had high risk of bias in subject selection or index-test domains, often because sample sources, scanner models, preprocessing, or external validation were incomplete. The practical conclusion is conservative: gynecologic ultrasound AI is a research and planning adjunct until models are externally validated across scanners, operators, and representative patient populations.
AI in Assisted Reproduction and IVF
Clinical problem: IVF success rates remain under 40% per cycle even in optimal candidates. Embryo selection by morphology alone is highly subjective; time-lapse incubator-based AI selection systems aim to improve blastocyst selection and cumulative live birth rates.
Evidence:
- A multicenter, double-blind, noninferiority randomized trial across 14 IVF clinics in Australia and Europe assigned 1,066 patients to blastocyst selection by a deep-learning score (iDAScore) or by standard morphological assessment. Clinical pregnancy was 46.5% (248/533) with the algorithm and 48.2% (257/533) with embryologist morphology, and the trial did not demonstrate noninferiority against its 5% margin (Illingworth et al., Nature Medicine, 2024)
- The three-armed SelecTIMO trial across 15 Dutch clinics compared time-lapse-based embryo selection, uninterrupted time-lapse culture, and interrupted standard culture. Neither time-lapse-based selection nor uninterrupted culture improved cumulative ongoing pregnancy or live birth rates over routine methods (Kieslinger et al., The Lancet, 2023)
Limitations:
- Most studies underpowered for live birth as primary endpoint
- Heterogeneous AI algorithms compared across trials
- Selection effect: AI advantage may be largest in high-complexity cycles (poor prognosis patients)
Current RCT evidence does not support AI embryo selection as improving live birth rates over expert morphological assessment. The technology may still offer value for standardization across centers and reduction of inter-observer variability in embryo grading, but claims of improved IVF success rates require patient-level caution until prospective live birth data accumulates.
Consumer Wearables and Menstrual Cycle Tracking
Ring-form wearables that continuously record distal finger temperature, resting heart rate, and HRV have generated peer-reviewed evidence for ovulation detection that outperforms traditional calendar-based methods. Physicians are increasingly encountering patients who track their cycles with these devices and present the data during fertility, contraception, or endocrine consultations.
Oura Ring: Temperature-Based Ovulation Detection
The Oura Ring Gen 3 captures overnight distal skin temperature and identifies the luteal phase temperature shift associated with ovulation. A manufacturer-conducted validation drawing on the company’s own user database, covering 1,155 ovulatory cycles from 964 participants aged 18-52, reported that the physiological method:
- Detected ovulation in 96.4% of cycles with a 1.26-day average timing error (Thigpen et al., JMIR, 2025)
- Calendar method produced 3.44-day average error on the same dataset
- Accuracy did not differ by age or by typical cycle variability, but detection was lower in short cycles and accuracy was lower in abnormally long ones
A longitudinal study of 116 women ages 18-55 confirmed that distal finger temperature follows a reliable oscillatory pattern across menstrual phases, with resting heart rate lowest during menses in both younger and midlife groups (Alzueta et al., J Biol Rhythms, 2024). Sleep efficiency and duration showed no significant cycle-phase variation in this cohort, a finding worth noting when patients attribute sleep disturbance to their cycle.
Regulatory and clinical context:
The cited Oura study was conducted using the company’s user database and evaluated agreement with an ovulation-day reference, not pregnancy, infertility-treatment, or contraceptive outcomes. Clinical use should therefore be bounded by the study design and current product labeling rather than inferred from consumer availability.
The cited study does not establish that Oura-guided timing improves pregnancy rates or provides effective contraception. Patients pursuing fertility treatment, those with irregular or anovulatory cycles, and those relying on cycle tracking for contraception warrant established clinical evaluation rather than consumer wearable data alone. When patients present this data:
- Cycle-phase temperature and HRV trends can provide useful physiological context for symptom discussions
- Ovulation timing estimates carry uncertainty; confirmation with serum LH, basal body temperature, or ultrasound remains appropriate for fertility planning
- Temperature deviations may reflect intercurrent illness, sleep disruption, or travel across time zones rather than cycle changes
Equity and Racial Bias in OBGYN AI
Documented Disparities in Maternal Outcomes:
1. Maternal Mortality:
- Final 2024 maternal mortality rates were 44.8 deaths per 100,000 live births among non-Hispanic Black women and 14.2 among non-Hispanic White women (CDC/NCHS, 2026)
- Pregnancy-related mortality and maternal mortality use different surveillance definitions and should not be mixed without qualification
- Persists across education and income levels
- Driven by structural racism, implicit bias, access to quality care
2. Preterm Birth:
- Black women 50% higher rate of preterm birth (March of Dimes, 2015)
- Not explained by traditional medical risk factors
- Chronic stress from racism implicated (weathering hypothesis)
3. Cesarean Section Rates:
- Black and Hispanic women have higher cesarean rates after controlling for medical indications (Edmonds et al., 2013)
How AI Can Worsen Disparities:
Training Data Bias:
- Dataset composition, missingness, care access, and label quality may differ across racial, ethnic, language, payer, and geographic groups
- Many studies do not provide enough subgroup cases or uncertainty estimates to evaluate transportability
- Models can learn site-specific care patterns that do not generalize
Examples:
- Preterm birth prediction models trained predominantly on white women may underperform in Black women
- Preeclampsia risk calculators may misclassify Hispanic women (different biomarker distributions)
- Fetal monitoring algorithms may have different accuracy across racial groups (not well studied)
Pulse Oximetry Bias:
- In one hospital cohort, occult hypoxemia occurred more often among Black patients than White patients when pulse oximetry readings appeared reassuring (Sjoding et al., 2020)
- AI relying on pulse ox data inherits this measurement bias
- Affects intrapartum monitoring of mothers and neonates
Mitigation Strategies:
- Require diverse datasets for training and validation
- Stratify performance reporting by race/ethnicity
- Address social determinants of health in models
- Engage community members in AI development
- Monitor for bias after deployment
- Invest in addressing root causes of disparities, not just prediction
Ethical Challenges in Maternal-Fetal Medicine
Maternal-Fetal Conflict:
- AI recommendations may optimize fetal outcomes at maternal expense (e.g., early cesarean delivery)
- Maternal autonomy must be preserved
- Cannot treat fetus as independent patient against maternal wishes
Hypothetical teaching scenario:
- AI predicts 30% risk of stillbirth if pregnancy continues
- Recommends immediate cesarean delivery at 35 weeks
- Mother prefers expectant management to avoid surgery and prematurity risks
- The decisionally capable patient retains decisional authority after informed discussion
- ACOG states that pregnancy is not an exception to the right of a decisionally capable patient to refuse recommended treatment (ACOG Committee Opinion 664)
Informed Consent Challenges:
- How to communicate AI-generated risk predictions?
- Uncertainty in predictions (confidence intervals rarely provided)
- Risk of coercing decisions through statistical intimidation
Liability Concerns:
- If AI predicts complication and physician does not act, malpractice risk?
- If AI wrong and intervention causes harm, who is liable?
- Documentation burden: Must explain AI role and rationale for agreement/disagreement
Reproductive Autonomy:
- Prenatal screening AI may influence pregnancy termination decisions
- Access to screening differs by geography, insurance (equity)
- Disability rights advocates concerned about selective termination
Principles for Ethical Obstetric AI:
- Maternal autonomy paramount
- AI provides information, not directives
- Shared decision-making framework
- Transparent communication of uncertainty
- Respect diverse values and preferences
- Address disparities, do not worsen them
Professional Society Guidelines
Society for Maternal-Fetal Medicine (SMFM)
SMFM meetings and news releases increasingly feature AI research, but a conference session, abstract, or society news item is not a clinical-practice recommendation. The source type and study design should remain explicit.
SMFM 2025 Pregnancy Meeting Highlights:
Luncheon Roundtable: Practical Tips for Incorporating AI and Clinical Informatics Into an MFM Practice explored how AI and clinical informatics enhance documentation, clinical decision-making, and workflow efficiency in MFM settings.
AI for Postpartum Hemorrhage Prediction: Presented research on AI models using routinely collected clinical data to generate PPH risk predictions at admission, during labor, and after delivery. Trained and validated on data from 12,807 women (386 with severe PPH morbidity).
BrightHeart fetal cardiac screening research: SMFM highlighted conference research on AI-assisted congenital-heart-defect detection (SMFM, 2025). Separately, FDA cleared BrightHeart Fetal EchoScan on November 14, 2024, as an adjunct for defined second-trimester fetal-heart ultrasound examinations (FDA K242342). The conference report and FDA record support different claims.
SMFM-Featured AI Applications:
Congenital Heart Defect Detection: AI-assisted fetal echocardiography to improve detection rates of CHDs, which are missed in 50% of cases prenatally.
Maternal Risk Stratification: Integration of social determinants, clinical factors, and biomarkers for personalized risk assessment.
For current SMFM guidance and publications, see: smfm.org
American College of Obstetricians and Gynecologists (ACOG)
The current ACOG sources most directly relevant to this chapter address fetal-heart-rate monitoring, prenatal screening, informed consent, and refusal of treatment. They do not constitute a blanket endorsement of obstetric AI.
Directly applicable ACOG guidance:
Intrapartum monitoring: Clinical Practice Guideline 10 provides the current evidence-based framework for interpreting and managing intrapartum fetal-heart-rate patterns
Prenatal screening: The January 2026 ACOG Practice Advisory endorses updated SMFM cfDNA guidance and states that positive cfDNA results should be followed by genetic counseling, detailed anatomic evaluation, and a recommendation for diagnostic testing
Informed consent: ACOG Committee Opinion 819 emphasizes voluntary, understandable, patient-centered decision-making
Refusal during pregnancy: ACOG Committee Opinion 664 centers the decisionally capable pregnant patient’s authority and rejects coercive treatment
The 2026 Practice Advisory replaces older screening guidance for this purpose. It explicitly distinguishes cfDNA screening from diagnostic testing and routes positive results to genetic counseling, detailed anatomic evaluation, and an opportunity for CVS or amniocentesis.
For ACOG resources: acog.org
Clinical Guidelines for AI Adoption
This framework applies the evidence and professional guidance above. It is not an ACOG AI position statement.
Before Adopting OBGYN AI:
- Demand robust validation:
- Prospective studies showing improved outcomes (not just prediction accuracy)
- Validation in diverse populations (race, ethnicity, SES, geography)
- Transparent reporting of performance by subgroups
- Assess impact on maternal autonomy:
- Does this support or constrain patient decision-making?
- How will recommendations be communicated?
- Can patients decline AI-assisted care?
- Evaluate equity implications:
- Will this widen or narrow disparities?
- Is it accessible regardless of insurance/geography?
- Does training data reflect population diversity?
- Consider medico-legal landscape:
- What does malpractice carrier advise?
- Documentation requirements?
- Informed consent necessary?
- Ensure multidisciplinary review:
- Obstetricians, midwives, nurses, patients, ethicists
- Diverse perspectives on benefits and risks
Safe Implementation:
- Pilot testing in low-stakes scenarios first
- Enhanced informed consent process
- Clear protocols for discordance (AI says X, clinician thinks Y)
- Systematic bias monitoring (outcomes by race/ethnicity)
- Patient feedback mechanisms
- Regular performance audits
Red Flags (Avoid These Systems):
No validation in diverse populations Claims to replace clinical judgment in high-stakes decisions Black-box models without explanation Vendor resists equity audits Recommends interventions without evidence base Does not account for patient preferences/values
Future Directions
Demonstrated or in prospective evaluation:
- Enhanced preeclampsia screening with targeted prevention
- AI-assisted cervical cancer screening in low-resource settings
- Standardized fetal biometry and anatomic survey
- Improved endometriosis detection on MRI
Theoretical or early-stage applications:
- Integration of social determinants of health into risk prediction
- Real-time intrapartum decision support (with appropriate safeguards)
- Personalized cesarean delivery risk counseling
- AI-guided fertility treatment optimization
Beyond current clinical evidence:
- Predictive models for pregnancy complications incorporating genetics, environment, social factors
- Continuous remote monitoring for high-risk pregnancies
- AI-assisted robotic gynecologic surgery (surgeon-in-the-loop)
Unlikely Despite Hype:
- AI replacing clinical judgment for delivery timing/mode
- Fully automated prenatal diagnosis
- Elimination of health disparities through AI alone (requires addressing root causes)
Conclusion
AI in obstetrics and gynecology must navigate unique ethical challenges: two-patient scenarios, profound health disparities, and reproductive autonomy. While AI shows promise for improving prenatal screening, cervical cancer detection, and risk stratification, deployment must center maternal autonomy, address rather than worsen disparities, and prove clinical benefit beyond prediction accuracy.
Obstetricians and gynecologists should demand robust evidence, transparent algorithms, and equity analyses before adopting AI systems. The goal is not just more accurate predictions, but healthier mothers and babies, especially those from communities bearing disproportionate burdens of maternal morbidity and mortality.
ACOG’s current guidance on informed consent and refusal of treatment supplies the controlling principle: technology does not displace a decisionally capable patient’s authority or the need for voluntary, understandable shared decision-making.
Check Your Understanding
The following scenarios are fictional teaching exercises. Product names, algorithm outputs, institutional audit results, patient outcomes, legal arguments, and dialogue are illustrative. They are not reports of actual events, FDA-authorized products, legal holdings, verdicts, or settlement outcomes.
Scenario 1: NIPT False Positive and Pregnancy Termination
You’re an obstetrician managing prenatal care for a 34-year-old G1P0 woman at 12 weeks gestation.
Patient presents: Requesting non-invasive prenatal testing (NIPT) for aneuploidy screening. She has private insurance that covers the test. The clinician counsels her about the test’s purpose (screening, not diagnostic) and she consents.
Results at 13 weeks: NIPT positive for Trisomy 18 (Edward syndrome). The lab report states “High probability of Trisomy 18. Diagnostic testing recommended.”
Counseling: The clinician explains that cfDNA is the most sensitive and specific screening test for common aneuploidies but remains a screening test. Diagnostic confirmation with CVS or amniocentesis is offered, with condition-specific and values-sensitive counseling.
Patient’s decision: “I cannot continue this pregnancy knowing my baby has a fatal condition. I want to terminate. I don’t want an amniocentesis. The NIPT result is clear enough.”
Clinician response: “I understand this is devastating news. The positive predictive value depends on the condition, the test, the population, and the prior probability. A positive screening result is not diagnostic. The laboratory’s validated, condition-specific PPV estimate and the option of diagnostic testing should be reviewed before an irreversible decision.”
Patient: “50-60% is high enough. I cannot take the risk. Please refer me for pregnancy termination.”
You refer her to a maternal-fetal medicine specialist who performs detailed ultrasound: normal nuchal translucency, normal anatomy visible at this gestational age, no features concerning for Trisomy 18.
MFM recommendation: Strong recommendation for amniocentesis before making decisions. Patient again declines.
Pregnancy is terminated at 14 weeks at patient’s request. She requests tissue karyotyping.
Karyotype result: 46,XX. Chromosomally normal female fetus.
Patient calls 3 weeks later: “The genetics counselor told me the baby was normal. How could the NIPT be wrong? You told me it was 99% accurate. I terminated a healthy pregnancy because of your test.”
Question 1: What went wrong in this case?
NIPT false positive for Trisomy 18 with inadequate counseling about positive predictive value and the critical importance of diagnostic confirmation before irreversible decisions.
Root causes:
- Misunderstanding of screening vs. diagnostic testing
- Sensitivity is not the same as positive predictive value
- At low prevalence, even highly sensitive tests have significant false positive rates
- Laboratory reporting ambiguity
- “High probability” language may be interpreted as diagnostic certainty
- PPV not clearly stated on lab report
- Recommendation for diagnostic testing was present but not sufficiently emphasized
- Insufficient counseling infrastructure
- Pre-test counseling did not adequately prepare patient for possibility of false positive
- Post-test counseling did not successfully communicate PPV vs. sensitivity distinction
- Genetics counselor not involved until after termination
- Time pressure and patient anxiety
- Patient wanted rapid decision-making
- Anxiety about carrying potentially affected fetus overrode statistical reasoning
Question 2: Are you liable for malpractice?
Legal analysis:
Current counseling framework (per the ACOG 2026 Practice Advisory):
- Pre-test counseling requirements:
- NIPT is screening, not diagnostic
- False positives and false negatives occur
- Diagnostic testing (CVS or amniocentesis) required for definitive diagnosis
- Decisions about pregnancy continuation should not be made based on screening alone
- Post-positive counseling requirements:
- Explain positive predictive value (not just sensitivity)
- Strongly recommend diagnostic confirmation
- Involve genetics counselor for high-risk results
- Document detailed counseling and patient decision-making
Plaintiff’s argument:
- “Dr. Smith told me NIPT was 99% accurate, leading me to believe a positive result meant certainty”
- “Dr. Smith did not adequately explain positive predictive value”
- “I was not required to see a genetics counselor before termination decision”
- “If I had understood there was 40-50% chance of false positive, I would have gotten amniocentesis”
- Damages: Wrongful termination of healthy pregnancy, emotional distress, need for psychiatric care
Defense arguments:
- Pre-test consent documented: Screening vs. diagnostic distinction documented in chart
- Post-positive counseling documented: Chart note states “explained PPV ~50-60%, strongly recommended amnio before any decisions”
- MFM referral made: Specialist also counseled patient and recommended amniocentesis
- Autonomous patient decision: Patient capacity intact, made informed refusal of diagnostic testing despite counseling
- Causation question: Patient chose termination despite appropriate counseling. No breach of standard of care
Legal and risk questions for jurisdiction-specific review:
Liability cannot be predicted from this hypothetical. Relevant questions may include the applicable standard of care, the content and timing of counseling, informed refusal, access to diagnostic testing, documentation, causation, damages, and local law.
- If counseling is well documented: The record can demonstrate that screening limitations, diagnostic options, uncertainty, and patient choice were addressed
- If counseling is poorly documented: Reviewers may be unable to determine what information was communicated or understood
- Informed refusal documentation: The record should reflect the patient’s understanding and voluntary decision to decline diagnostic testing
Current ACOG guidance states that a positive cfDNA result should be followed by genetic counseling, detailed anatomic evaluation, and a recommendation for diagnostic testing with CVS or amniocentesis. It also preserves the patient’s right to accept or decline screening and diagnostic testing (ACOG, 2026).
No responsible conclusion about verdict or settlement can be drawn from a fictional scenario. Those outcomes are fact- and jurisdiction-specific.
Question 3: How should NIPT counseling be structured to prevent this scenario?
Best practices for NIPT implementation:
1. Pre-Test Counseling Framework
Required elements (document in chart):
☐ NIPT is screening, not diagnostic
☐ False positives occur, especially for rarer aneuploidies (T18, T13)
☐ Positive predictive value varies by condition and maternal age
☐ Diagnostic testing required before pregnancy decisions
☐ Incidental findings possible (maternal malignancy, vanishing twin)
☐ Patient verbalized understanding of above
2. Post-Positive Result Protocol
IMMEDIATE STEPS:
- Design a safe disclosure workflow that provides timely access to a knowledgeable clinician or genetic counselor, including when results appear in a patient portal
- Offer a counseling visit by an appropriate modality without creating avoidable delay
- Prepare counseling materials with the laboratory’s validated condition-specific PPV and diagnostic options
IN-PERSON COUNSELING:
- Review what NIPT is (screening) and is not (diagnostic)
- Provide the laboratory’s validated condition-specific PPV estimate, its assumptions, and the uncertainty around that estimate
- Avoid transferring a PPV from another laboratory, population, age group, or condition
- Review diagnostic options (CVS if <14 weeks, amniocentesis if >15 weeks)
- Explain procedure-specific benefits and risks using current local consent materials and the patient’s gestational context
Provide timely access to genetic counseling after a positive result, consistent with current ACOG guidance
INFORMED REFUSAL PROCESS if patient declines diagnostic testing:
☐ Patient counseled about PPV for specific condition
☐ Patient understands significant false positive rate
☐ Patient offered amniocentesis/CVS, risks and benefits explained
☐ Patient declines diagnostic testing
☐ Patient understands pregnancy decisions based on screening alone carry risk of false positive
☐ Patient signature acknowledging understanding
☐ Physician signature
☐ Genetics counselor signature (if involved)
3. System-Level Safeguards
Laboratory reporting standards:
- PPV must be clearly stated on result reports (not just sensitivity)
- Avoid language like “high probability” without quantification
- Flagged recommendation: “Diagnostic testing required before pregnancy decisions”
Clinical decision support:
- EHR alert or order-set support that prompts counseling and diagnostic options without coercing the patient’s decision or obstructing time-sensitive care
- Automated offer or routing for genetics consultation after a positive result
Informed consent for termination:
- Services should confirm that the patient understands the screening result, diagnostic options, uncertainty, and alternatives
- Institutions should avoid rigid requirements that override patient autonomy or create inequitable delay
4. Vendor Accountability
When evaluating NIPT vendors:
REQUIRED: - PPV reported on all result letters (not just sensitivity/specificity) - Age-stratified performance data provided - Clear statement that test is screening, not diagnostic - Support for genetics counseling (materials, hotline)
RED FLAGS: - Marketing emphasizes “99% accurate” without PPV context - No mention of false positives - Results reported through patient portal without counseling - Pressure for rapid turn-around without counseling infrastructure
Lesson: NIPT is a powerful screening tool, but positive results require diagnostic confirmation before irreversible pregnancy decisions. Counseling must emphasize positive predictive value (not sensitivity), and systems must ensure genetics counseling is integrated into the care pathway. Documentation of informed refusal is essential if patients decline diagnostic testing.
Scenario 2: Electronic Fetal Monitoring AI and Cesarean Section
A fictional laborist covers a busy community hospital labor and delivery unit. The hospital has implemented a hypothetical CTG decision-support system called “FetalGuard AI.” The product is invented for this exercise and is not represented as FDA-authorized.
System description: Real-time AI analysis of fetal heart rate tracings with categorization: - Category I (normal): Reassuring, continue labor - Category II (indeterminate): Close monitoring, consider interventions - Category III (abnormal): prompt evaluation, intrauterine resuscitative measures when appropriate, and management under current guidance based on feasibility and maternal-fetal status
Your experience: First 2 months, system seems helpful. It flags concerning tracings and reduces cognitive burden during busy shifts.
Case presentation: 28-year-old G2P1 at 39 weeks, spontaneous labor, epidural analgesia, oxytocin augmentation for slow progress.
Labor course: - Cervix: 4 cm → 7 cm over 4 hours (adequate progress) - Fetal heart rate: Baseline 140s, moderate variability, no decelerations - AI categorization: Category I (reassuring)
Hour 5 of labor: - AI alert: “Category II - Variable decelerations detected. Consider intervention.” - Your review: Tracing shows occasional mild variable decelerations (<30 seconds, <60 bpm drop), moderate variability maintained - Your assessment: Category II, likely cord compression, acceptable for vaginal delivery with close monitoring - Your plan: Continue labor, nurse to notify you of any Category III features
Hour 6 of labor: - AI escalation: “Category III - Concerning tracing. Immediate delivery recommended.” - Your review: Moderate variability still present, variable decelerations now more frequent (every 2-3 contractions), one prolonged deceleration to 90 bpm × 90 seconds (returned to baseline 140s) - Your assessment: Borderline Category II vs. III. Not clearly Category III by ACOG criteria (would need absent variability + recurrent decelerations)
You perform scalp stimulation: Accelerates to 160 bpm (reassuring, suggests no acidemia)
Your decision: Continue labor, very close monitoring. Cervix now 9 cm, pushing likely within 1 hour.
AI continues to alarm: “Category III - Immediate delivery recommended” every 5 minutes
Nurse: “Dr. Jones, the AI keeps saying Category III. Should we do a C-section?”
You: “I’m watching the tracing closely. This is Category II. The baby is responding to stimulation. She’s almost complete. Let’s get her through to delivery.”
30 minutes later: - Cervix complete (10 cm) - Patient begins pushing - AI: “Category III - Immediate delivery recommended” - Fetal heart rate: Baseline 130s with moderate variability, variable decelerations to 80s with pushing (common, usually benign)
Delivery after 45 minutes of pushing: - Live male infant, Apgar 8 at 1 minute, 9 at 5 minutes - Cord pH 7.28 (normal, >7.20) - No neonatal complications
Next day: Chart review by peer reviewer (standard QA process)
Peer reviewer’s note: “AI system categorized as Category III for 75 minutes prior to delivery. Physician did not perform cesarean section. Prolonged Category III exposure concerning. Recommend M&M review.”
M&M Committee: “Why did the clinician disagree with the system’s recommendation for expedited delivery, and how was the discordance assessed and documented?”
Question 1: Did you breach the standard of care by not performing cesarean section when AI recommended immediate delivery?
Standard of care analysis:
Current ACOG guidance:
The governing source is ACOG Clinical Practice Guideline 10, Intrapartum Fetal Heart Rate Monitoring: Interpretation and Management. The simplified definitions below support the teaching example but do not replace the full guideline and its management algorithm.
Simplified Category III criteria: - Absent baseline variability AND any of: - Recurrent late decelerations - Recurrent variable decelerations - Bradycardia
OR - Sinusoidal pattern
Key element: Absent baseline variability is required for Category III classification (except sinusoidal pattern)
Your tracing: - Moderate variability present throughout - Variable and prolonged decelerations present - Does NOT meet ACOG definition of Category III (would be Category II)
AI misclassification: The AI system categorized tracings with moderate variability as Category III, which contradicts ACOG criteria.
Physician responsibility: - Standard of care requires physician interpretation of fetal monitoring, not blind adherence to AI - AI is clinical decision support, not a replacement for clinical judgment - Physicians must understand ACOG fetal monitoring definitions and apply them - Scalp stimulation with acceleration is reassuring and supports continuing labor
Defense argument:
- Adherence to current ACOG guidance: The fictional tracing retains moderate variability and therefore does not meet the simplified Category III definition used in the exercise
- Appropriate fetal assessment: Scalp stimulation showed reassuring response
- Excellent outcome: Normal Apgar scores, normal cord pH
- AI error, not physician error: AI system misclassified Category II as Category III
- Standard of care met: Physician correctly interpreted tracing and made appropriate clinical decision
Teaching conclusion: The scenario supports a guideline-based review of the tracing and the complete clinical context. It does not establish a legal verdict. FDA authorization, if present for a real product, would not make every output correct or determine negligence.
Question 2: What if the outcome had been bad (low Apgar, neonatal encephalopathy)?
Different legal landscape with adverse outcome:
Plaintiff’s argument: - “The AI system said Category III and recommended immediate delivery” - “Dr. Jones ignored the AI warning” - “If cesarean had been performed when AI first recommended (75 minutes earlier), baby would not have brain injury” - Causation: Delay in delivery caused hypoxic-ischemic encephalopathy
Defense argument: - ACOG guidelines do not define this as Category III - Cord pH of 7.05 (hypothetical) may reflect chronic placental insufficiency, not acute intrapartum event - Cesarean 75 minutes earlier would not have prevented outcome if chronic process - AI false positive led to inappropriate recommendation
Key legal issue: Competing standards
- ACOG guidelines (physician interpretation) vs.
- An authorized product’s labeled output, if the actual product and use were within scope
Jury question: “Which standard should the physician follow when they conflict?”
Risk: Jury may be swayed by “FDA-cleared” designation and believe AI is authoritative
Expert witness battle: - Plaintiff expert: “Hospital implemented this system and physician should follow it” - Defense expert: “ACOG guidelines are standard of care, AI is adjunct only”
Legal uncertainty: The result would depend on the facts, jurisdiction, expert evidence, causation, hospital policy, product labeling, and the reasonableness of the clinical response. The scenario cannot predict a verdict. Documentation should explain the clinical interpretation and response rather than portray an AI output as an order.
Question 3: How should hospitals implement fetal monitoring AI to avoid this liability trap?
Best practices for EFM AI implementation:
1. Validation Before Deployment
Require vendor to demonstrate: - Sensitivity/specificity for Category III detection against ACOG criteria (not vendor’s internal definitions) - False positive rate (critical for alert fatigue) - Validation in diverse populations (race/ethnicity, BMI, electrode type) - Prospective trial data showing improved neonatal outcomes (not just classification accuracy)
If vendor definitions differ from ACOG: - Do not implement until alignment achieved - Insist on ACOG-compliant categorization
2. Physician Training (Mandatory)
All laborists/obstetricians must complete: - ACOG fetal monitoring workshop - Review of Category I/II/III definitions - Understanding that AI is adjunct, not replacement - Policy: Define accountable human decision ownership and escalation when clinicians and software disagree - Scenarios where AI may misclassify (examples reviewed)
Key principle: “AI provides a second opinion, not an order”
3. Clinical Protocols
When AI and physician interpretation disagree:
PROTOCOL: AI-Physician Discordance
If AI categorizes as Category III but physician assessment is Category I/II:
1. Physician documents rationale for disagreement in chart
2. Physician performs additional fetal assessment (scalp stimulation, consider scalp pH if available)
3. Physician discusses with patient: "The AI system is concerned, but I believe the tracing is reassuring based on [rationale]. I recommend continuing labor with close monitoring."
4. Physician considers second opinion from colleague if time permits
5. AI alert acknowledged but clinical judgment prevails
Documentation template:
"AI system categorized tracing as Category III. My interpretation: Category II based on moderate variability present, decelerations consistent with cord compression, reassuring response to scalp stimulation. Continuing labor with continuous monitoring per ACOG guidelines."
4. Informed Consent
Patient communication should be proportionate to the system’s material role and local policy. One possible plain-language explanation is:
“Our hospital uses an AI system to assist with fetal monitoring interpretation. This system provides real-time analysis of your baby’s heart rate. However, your physician makes all final decisions about your care. The AI is a tool to assist your physician, not a replacement for their judgment.”
5. Quality Assurance
Regular audits: - AI false positive rate (Category III alerts that did not meet ACOG criteria) - AI false negative rate (missed Category III tracings) - Cesarean section rate before/after AI (watch for increase from false positives) - Neonatal outcomes stratified by AI alert status - Bias assessment: AI performance by race/ethnicity, BMI
Triggers for re-evaluation should be prespecified from baseline variation, clinical consequences, statistical uncertainty, and governance policy. Material deterioration, unexplained intervention increases without benefit, or subgroup harm should trigger review without relying on arbitrary universal percentages.
6. Vendor Evaluation Questions
MUST ANSWER before purchase:
- “Does your system use ACOG Category I/II/III definitions exactly, or internal definitions?”
- “What is your false positive rate for Category III alerts?”
- “Provide data on cesarean section rates in hospitals using your system vs. controls”
- “Provide neonatal outcome data (Apgar, cord pH, NICU admission) in prospective trials”
- “Provide performance data stratified by maternal race, BMI, electrode type”
- “What happens legally if a physician overrides your recommendation and outcome is bad?”
- “Will you indemnify physicians for adverse outcomes when following ACOG guidelines that differ from your recommendations?”
RED FLAGS: - Vendor cannot provide false positive rate data - Vendor claims “physicians should always follow AI recommendations” - Vendor definitions of Category I/II/III differ from ACOG - No prospective outcome data, only retrospective classification accuracy - Vendor resists bias audits
Lesson: FDA clearance applies to a defined intended use and does not make every output correct. Fetal-monitoring decision support should be evaluated against current ACOG guidance and the exact clinical workflow. Institutions need clear discordance protocols, independent clinician competence, audit data, escalation pathways, and documentation of the reasoning behind the clinical response.
Questions About AI in Obstetrics and Gynecology
Not in the large INFANT randomized trial. Among 47,062 randomized women, computerized CTG decision support did not reduce the composite poor neonatal outcome, which occurred in 0.7% of both groups. Newer systems still require product-specific prospective evaluation.
A combined competing-risks screening model detected 90% of early and 75% of preterm preeclampsia at a 10% screen-positive rate in one large study. It is not a machine-learning result. The separate ASPRE randomized trial showed benefit from aspirin 150 mg in patients identified as high risk by a defined screening pathway.
In a cohort of 112,963 nulliparous women, the best model reached AUC 0.60 with first-trimester data and 0.65 with second-trimester data. Performance and clinical utility vary by target, gestational horizon, population, threshold, and available intervention.
Evidence includes retrospective image studies, meta-analysis, a 2026 prospective diagnostic-accuracy study of 24,447 women in five African countries, and an AI-assisted cytopathology crossover trial. These studies support task-specific accuracy and reader assistance, but do not establish outcome benefit for every product or screening program.
Current RCT evidence does not support AI embryo selection over expert morphology. A 1,066-patient noninferiority trial (Illingworth, Nature Medicine 2024) failed to show noninferiority against embryologist morphology, and the three-armed SelecTIMO trial (Lancet 2023) found no improvement in ongoing pregnancy or live birth rates.
A manufacturer-conducted 2025 study of 1,155 ovulatory cycles from 964 participants reported detection in 96.4% of cycles and a mean timing error of 1.26 days. The study did not evaluate pregnancy, infertility treatment, or contraceptive outcomes.
Scenario 3: Preeclampsia Prediction Algorithm and Racial Bias
A fictional maternal-fetal medicine specialist works at an academic medical center that has implemented a hypothetical preeclampsia prediction algorithm called “PreeclampSafe AI.” The product, audit data, cases, and outcomes are invented to illustrate an equity audit.
Algorithm inputs (collected at 11-13 weeks): - Maternal age, BMI, race, obstetric history - Mean arterial pressure (MAP) - Uterine artery Doppler pulsatility index (PI) - Serum PAPP-A and PlGF
Algorithm output: - High risk (≥10% risk of early-onset preeclampsia <34 weeks): Aspirin 81mg daily + enhanced monitoring - Low risk (<10% risk): Routine prenatal care
Your experience: First 6 months, algorithm seems effective. It identifies high-risk patients, and aspirin compliance is good.
Month 7 - Quality review:
Your fellow presents data at division meeting:
Preeclampsia outcomes stratified by race:
| Race/Ethnicity | Patients | Algorithm “High Risk” | Developed Preeclampsia <34 wks | Preeclampsia Cases Missed by Algorithm |
|---|---|---|---|---|
| White | 450 | 68 (15%) | 12 cases | 2/12 (17% missed) |
| Black | 280 | 28 (10%) | 16 cases | 8/16 (50% missed) |
| Hispanic | 220 | 25 (11%) | 9 cases | 3/9 (33% missed) |
Alarming finding: Algorithm identifies only 50% of early-onset preeclampsia cases in Black women, compared to 83% in White women.
Outcome of missed cases:
Case 1: 29-year-old Black G1P0, algorithm “low risk” (7% predicted risk) - Developed severe preeclampsia at 29 weeks - Delivered emergently for HELLP syndrome - Infant: 1200g, 8-week NICU stay, severe ROP requiring laser, chronic lung disease - Mother: ICU admission for eclampsia, 2 seizures despite magnesium sulfate
Case 2: 34-year-old Black G2P1, algorithm “low risk” (8% predicted risk) - Developed early-onset preeclampsia at 31 weeks - Placental abruption, emergency cesarean - Infant: 1400g, intraventricular hemorrhage grade III, long-term neurodevelopmental concerns - Mother: postpartum hemorrhage, required transfusion
Both hypothetical patients were not prescribed aspirin because the invented algorithm classified them as low risk. The ASPRE trial tested aspirin 150 mg in a population identified as high risk by a defined screening pathway and reduced preterm preeclampsia at the trial level (Rolnik et al., 2017). It cannot establish that aspirin would have prevented either invented outcome. Current U.S. guidance recommends 81 mg based on defined clinical risk factors (ACOG and SMFM, 2021).
Question 1: Why is the algorithm underperforming in Black women?
Hypotheses the audit team should investigate:
1. Training data composition - The training cohort may not represent the deployed population or may include too few outcome cases for reliable subgroup estimates - Biomarker distributions differ by race: - PAPP-A levels lower in Black women (biology, not pathology) (Spencer et al., 2005) - PlGF levels may differ - Uterine artery Doppler cutoffs developed in predominantly White cohorts
2. Algorithm inappropriately adjusts for race - “Race” entered as input variable - The model’s use of race may encode site-specific care patterns or inequities; this requires inspection rather than assumption - Confounding: Training data may have systematically under-diagnosed preeclampsia in Black women (due to bias in care access, delayed presentation)
3. Social determinants not captured - Chronic stress from racism (weathering hypothesis) - Food insecurity, housing instability - Limited prenatal care access - Neighborhood-level factors - These factors can affect risk and access but may be absent, poorly measured, or inappropriately represented in available data
4. Measurement bias - Blood-pressure measurement error related to cuff size, technique, device, or workflow can propagate through the model - The audit should test measurement quality directly rather than attribute error to race
Question 2: Are you liable for adverse outcomes in patients misclassified by biased algorithm?
Legal analysis:
Current clinical guidance for aspirin prophylaxis: - ACOG and SMFM recommend low-dose aspirin for patients meeting defined high-risk criteria and for combinations of moderate-risk factors (ACOG and SMFM, 2021) - High-risk criteria include: - History of preeclampsia, especially early-onset - Multifetal gestation - Chronic hypertension, diabetes, renal disease, autoimmune disease - Combination of moderate risk factors
Plaintiff’s argument (Case 1: 29-year-old with HELLP at 29 weeks):
- “Dr. Smith’s hospital uses a preeclampsia algorithm that is racially biased”
- “Algorithm missed 50% of Black women who developed preeclampsia, but only 17% of White women”
- “The institution failed to evaluate whether the deployed model systematically missed patients in my subgroup”
- “Hospital knew or should have known algorithm was biased and continued using it”
- Civil-rights question: Counsel would assess whether the facts implicate applicable nondiscrimination law and whether the institution responded reasonably to identified disparate performance
- Damages: Neonatal complications, maternal morbidity, NICU costs, long-term disability
Defense arguments:
1. Standard of care met: - Algorithm is one tool among many - ACOG risk criteria also applied (patient had no traditional high-risk factors) - Clinical judgment incorporated
2. Causation uncertain: - Not all high-risk women develop preeclampsia even without aspirin - ASPRE demonstrated a trial-level reduction in preterm preeclampsia with aspirin 150 mg in a defined high-risk population, but no trial can establish that treatment would have prevented either invented individual outcome - Cannot prove aspirin would have prevented this specific case
3. Algorithm bias unknown at time: - Hospital identified bias through QI process - Acting now to address it
4. Race-based differences may reflect biology, not bias: - Different biomarker distributions by race may be valid - Algorithm optimized for overall population performance
Plaintiff’s rebuttal:
Title VI and other law: Hospitals receiving federal financial assistance have nondiscrimination duties. Application to an algorithmic fact pattern is jurisdiction- and fact-specific and should be assessed by qualified counsel.
- Disparate performance can trigger clinical, ethical, civil-rights, contractual, and regulatory review even without evidence of intentional discrimination
- Governance should define predeployment subgroup review and a response to newly identified harm
- Continuing unchanged deployment after a credible safety signal increases clinical and legal risk
Risk assessment:
The scenario raises significant clinical and legal questions, especially after the disparity is discovered, but it cannot predict a verdict, regulatory action, class certification, or settlement.
- Before discovery: Review whether the institution performed reasonable predeployment validation and continued to apply current clinical risk criteria
- After discovery: Pause or constrain use while the signal is investigated, with patient-safety, equity, legal, and communication teams involved
- Collective claims: Whether a proposed class could proceed depends on legal requirements and shared facts, not the existence of an algorithm alone
- Regulatory review: The relevant agency and reporting pathway depend on federal funding, device status, intended use, and the facts
No settlement prediction is supportable from this hypothetical.
Question 3: How should hospitals address AI racial bias when discovered?
Immediate actions (within 1 week of discovery):
1. Suspend or constrain algorithm use while the disparity is investigated
Option A: Suspend entirely - Revert to ACOG clinical risk criteria - Notify all clinicians of suspension - Notify vendor of bias discovery
Option B: Controlled interim workflow (if complete suspension is not feasible) - Route decisions through current ACOG and SMFM clinical risk criteria and specialist review - Do not invent race-specific thresholds at the bedside without validation, ethical review, and legal review - Treat the model as nonauthoritative until the subgroup signal is resolved
2. Patient notification
Define the affected population through a documented lookback analysis. Any patient notification should be coordinated through patient safety, clinical leadership, ethics, communications, and legal review. A possible plain-language starting point is:
“Our hospital has identified that the preeclampsia prediction algorithm used during your pregnancy may have underestimated risk for Black and Hispanic women. If you are still pregnant, we recommend re-evaluation for aspirin prophylaxis. If you have delivered, we apologize and are reviewing your care.”
Offer: - Re-evaluation by MFM specialist - Aspirin initiation if still pregnant and appropriate - Documentation review and outcomes analysis
3. Root cause analysis
Convene multidisciplinary team: - Maternal-fetal medicine specialists - Health equity experts - Biostatisticians - Patient advocates from affected communities - Hospital legal/risk management
Investigate: - Training data composition (% by race) - Performance metrics stratified by race (should have been done before deployment) - Algorithm decision tree (how is race variable used?) - Comparison to ACOG clinical criteria
4. Vendor accountability
Demand from vendor (PreeclampSafe AI):
- Full performance data by race/ethnicity from all deployment sites
- Explanation of bias source (training data? feature engineering? race adjustment?)
- Correction plan with timeline (re-training? different features?)
- Validation in diverse cohort before re-implementation
- Contractual responsibilities for investigation, correction, notice, support, and potential losses
If vendor is unresponsive or defensive: - Terminate contract - Use the applicable FDA medical-device reporting or safety-communication pathway if the function is an authorized device and reporting criteria are met - Coordinate broader safety communication through the vendor, regulator, and institutional counsel as appropriate
5. Policy changes
New institutional policy for all AI systems:
Equity Impact Assessment (Required Before Deployment):
☐ Validation data includes adequate representation of patient populations served
- Sufficient outcome cases in clinically relevant subgroups to estimate performance with useful uncertainty
☐ Performance metrics (sensitivity, specificity, PPV, NPV) reported separately by:
- Race/ethnicity
- Age
- BMI
- Insurance status
- Language
☐ Any use or exclusion of race and ethnicity is justified, documented, and tested for clinical and equity consequences
☐ Decision thresholds and downstream actions are evaluated for clinically important subgroups
☐ Monitoring plan for ongoing bias detection (quarterly audits)
☐ Health equity committee approval obtained
6. Ongoing monitoring (post-implementation)
Quarterly audits required:
| Metric | Overall | White | Black | Hispanic | Asian | Other |
|---|---|---|---|---|---|---|
| % Classified High Risk | ||||||
| % Developed Early Preeclampsia | ||||||
| Sensitivity (% cases detected) | ||||||
| Specificity | ||||||
| PPV |
Trigger for intervention: Predefine clinically meaningful differences and statistical signals using baseline variation, uncertainty, outcome severity, and the intended intervention. Do not use an unvalidated universal 10% rule.
7. Community engagement
Engage Black and Hispanic patient communities: - Explain what happened (transparent communication) - Apologize for bias and harm - Describe corrective actions - Invite feedback on AI governance processes - Ensure representation on AI ethics committee
8. Alternative approaches that reduce bias
Feature and model review: - Removing race does not necessarily remove structural bias, because other variables may encode the same inequities - Compare models with and without contested features and evaluate calibration, error rates, clinical actions, and outcomes across groups - Prefer the approach with the strongest benefit-harm and equity profile, not simply the highest overall AUC
Social determinants integration: - Add neighborhood-level deprivation index - Screen for food insecurity, housing instability - Incorporate chronic stress measures
Validated broader clinical pathway: - Consider whether current ACOG and SMFM clinical criteria identify a broader eligible population than the model - Any threshold change should be prospectively evaluated for benefit, false positives, contraindications, workload, and equity - The tradeoff is not merely more prescriptions versus fewer missed cases; it includes patient preference, adverse effects, and downstream monitoring
Hybrid approach: - Use algorithm as one input, not sole decision-maker - Require clinician review of all “borderline” cases (e.g., 8-12% risk) - Clinical judgment can override algorithm
Lesson: High-stakes obstetric algorithms require subgroup validation, a defined response to emerging disparities, and continued application of current clinical guidance. A model should not narrow access to an effective prophylactic pathway unless that use has been directly validated. Hospitals have patient-safety, ethical, contractual, and potentially legal duties to investigate and address credible bias signals. Community engagement and transparent communication are essential when algorithmic harm is suspected.