Hematology-Oncology and Precision Medicine
Oncology and hematology generate imaging, pathology, flow-cytometry, genomic, laboratory, treatment, and outcome data that support narrowly defined AI tasks. In the randomized MASAI trial of 105,934 women, AI-supported mammography had higher sensitivity than standard double reading at the same specificity (Gommers et al., 2026). Detection, grading, prognosis, treatment selection, workflow efficiency, and patient benefit are separate claims and require different evidence.
After reading this chapter, you will be able to:
- Evaluate AI systems for cancer screening and early detection
- Understand AI applications in pathology and radiology for cancer diagnosis
- Assess AI tools for peripheral blood smear and bone marrow analysis
- Navigate AI-assisted flow cytometry interpretation for hematologic malignancies
- Apply genomic AI tools for treatment selection in solid tumors and leukemias
- Evaluate AI in radiation therapy planning and CAR-T response prediction
- Recognize limitations and failure modes of hematology-oncology AI
- Apply evidence-based frameworks for adopting AI in cancer and blood disorder care
Introduction: AI Across Hematology and Oncology
Hematology and oncology sit at the forefront of AI in medicine. These fields generate enormous volumes of complex data: peripheral blood smears, bone marrow aspirates, flow cytometry panels, imaging studies, pathology slides, genomic sequencing, treatment responses, survival outcomes. Each cancer type and hematologic disorder has distinct biology, staging systems, treatment algorithms, and prognoses. This complexity creates both opportunity and challenge for AI applications.
The opportunity: AI excels at pattern recognition in complex, high-dimensional data. Classifying cells on blood smears, detecting blast populations in bone marrow, interpreting flow cytometry patterns, identifying genomic variants, and predicting treatment response are tasks where machine learning can potentially augment human expertise. Hematology is particularly suited to AI given the structured nature of cell morphology and immunophenotyping data.
The challenge: Cancer is not one disease but hundreds. Hematologic malignancies add layers of complexity: lineage determination, blast quantification, measurable residual disease assessment. Sample sizes for rare disorders are limited. Treatment decisions involve nuanced tradeoffs between efficacy and toxicity, survival and quality of life, evidence-based guidelines and individual patient preferences. These decisions require expertise, empathy, and judgment that current AI cannot replicate.
AI applications span screening, diagnosis, blood and bone marrow morphology, flow cytometry, genomic analysis, treatment planning, and prognostication. The IBM Watson for Oncology failure remains the field’s most prominent cautionary tale for premature deployment.
Cancer Screening and Early Detection
Lung Cancer Screening CT Analysis
Clinical context: Low-dose CT screening reduces lung cancer mortality by 20% in high-risk smokers (NLST, NEJM 2011), with the European NELSON trial reporting a 24% reduction (NEJM 2020). But screening programs face challenges: radiologist workload, inter-reader variability, and false positives requiring invasive follow-up.
AI enhancement: - Automated lung nodule detection and volumetry - Lung-RADS classification assistance - Reduction in false positives
Evidence boundary: - Regulatory authorization and performance are product-specific. Evidence for one product cannot support a class-wide sensitivity, reading-time, or false-positive reduction claim. - Nodule detection, nodule characterization, Lung-RADS support, screening-program outcomes, and lung-cancer mortality are separate endpoints. - The Radiology chapter maintains the modality-specific evidence and FDA evaluation standard.
New FDA clearances (2026):
- Median Technologies eyonis LCS 1.1: FDA cleared K251474 on February 6, 2026 for detection, localization, and characterization of solid and part-solid pulmonary nodules on low-dose chest CT (FDA K251474). The record does not establish a mortality benefit for a screening program.
- RevealDx RevealAI-Lung: FDA cleared K251769 on January 30, 2026 as computer-assisted diagnostic software for characterizing lesions suspicious for cancer (FDA K251769). At a prespecified example threshold in the reader study, sensitivity and specificity increased with device assistance, but FDA cautions that the values are sample-dependent and should not be interpreted as population statistics.
These clearances establish defined assistive intended uses, not autonomous screening, fewer invasive procedures, or lower lung-cancer mortality. Institutions should evaluate the exact version, threshold, CT protocol, prevalence, radiologist interaction, and downstream diagnostic pathway.
Implementation considerations: - AI assists radiologist interpretation, does not replace it - False negatives still occur (especially for ground-glass opacities, small nodules) - Integration with Lung-RADS reporting systems essential - Patient communication about AI-assisted interpretation
Breast Cancer Mammography AI
Evidence: - Deep learning matches or exceeds radiologist performance - Published in Nature - The initial MASAI safety analysis randomized 105,934 women and found a 44.3% reduction in screen-reading workload with AI-supported screening (Lång et al., 2023). - Full follow-up found 80.5% sensitivity with AI-supported screening versus 73.8% with standard double reading, at 98.5% specificity in both groups (Gommers et al., 2026). Interval and aggressive cancer differences favored AI numerically but were not statistically significant.
MASAI provides strong program-level evidence for its defined Swedish screening workflow. It does not establish that every mammography model, threshold, reader configuration, or population will produce the same effect.
Limitations: - Performance varies by breast density (dense breasts more challenging) - Most training data from screening populations (may not generalize to diagnostic mammography) - Racial bias: many systems trained predominantly on white women - Does not replace clinical judgment for complex cases
Colorectal Cancer AI Colonoscopy
Application: Real-time polyp detection during colonoscopy
Evidence: A systematic review of 44 randomized trials and 36,201 colonoscopies found a higher adenoma detection rate with computer-aided detection than with conventional colonoscopy (44.7% vs 36.7%; rate ratio 1.21, 95% CI 1.15–1.28), but no difference in advanced colorectal neoplasia per colonoscopy and nearly two additional nonneoplastic polyp resections per 10 procedures (Soleymanjahi et al., 2024).
Limitations: - Does not improve detection of flat or subtle serrated polyps - May increase procedure time - Cost-effectiveness debated
AI-assisted colonoscopy has randomized evidence for adenoma detection, but that intermediate quality metric should not be promoted to proven prevention of interval cancer or mortality. More detections can also mean more nonneoplastic resections, so benefits and burdens must be evaluated together.
Multi-Cancer Early Detection (MCED) Blood Tests
MCED tests use cell-free DNA methylation patterns and machine learning to detect signals from multiple cancers simultaneously from a single blood draw. Two tests are now commercially available in the US; neither has established mortality benefit.
NHS-Galleri Trial (February 2026): Primary endpoint missed
The NHS-Galleri trial randomized 142,924 asymptomatic adults aged 50–79 years in England. A 2026 ASCO meeting abstract reported that the primary endpoint, a reduction in combined stage III and IV cancers, was not met (incidence rate ratio 1.03, 95% CI 0.92–1.14; P = .6324) (Swanton et al., 2026, meeting abstract). A full peer-reviewed trial report is still needed for complete appraisal.
The abstract reported fewer stage IV diagnoses for 12 prespecified cancers in later screening rounds and an overall stage IV incidence rate ratio of 0.86 (95% CI 0.744–0.998). These secondary results do not convert a null primary endpoint into established mortality benefit (Swanton et al., 2026, meeting abstract).
Commercially available tests:
- Galleri: The test is commercially available as a laboratory-developed test, but availability does not establish FDA authorization for population screening or mortality benefit. Performance estimates should be tied to the exact study and intended-use population.
- CancerGuard: Company availability and registry announcements are not substitutes for a peer-reviewed population-screening outcome trial or an FDA record.
Clinical guidance: Neither a commercial launch nor a favorable secondary endpoint establishes that MCED screening reduces cancer mortality. Any use should preserve recommended standard screening, disclose regulatory and evidence status, and define the diagnostic pathway and potential harms after a positive result.
Skin Cancer AI
AI-supported assessment of suspicious skin lesions is covered in Dermatology. In the United States, DermaSensor is a Class II De Novo device for a narrow referral decision by non-dermatologist physicians, for qualifying suspicious lesions in patients aged 40 years or older. It is not a screening or standalone diagnostic device (FDA, DEN230008). Claims about image-level performance should not be interpreted as evidence of improved cancer outcomes.
Pathology AI
Prostate Cancer Detection, Planning, and Prognosis
Application: AI analysis of prostate biopsy whole-slide images for distinct detection, planning, and prognostic tasks
Authorized products have different intended uses:
- Paige Prostate: FDA granted De Novo authorization under DEN200080 as an adjunctive aid for identifying foci suspicious for prostate cancer on digitized hematoxylin and eosin slides. The pathologist remains responsible for the primary diagnosis, and the authorization does not cover autonomous diagnosis or Gleason grading (FDA, DEN200080).
- Avenda Unfold AI: FDA cleared K221624 to support three-dimensional cancer estimation and treatment planning in patients with known prostate cancer. It is not a screening or primary-diagnosis system (FDA, K221624).
- ArteraAI Prostate: FDA granted De Novo authorization under DEN240068 for image-only prognostic risk assessment in a defined population with localized prostate cancer. The authorized device uses digitized biopsy images, not the multimodal clinical-data architecture described in earlier research, and its authorization does not itself establish treatment-effect prediction (FDA, DEN240068).
Detection, grading, treatment planning, prognosis, and prediction of treatment benefit are different claims. Evidence or authorization for one must not be transferred to another. Gleason pattern remains clinically consequential, but the pathologist, not an adjunctive detection system, assigns the diagnosis and grade.
Implementation: - Match the exact device record to the task being performed - Preserve a pathologist-led primary review when required by labeling - Record disagreements between the system and final interpretation - Monitor scanner, stain, specimen, and site-specific performance after deployment
Breast Cancer Pathology
Applications: - HER2 scoring from IHC - Ki-67 quantification - Lymph node metastasis detection
Evidence boundary: Digital quantification may improve reproducibility for a defined stain, scanner, scoring protocol, and population. Performance for HER2 scoring cannot be transferred to Ki-67 quantification or lymph-node metastasis detection. Before deployment, confirm the reference standard, reader study, external validation, and exact regulatory record for the intended task.
HER2-low and HER2-ultralow AI: ADC eligibility
The clinical stakes of HER2 scoring increased as treatment eligibility expanded to HER2-low and HER2-ultralow disease. A multinational reader study reported at ASCO 2025 evaluated AI assistance across 105 pathologists from 10 countries and 20 digital cases per reader. The reported reader-study results should be interpreted as conference evidence about scoring agreement, not proof of improved treatment access, response, or survival. A peer-reviewed full report and prospective implementation study are needed before the exact percentages are treated as generalizable clinical-performance estimates.
Expanding Prostate Pathology AI Landscape
The current landscape illustrates why product identity and intended use matter:
- Galen Second Read v3.1-US: FDA cleared K241232 as a second-read aid for prostate biopsy whole-slide images. Its labeling should be consulted before making claims about workflow, detection, or performance (FDA, K241232).
- Paige Prostate, Avenda Unfold AI, and ArteraAI Prostate: These products address suspicious-focus detection, treatment planning for known disease, and prognosis, respectively. They are not interchangeable.
- Research and commercial pathology models: Company announcements about acquisitions, slide counts, biomarker prediction, or product launches are leads for appraisal. They do not replace an exact FDA record, a peer-reviewed external validation, or a prospective clinical-utility study.
Product-level evidence mapping means connecting the precise model version, input, output, intended user, intended population, endpoint, and regulatory record before making a clinical claim.
Pathology Foundation Models
Pathology AI research is shifting from task-specific algorithms toward foundation models such as UNI, CONCH, Virchow, and GigaPath. These models are pretrained on large slide collections and then adapted to downstream tasks such as cancer detection, biomarker estimation, prognosis, or treatment-response research. The same architectural shift is occurring in radiology and language (see Emerging AI Technologies). Dataset scale alone does not establish external validity or clinical utility.
A 2025 Journal of Clinical Oncology multinational cohort study trained a pathology foundation model on more than 130 million patches from 104,876 whole-slide images, then fine-tuned models for prognosis in gastric, esophageal, and colorectal cancers across seven cohorts. The models predicted disease-free and disease-specific survival and were also evaluated for adjuvant chemotherapy benefit prediction (Wang et al., 2025). This is clinically more consequential than cancer detection alone, but it sits in the ESMO EBAI C1/C2 territory below: prognostic and predictive AI biomarkers need high-quality retrospective validation and, for treatment-effect claims, prospective trial validation before they change adjuvant therapy decisions.
Radiology AI for Cancer Staging
Automated Tumor Segmentation and Measurement
Applications: - Automated RECIST measurements for treatment response - Tumor volumetry (more accurate than 2D diameter) - Longitudinal tracking
Evidence boundary: Automated segmentation can improve measurement reproducibility for a defined cancer, modality, acquisition protocol, reader workflow, and endpoint. A volumetric model may support longitudinal review or research without proving earlier response detection, better treatment selection, or improved patient outcomes. Any authorized product must be assessed through its exact FDA record rather than a class-wide statement.
Lymph Node Metastasis Detection
Applications: - Automated detection of suspicious lymph nodes on CT/MRI - PET-CT analysis for staging
Evidence: Research models show variable performance across cancer types, imaging protocols, nodal stations, and reference standards. A model trained for one disease or modality should not be assumed to generalize to another.
A false-negative nodal assessment can understate disease extent and change treatment planning. These systems should therefore remain adjunctive unless the exact product, population, and clinical workflow have been prospectively validated.
Genomic and Molecular AI
Tumor Mutation Profiling and Treatment Selection
Applications: - NGS data processing and variant interpretation - Identification of variants linked to authorized companion-diagnostic indications - Tumor mutational burden calculation under the exact assay labeling
Evidence: Comprehensive genomic profiling can identify alterations with established, investigational, or uncertain clinical significance. Some interpretation layers use computational methods, but a molecular assay is not automatically an AI intervention, and authorization of the assay does not validate every downstream software recommendation.
Regulatory examples: - FoundationOne CDx: FDA approved this NGS-based comprehensive genomic profiling assay through PMA P170019 in 2017 with specified companion-diagnostic indications. Current indications must be read from the current labeling and supplements rather than summarized with a fixed therapy count (FDA, 2017). - Other profiling assays: Verify the exact submission number, specimen type, genes, intended population, report language, and companion-diagnostic claims in the primary FDA record before clinical use.
Limitations: - Interpretation of variants of unknown significance (VUS) remains challenging - Off-label treatment recommendations not always evidence-based - Insurance coverage variable
AI tools accelerate variant interpretation, and platforms like Foundation Medicine and Tempus are valuable for precision oncology. But these must be integrated with multidisciplinary tumor board review, a requirement unchanged by a 20-case ChatGPT 4.0 comparison that measured recommendation volume, evidence level, and time rather than outcomes (Schmutz et al., 2025). Genomic data without clinical context leads to off-label recommendations that may not benefit patients.
ESMO EBAI: A Framework for AI-Based Biomarkers
The European Society for Medical Oncology published the first international consensus framework for AI-based biomarkers in oncology, the ESMO Basic Requirements for AI-based Biomarkers in Oncology (EBAI), developed by a 37-expert multidisciplinary panel using modified Delphi methodology (Aldea, Kather et al., 2026, Annals of Oncology, March 2026).
EBAI classifies AI biomarkers into four tiers:
| Class | Description | Validation Required |
|---|---|---|
| A | AI quantification of established biomarkers (e.g., Ki-67 by AI, HER2 scoring) | Concordance studies against reference standard |
| B | AI as indirect or pre-screening measure of a known biomarker (e.g., HER2 estimation from H&E without IHC) | Analytical validation |
| C1 | Novel AI-derived prognostic biomarker (e.g., pathology foundation model survival predictor) | High-quality retrospective real-world or clinical trial data |
| C2 | Novel AI-derived predictive biomarker of treatment effect | Prospective clinical trial validation |
EBAI identifies ground truth, performance, and generalisability as essential criteria; fairness is recommended. Generalisability must be demonstrated across intended use settings including variations in data acquisition, post-processing, and patient population.
Why this matters: Most oncology AI is pitched as Class C2 (predictive) but validated only at Class A or B levels. EBAI gives clinicians a language to interrogate vendor claims and research publications: “What class of biomarker is this, and has it met the corresponding validation standard?”
Liquid Biopsy and Minimal Residual Disease (Non-AI Comparator)
The assays and decision rules in the ctDNA trials below were not AI interventions. They are included as a comparator because they show the level of evidence required before any biomarker-guided strategy changes care. Prognostic evidence estimates risk, whereas predictive evidence shows that treatment effect differs according to biomarker status.
Application: ctDNA analysis for MRD detection after curative-intent surgery/treatment
Evidence: Multiple ctDNA platforms have shown associations between molecular residual disease and recurrence risk. Lead time, sensitivity, and clinical meaning depend on assay, tumor type, sampling schedule, treatment, and reference standard.
First RCT evidence: the DYNAMIC trial
The DYNAMIC trial in stage II colon cancer provided randomized evidence for a ctDNA-guided management strategy. Five-year follow-up showed 88% versus 87% recurrence-free survival with less chemotherapy use in the ctDNA-guided arm (Tie et al., 2025). The result supports noninferiority for this defined strategy and population. It does not establish that every ctDNA-positive patient benefits from chemotherapy or that findings transfer to other cancers and assays.
DYNAMIC supports less chemotherapy without inferior recurrence-free survival in the studied stage II colon-cancer strategy. It does not turn ctDNA into a universal treatment-selection rule.
IMvigor011: First Phase III RCT showing MRD-directed therapy improves survival
IMvigor011 enrolled 761 patients after surgery for muscle-invasive bladder cancer. The 250 eligible patients who became ctDNA-positive were randomized 2:1 to atezolizumab or placebo. Median disease-free survival was 9.9 versus 4.8 months (hazard ratio 0.64, 95% CI 0.47–0.87; P = .005), and median overall survival was 32.8 versus 21.1 months (hazard ratio 0.59, 95% CI 0.39–0.90; P = .01) (Powles et al., 2025). Persistently ctDNA-negative patients were observed rather than randomized, so their outcomes do not estimate the effect of withholding treatment.
Negative trials: DYNAMIC-III and COBRA
Not all ctDNA-guided trials succeeded. DYNAMIC-III tested de-escalation for ctDNA-negative stage III colon cancer and did not establish noninferiority: 3-year recurrence-free survival was 85.3% with ctDNA-guided de-escalation and 88.1% with standard management (Tie et al., 2025). COBRA, a stage IIA colon-cancer trial, stopped after its prespecified interim analysis when chemotherapy did not improve ctDNA clearance in the ctDNA-positive subgroup (Morris et al., 2024, meeting abstract). A strong prognostic association does not prove that acting on the marker improves outcomes.
Blood-based cancer screening: Blood-based colorectal screening assays address a different use case from postoperative MRD. FDA authorization, guideline positioning, sensitivity, and specificity must be tied to the exact assay version and intended population. Screening performance cannot be transferred to recurrence surveillance or treatment selection.
Current state of MRD evidence: ctDNA is prognostic in multiple settings, but clinical utility is cancer-, assay-, treatment-, and strategy-specific. IMvigor011 supports atezolizumab for the randomized ctDNA-positive post-cystectomy population. DYNAMIC supports a stage II colon-cancer management strategy with less chemotherapy, while DYNAMIC-III and COBRA show that de-escalation or escalation rules cannot be assumed to work. Each decision strategy needs its own comparative evidence.
Treatment Planning and Delivery
Radiation Therapy AI
Auto-Contouring: - AI-automated organ-at-risk (OAR) and tumor volume delineation - Reduces planning time from hours to minutes - Commercially available systems widely deployed
Evidence boundary: Geometric agreement depends on anatomy, cancer type, image quality, contouring convention, and reference observer. A high similarity score does not establish that a contour is clinically acceptable or that a plan improves tumor control or toxicity.
Treatment Plan Optimization: - AI-generated IMRT/VMAT plans - Knowledge-based planning using historical data - Plan quality improvements (better OAR sparing)
Commercial auto-contouring and planning products are in use, but regulatory status and workflow evidence are product-specific. The strongest current evidence supports defined contouring and planning-workflow endpoints, not a class-wide claim of better cancer outcomes.
Prospective multicenter observational evidence (2026)
A prospective multicenter observational study (NCT05787522) evaluated AI-assisted organ-at-risk delineation in 500 patients across five centers, with 37 physicians using manual, AI-generated, and AI-assisted workflows (Niu et al., 2026). AI-assisted delineation had higher geometric agreement and shorter physician time than manual annotation in the study workflow. Because the study was not a randomized patient-outcome trial, it supports contouring consistency and efficiency, not improved tumor control or toxicity. A separate multicenter retrospective planning study reported that many AI-generated plans met study criteria and were often preferred in blinded review (Yu et al., 2025).
Intraoperative Molecular Imaging AI (Research Stage)
Application: ML algorithms combined with intraoperative molecular imaging (IMI) to assess malignancy potential of lung nodules during surgery in real time.
Emerging evidence:
A 2026 cohort study from the University of Pennsylvania developed an ML algorithm combining tumor-to-background fluorescence ratios (using pafolacianine, a folate receptor-targeted tracer) with clinical variables including smoking history (Azari et al., 2026):
- Development cohort: 279 patients with indeterminate lung nodules
- Prospective validation (n=61, 74 lesions): AUC 0.864 (95% CI 0.789-0.912) for malignancy assessment
- 93.8% sensitivity, 100% specificity, 100% PPV, 71% NPV
- Algorithm produced results in under 2 minutes versus 34 minutes for frozen section analysis
Current status: Research stage, not FDA-cleared. Algorithm is proprietary (patent pending). Single-center study requiring external validation before clinical adoption.
Known limitations:
- Fails in heavy smokers (>60 pack-years) with significant anthracosis due to light-absorbing carbons producing false fluorescence
- Blood products dampen fluorescence signal, affecting accuracy in highly inflamed cases
- Granulomas produce false positives (TBR >8.5) due to folate receptor expression
- Specific to pafolacianine tracer; generalizability to other molecular imaging agents unknown
Clinical relevance: If validated externally and cleared by FDA, this approach could reduce intraoperative decision time for indeterminate pulmonary nodules. However, the technology remains investigational, and frozen section analysis remains the current standard for intraoperative diagnosis.
Systemic Therapy Selection AI
Challenge: Complex decision-making involving tumor characteristics, patient factors, evidence quality, goals of care
AI approaches: - IBM Watson for Oncology (discontinued due to poor performance, see History of AI in Medicine) - Newer systems integrating guidelines + patient data
Evidence: - Mixed at best - Concordance with oncologist decisions 50-90% depending on cancer type - Does not account for patient preferences, quality of life considerations, financial toxicity
Watson for Oncology illustrates the danger of promoting variable concordance as clinical utility (see History of AI in Medicine for full details). Complex treatment decisions involve efficacy, toxicity, goals, access, comorbidity, and uncertainty. A treatment recommendation requires verification against current evidence and patient context, regardless of whether it came from a rules engine, predictive model, or language model.
LLM treatment concordance: stage matters
A 2026 nationwide retrospective study of 13,614 patients with hepatocellular carcinoma compared three LLM outputs with recorded physician decisions (Yang et al., 2026). Concordance was low overall. Associations between concordance and survival differed by stage, but this observational design cannot show that following an LLM caused better or worse outcomes. Confounding by disease complexity, treatment eligibility, and physician selection remains possible. The study therefore supports an evidence-verification warning, not an efficacy claim for or against LLM-guided care.
A single-center retrospective comparison of ChatGPT 4.0 with a human molecular tumor board (N = 20) found a median of 3 versus 1 therapeutic recommendations per case, similar information-density (IDM) scores, moderate run-to-run consistency (median Fleiss κ 0.51), more recommendations at LoE 3–4, and shorter case-processing time (median 15.2 versus 34.7 minutes) (Schmutz et al., 2025). Those endpoints do not measure patient outcomes and do not show that the model outperformed the board.
A 2026 evaluation applied five zero-shot role-prompting frameworks (simulated MDT, multi-expert deliberation, and surgical, medical, and radiation oncologist personas) and a majority-vote ensemble to GPT-5 on 100 gastrointestinal oncology cases with institutional MDT-validated treatment-category decisions; concordance ranged from 78%–87% with no significant inter-framework differences, and specialty-characteristic language was near-universal (97%–100%) but uncorrelated with accuracy (Stifini et al., 2026). Specialty-sounding prose is not specialty-specific reasoning and is not a tumor-board replacement.
Clinical Trial Matching
AI-Assisted Trial Eligibility Screening
Application: Analyze EHR data to identify patients potentially eligible for clinical trials
Evidence boundary: Trial-matching studies measure different endpoints, including criteria extraction, prescreening accuracy, review time, referral, contact, consent, and enrollment. A system that shortens prescreening does not necessarily increase enrollment or improve representativeness.
Limitations: - Eligibility algorithms only capture structured EHR data (miss nuanced exclusions) - Requires human review for final determination - Does not solve root problems (trial design, access barriers, mistrust)
AI can help identify potentially eligible patients, but detailed eligibility assessment still requires source-document review and clinical judgment. Nuanced exclusions, current disease status, patient preference, site access, and protocol interpretation may not be captured in structured data.
Human+AI prescreening outperforms either alone
A randomized noninferiority trial (n=355 patients with NSCLC or colorectal cancer) compared three prescreening approaches: autonomous AI, human coordinator alone, and human+AI. Human+AI was noninferior to human alone for accuracy (78.7% vs 76.7%) and significantly faster (34.1 vs 43.9 minutes per review, P = 0.05). AI alone was substantially less accurate (63.5%), confirming AI cannot yet replace clinical research staff. The human-in-the-loop model (AI handles initial EHR screening, human confirms nuanced eligibility) is the current evidence-based approach (Parikh et al., 2026, Nature Communications).
Prognostication and Survival Prediction
ML-Based Survival Models
Applications: - Predict overall survival, progression-free survival - Integrate clinical + genomic + imaging features
Evidence boundary: Some models outperform selected comparators within retrospective datasets, but discrimination varies by population and endpoint. Calibration, external validation, missing-data handling, and decision consequences determine whether a model is clinically useful.
Critical limitations: - Predictions at individual patient level uncertain (wide confidence intervals) - Cannot capture all relevant factors (patient goals, social support, unmeasured confounders) - Risk of self-fulfilling prophecies (predicted short survival leads to less aggressive treatment leads to shorter survival)
Ethical concerns: - Prognostic algorithms may influence treatment intensity, hospice referral - Vulnerable to bias (if training data underrepresents certain populations) - Must not be sole basis for withholding treatment
ML models often outperform traditional nomograms, but individual predictions remain uncertain. These may inform discussions about prognosis, but they should never dictate treatment decisions. Predicted survival is not actual survival. Communicate the uncertainty transparently, and recognize the risk of self-fulfilling prophecies.
Electronic Patient-Reported Outcomes: The Overlooked Evidence Base
Before deploying LLM symptom chatbots, consider what structured electronic patient-reported outcome (ePRO) monitoring has already proven. The PRO-TECT cluster-randomized trial (52 oncology practices, 1,191 patients with metastatic cancer) assigned practices to weekly electronic symptom surveys with severe-symptom alerts to the care team versus usual care.
Final results (Nature Medicine, 2025) showed no overall survival difference (HR 0.99), but meaningful benefits:
- 12.6 vs 8.5 months to functional decline (delayed by 4 months)
- 31% delay in symptom deterioration
- 28% improvement in health-related quality of life
- 6.1% reduction in emergency visits
91% of patients would recommend the system to others. The technology is low-complexity (validated symptom questionnaires, alert logic) compared to generative AI, yet delivers RCT-level evidence of benefit. Symptom monitoring ePRO systems have stronger evidence than most AI tools in oncology. (Basch et al., 2025, Nature Medicine)
Hematology AI: Blood, Bone Marrow, and Beyond
The American Society of Hematology maintains a Subcommittee on Artificial Intelligence whose mandate includes education, collaboration with regulators and other societies, and responsible testing and use of AI in hematology (ASH, Subcommittee on Artificial Intelligence). ASH’s 2026 comments to HHS emphasize transparency, continuous validation, governance, and human review, especially for complex and rare hematologic conditions (ASH, 2026).
Peripheral Blood Smear Analysis
Clinical context: Manual peripheral blood smear review is labor-intensive, requires expertise, and suffers from inter-observer variability. Automated digital morphology systems promise standardization and efficiency.
Regulatory examples:
CellaVision systems: Digital morphology products use image analysis to preclassify cells for trained-user review. The exact model, software version, specimen type, and FDA record should be verified before making a regulatory or performance claim.
Scopio X100 family: Full-field imaging applications support specified morphology workflows. Each application and software version has its own labeling; later clearances should not be treated as evidence for every claimed cell type or use.
Evidence boundary: Digital morphology can support remote review and workflow standardization. Turnaround-time and staffing effects depend on laboratory volume, preparation quality, network performance, review policy, and comparator workflow. Atypical cells and rare populations still require expert verification.
Limitations:
- Performance varies with smear quality and staining consistency
- Atypical lymphocytes, reactive changes, and rare cell populations require human verification
- Training data predominantly from reference laboratories may not generalize to all settings
Digital morphology represents one of the more mature hematology AI applications, with genuine workflow improvements. The technology augments rather than replaces morphologist expertise.
Bone Marrow Aspirate Analysis
Clinical context: Bone marrow morphology is essential for diagnosing leukemias, myelodysplastic syndromes, lymphomas, and other hematologic disorders. Manual review is time-consuming and subject to inter-pathologist variability, particularly for blast enumeration and dysplasia assessment.
FDA-cleared systems:
- Scopio X100 and X100HT with Full Field Bone Marrow Aspirate application: FDA granted De Novo authorization under DEN230034 on March 22, 2024. The Class II device supports specimen-quality assessment, blast and plasma-cell estimation, and myeloid-to-erythroid ratio estimation. Its labeling does not authorize an autonomous leukemia diagnosis (FDA, DEN230034).
Research systems:
- Morphogo (convolutional neural network trained on 2.8 million bone marrow cell images): Achieves high accuracy for cell classification
- Mayo Clinic Histogram of Cell Types (HCT): Deep learning system generating automated cytological fingerprints from bone marrow aspirates (published in Communications Medicine 2022). Achieved 0.97 accuracy for region detection and 0.75 mean average precision for cell classification.
Evidence from reviews:
A 2024 review of AI-based cell classification in bone marrow aspirate smears found heterogeneous datasets, labels, tasks, and validation practices. Individual studies reported high within-dataset accuracy, but those percentages cannot be combined into a class-wide estimate. External validation was uncommon, so transportability across laboratories, stains, scanners, and patient populations remains uncertain.
Flow Cytometry AI
Clinical context: Flow cytometry is essential for diagnosing and classifying hematologic malignancies, requiring expert pattern recognition across high-dimensional immunophenotyping data. Manual gating is time-intensive and subject to inter-operator variability.
Evidence: Published recommendations for AI in clinical flow cytometry describe promising automated classification studies, but performance depends on the antibody panel, instrument, preprocessing, disease mix, and reference diagnosis. A high result in one pipeline should not be generalized to all B-cell malignancies or laboratories.
Augmented Human Intelligence study (American Journal of Clinical Pathology, 2021):
- UMAP dimension reduction combined with random forest classification
- Achieves automated diagnosis without manual gating on specific populations
- Demonstrates potential for decision support in routine diagnostics
Implementation considerations:
- AI models are limited to the antibody panels they were trained on
- Low-risk applications (QA/QC flagging, panel ordering) appropriate for early adoption
- High-stakes diagnostic decisions still require expert review
- Sensitivity for small pathological populations (including measurable residual disease) remains a development priority
CAR-T Cell Therapy Response Prediction
Clinical context: CAR-T therapy achieves durable complete remission in approximately 40% of relapsed/refractory DLBCL patients, but 30-60% eventually relapse or progress. Predicting response and toxicity could guide patient selection and monitoring.
Evidence:
Multicenter AI model for early relapse prediction: A retrospective multicenter model used routine clinical and laboratory variables to estimate early relapse after axicabtagene ciloleucel. Such a model may support risk stratification research, but it requires prospective evaluation of calibration, action thresholds, and whether any intervention improves outcomes before clinical use.
Deep learning image analysis:
- Pre-treatment CT and PET imaging analyzed for 770 lymph node lesions from 39 patients
- Patient-level response prediction achieved 81% accuracy, 75% sensitivity, 88% specificity using 12-month outcomes
Computational modeling:
- Models calibrated on 209 leukemia patients dissect mechanisms behind heterogeneous responses
- Predict responders, non-responders, and CD19+/CD19- relapse patterns
These models are hypothesis-generating risk-stratification tools. Their performance cannot establish that changing CAR-T eligibility, monitoring, or therapy according to the prediction improves outcomes.
Sickle Cell Disease AI
Clinical context: Vaso-occlusive crises (VOCs) cause the majority of SCD hospitalizations. Predictive models could enable preventive intervention before acute deterioration.
Evidence:
Organ failure prediction: Multilayer perceptron models using signal processing features from continuous vital sign monitoring (heart rate, blood pressure, respiratory rate) achieved 96% sensitivity and 98% specificity for predicting acute organ failure one hour before onset in ICU patients with SCD, with performance declining at longer prediction windows (Mohammed et al., 2020)
Acute kidney injury prediction: XGBoost classifiers using continuous physiological monitoring data achieved AUROC 0.91 (95% CI 0.89-0.93) for prediction up to 12 hours before AKI onset, and AUROC 0.82 (95% CI 0.80-0.83) for 48-hour prediction in hospitalized SCD patients (Zahr et al., 2024)
VOC biomarker prediction: Combination of LDH >260 U/L and hemolysis index >12 UA/L demonstrated 90% sensitivity and 72.9% specificity for predicting VOC requiring hospitalization within 1 year (Feugray et al., 2023)
Hospital readmission prediction: Machine learning algorithms (Random Forest, Logistic Regression) achieved C-statistic 0.77 for predicting 30-day readmissions, outperforming standard LACE (0.60) and HOSPITAL (0.69) indices (Patel et al., 2021)
Limitations: Clinical implementation remains limited due to data quality challenges, need for prospective validation, and infrastructure requirements for continuous monitoring systems.
Anticoagulation Management as an Adjacent Use Case
Clinical context: Warfarin’s narrow therapeutic index and significant inter-individual variability make dosing challenging. ML models integrating clinical and pharmacogenomic data show promise.
Evidence boundary: Research models have used clinical, laboratory, and pharmacogenomic variables to estimate dosing or anticoagulation outcomes. Training on participants from major trials does not turn a retrospective model into randomized evidence of model-guided benefit. Any dosing system requires prospective validation, medication-safety controls, and performance assessment across ancestry, indication, interacting drugs, diet, and care setting.
ASH Guidance and Implementation Priorities
ASH’s public committee mandate and 2026 HHS comments identify critical implementation priorities. They should not be expanded into a more specific society position than the cited documents contain.
Current barriers:
- Data quality issues across institutions
- Equity concerns (training data predominantly from academic centers)
- Lack of regulatory frameworks and safety standards
- Limited external validation studies
Recommendations:
- Validation matched to the intended claim and risk
- Performance reporting stratified by demographics
- Clear delineation of AI limitations and failure modes
- Human oversight for complex clinical decisions and accessible escalation pathways
AI Agents in Cancer Research and Precision Medicine
Traditional AI in oncology usually performs a bounded task such as image classification or risk estimation. Agentic systems instead combine a language model with planning loops, external tools, memory, and intermediate checks to attempt multistep workflows. Current evidence is concentrated in simulations, retrospective analyses, and research prototypes, not autonomous cancer care. For foundational concepts, see AI Agents: From Chatbots to Autonomous Systems.
What distinguishes AI agents from traditional AI:
| Traditional Medical AI | AI Agents |
|---|---|
| Single task (classify image, predict risk) | Multi-step workflows (plan experiment, analyze results, propose next steps) |
| Human explicitly sequences most steps | Software can propose and execute a sequence within configured permissions |
| No planning capability | Plans sequences of actions to achieve goals |
| Cannot interact with external tools | Calls external software, databases, lab equipment |
| Outputs prediction or classification | Outputs intermediate artifacts, recommendations, or proposed actions for review |
How AI agents work: Large language models (LLMs) like GPT-4, Claude, and Med-PaLM serve as the “reasoning engine.” The agent uses the LLM to:
- Understand the goal (e.g., “design a drug targeting EGFR mutation”)
- Break down into subtasks (literature review, structure prediction, binding affinity calculation, toxicity screening)
- Execute each subtask by calling specialized tools (AlphaFold for structure, docking software for binding)
- Integrate results and propose next steps
- Iterate until goal achieved
Research applications under evaluation:
1. Drug-design workflow support - Search and prioritize candidate structures under explicit constraints - Call predictive models for pharmacokinetic, toxicity, or off-target estimates - Propose structural modifications for expert and experimental review - Maintain an auditable record of assumptions, tools, and failed steps
2. Treatment-strategy research - Integrate structured representations of genomics, imaging, and prior therapy - Retrieve potentially relevant literature and guidelines - Generate candidate strategies for multidisciplinary critique - Preserve source links so each factual premise can be verified
3. Clinical-trial prescreening - Extract candidate eligibility criteria and supporting chart evidence - Flag protocol exclusions in structured and unstructured records - Rank potential matches for research-staff review, not by unproven likelihood of benefit - Draft a prescreening summary that remains subject to source-document verification
The Virtual Biotech: A Multi-Agent Drug Discovery Framework
A February 2026 preprint from Stanford’s Zou lab introduced the Virtual Biotech, a coordinated multi-agent system structured to mirror a therapeutic research organization. A Chief Scientific Officer agent receives scientific queries, delegates to domain-specialized agents covering statistical genetics, functional genomics, chemoinformatics, disease biology, and clinical data, and integrates outputs through data-driven reasoning (Zhang et al., 2026, preprint).
Reported findings across three applications:
- Clinical trial analysis: More than 37,000 clinical-trialist sub-agents autonomously curated structured outcomes from 55,984 trials and linked drug targets to multi-omic annotations derived from single-cell RNA-sequencing atlases. The agents reported that drugs targeting cell-type-specific genes were 40% more likely to progress from Phase I to Phase II and 48% more likely to reach market, with 32% lower adverse event rates
- Target evaluation: The system evaluated B7-H3 as a lung cancer target, integrating multi-omic evidence to propose an antibody-drug conjugate strategy and identify potential liabilities, without independent clinical validation of those recommendations
- Trial failure analysis: For a terminated ulcerative colitis trial targeting OSMRb, the platform inferred potential failure mechanisms and proposed biomarker-guided enrollment strategies
Important limitations:
- All findings are from a single preprint not yet peer-reviewed
- Performance metrics are reported by the authors; no independent validation exists
- Clinical trial associations are retrospective and observational
- The framework has not been prospectively tested in a drug development pipeline
Relevance for oncologists: This work illustrates how coordinated agent architectures can perform evidence synthesis tasks at a scale not feasible with manual curation. Whether these systems accelerate actual drug approval or improve patient outcomes remains undemonstrated.
Evidence status (early 2026):
A Nature Reviews Cancer primer surveyed early oncology-agent research and the need for validation (Truhn et al., 2026). The defensible evidence boundary is:
- Drug design: Agents can coordinate computational tools and generate candidates, but experimental and clinical validation remains decisive.
- Therapeutic strategy: Simulated recommendations and concordance do not establish patient benefit.
- Complex workflows: Multistep benchmark performance does not prove reliability in a live oncology workflow.
- Autonomy: Permission boundaries, human review, audit trails, and failure recovery determine whether an agent can be used safely.
Critical limitations:
AI agents can hallucinate just like base LLMs. An agent proposing a treatment regimen can:
- Fabricate drug combinations never tested in humans
- Cite nonexistent clinical trials
- Misinterpret genomic variants
- Ignore critical contraindications
No agent should autonomously make treatment decisions without physician verification.
Regulatory and ethical frameworks:
- Regulatory status: Verify the exact product and intended use in the FDA record. No autonomous oncology-treatment agent authorization is identified in the sources cited in this chapter.
- Accountability: Duties can involve clinicians, institutions, developers, and vendors, depending on jurisdiction, product labeling, workflow, warnings, contracts, and the facts of the event.
- Transparency: Many agents are “black boxes” that cannot explain reasoning step-by-step
- Bias propagation: Agents trained on biased data perpetuate disparities
Physician perspective:
AI agents could shift some work from single-task assistance toward orchestrated research workflows. The evidence should be classified by maturity:
Demonstrated or actively studied: computational tool orchestration, literature retrieval, simulated reasoning tasks, retrospective data analysis, and prescreening support.
Plausible but not established: reliable tumor-board synthesis across multimodal records, prospective acceleration of drug development, and improved trial enrollment.
Beyond the evidence summarized here: autonomous treatment selection, dynamic protocol adjustment, or unsupervised patient-care actions.
Critical questions before clinical adoption:
- Validation: Has the agent been prospectively validated in clinical workflows?
- Failure modes: What happens when agent is wrong? Can errors be caught before patient harm?
- Explainability: Can the agent articulate its reasoning for audit and learning?
- Oversight: What level of physician supervision is required and feasible?
- Liability: Who bears responsibility for agent-generated recommendations?
Current recommendation: Treat AI agents as hypothesis-generating tools, not decision-making systems. Verify all agent outputs against clinical guidelines, literature, and multidisciplinary expertise before implementation.
For technical details on agent architectures and broader healthcare applications, see Emerging AI Technologies.
Key citation: Truhn, D., Azizi, S., Zou, J., et al. (2026). Artificial intelligence agents in cancer research and oncology. Nature Reviews Cancer. DOI: 10.1038/s41568-025-00900-0
IBM Watson for Oncology: The Cautionary Tale
(Covered extensively in History of AI in Medicine, summarized here)
What the evidence shows: - Published studies primarily measured concordance with local multidisciplinary recommendations. - Concordance varied by cancer type, stage, setting, and local practice. - Those studies did not establish improved patient outcomes. - Investigative reporting raised additional concerns about unsafe or inappropriate recommendations, which should be identified as reporting rather than converted into a trial result.
Lessons: - Precision oncology requires deep expertise, not just pattern matching - Black-box recommendations unacceptable for high-stakes decisions - Marketing does not equal clinical validation - Financial incentives can override evidence
Why oncologists must be skeptical: - Cancer treatment decisions involve tradeoffs (efficacy vs. toxicity, survival vs. QOL) - Guidelines provide frameworks, not algorithms - Patient preferences central - AI cannot replace nuanced judgment
Equity and Bias in Oncology AI
Documented Disparities in Cancer Outcomes:
- Black patients have higher cancer mortality across most cancer types despite similar incidence
- Hispanic patients less likely to receive guideline-concordant care
- Rural patients face access barriers to specialized oncology care
- Low-income patients experience financial toxicity limiting treatment adherence
How AI Can Worsen Disparities:
Training Data Bias: - Most cancer datasets from academic medical centers (affluent, insured patients) - Genomic databases overrepresent European ancestry - Imaging AI trained on specific scanner types and protocols
Examples: - Breast cancer screening AI trained predominantly on white women may have lower sensitivity in Black women - Genomic classifiers may misclassify variants in underrepresented populations - Treatment recommendation AI trained on insured patients may not account for financial toxicity concerns
Mitigation: - Require diverse training datasets - Validate across demographic subgroups - Report performance stratified by race, ethnicity, SES - Address root causes of disparities (access, bias, social determinants)
Implementation Guidelines for Oncology AI
ASCO lists its Principles for the Responsible Use of AI in Oncology among current policy statements. The society source should be consulted directly because policy language and dates can change.
Before Adopting Oncology AI:
- Demand high-quality evidence:
- Prospective validation studies
- External validation in diverse populations
- Clinical outcomes (not just prediction accuracy)
- Ensure transparency:
- Explainable AI (especially for treatment decisions)
- Clear description of training data
- Known failure modes disclosed
- Maintain human oversight:
- AI assists, never replaces, oncologist judgment
- Multidisciplinary tumor board review remains standard
- Assess equity:
- Performance in underrepresented populations
- Access considerations (cost, technology requirements)
- Consider patient preferences:
- Some patients prefer human-only decision-making
- Informed consent when AI significantly influences care
Safe Implementation:
- Pilot testing in low-stakes applications first
- Parallel validation (AI + standard approach)
- Clear escalation pathways for AI-human disagreement
- Systematic monitoring for bias and errors
- Patient feedback mechanisms
Red Flags:
- Claims of autonomous treatment decision-making
- No validation in diverse populations
- Black-box recommendations without rationale
- Vendor resistance to independent evaluation
- Replacing rather than augmenting tumor boards
AI in Prior Authorization (ASCO Position Statement, May 2025):
ASCO’s 2025 position statement on AI in prior authorization calls for government oversight, validation requirements, transparency when AI affects coverage determinations, human review, and protected patient and clinician appeal rights (ASCO, 2025). A widely cited 1.2-second review figure comes from an allegation in litigation described by ASCO, not an independently measured average for insurers as a class. Oncologists should document individual denials, preserve the clinical rationale and appeal record, and distinguish verified facts from allegations.
Conclusion
AI in oncology holds immense promise, from earlier cancer detection to personalized treatment selection to accelerating drug discovery. But IBM Watson’s failure demonstrates the perils of premature deployment. Oncologists must demand rigorous evidence, transparent algorithms, and proof of clinical benefit before integrating AI into high-stakes cancer care decisions.
The goal is not just better predictions, but better outcomes for patients, especially those from communities bearing disproportionate cancer burdens.
Check Your Understanding
The three cases below are fictional decision exercises. Patient details, product outputs, complications, costs, institutional responses, and legal arguments are illustrative. They are not reports of actual events, measured outcome rates, predicted verdicts, or estimates of recoverable damages. Clinical and legal decisions require current evidence, applicable guidelines, local policy, and jurisdiction-specific review.
Scenario 1: Genomic Classifier Misinterpretation and Overtreatment
You’re a medical oncologist at a comprehensive cancer center. Your practice routinely uses the Oncotype DX Breast Recurrence Score® for treatment decisions in hormone receptor-positive, HER2-negative breast cancer.
Patient: 52-year-old woman with newly diagnosed breast cancer - Pathology: T1c (1.8 cm), N0 (0/3 sentinel nodes), ER+ 95%, PR+ 80%, HER2-, Ki-67 15% - Stage: IA (T1cN0M0) - Performance status: Excellent, no comorbidities
Oncotype DX ordered: Tumor tissue sent for 21-gene expression assay
Result: Recurrence Score = 24 - Illustrative lab interpretation: “Integrate the recurrence score with age, clinical features, current evidence, and patient preferences.”
TAILORx trial reference (you review evidence): - Recurrence Score 21-25: Chemotherapy benefit for women ≤50 years old - Recurrence Score 21-25: No chemotherapy benefit for women >50 years old - Patient is 52 → Falls into “no chemotherapy benefit” group
Your initial recommendation: Endocrine therapy alone (tamoxifen or AI for 5-10 years). No chemotherapy benefit demonstrated in TAILORx for her age and Recurrence Score.
Patient responds: “I want to do everything possible. My neighbor had breast cancer and got chemo. Why aren’t you recommending it for me?”
You explain TAILORx findings: For women over 50 with Recurrence Score 21-25, chemotherapy did not improve survival compared to endocrine therapy alone.
Patient: “But the test said ‘intermediate risk’ and ‘consider chemotherapy.’ The lab wouldn’t say that if I didn’t need it.”
Patient sees second opinion oncologist (at outside institution):
Second oncologist’s interpretation: “Your Recurrence Score is 24, which I consider concerning. I recommend chemotherapy followed by endocrine therapy to maximize cure.”
Patient returns to you confused: “Dr. Johnson says I should get chemo. Why are you saying I don’t need it?”
You review: The second recommendation appears not to apply the randomized TAILORx evidence for women older than 50 with scores of 11–25. For this fictional patient, the score is 24.
Patient chooses second opinion oncologist’s recommendation: Receives TC chemotherapy (docetaxel + cyclophosphamide) × 4 cycles
Chemotherapy course: - Cycle 1: Febrile neutropenia, hospitalized 5 days, IV antibiotics - Cycle 2: Dose-reduced, severe fatigue, neuropathy developing - Cycle 3: Further dose reduction, persistent neuropathy - Cycle 4: Completed, but patient develops grade 2 peripheral neuropathy (permanent)
Patient outcome: - Completes endocrine therapy - 5-year follow-up: No recurrence (excellent prognosis, as expected) - Permanent peripheral neuropathy affecting hands and feet (difficulty with fine motor tasks) - Lasting anxiety and financial toxicity from chemotherapy (fictional case-cost assumptions, not general estimates)
Patient later learns (from support group discussion) that chemotherapy may not have been necessary for her specific situation.
Patient files complaint with state medical board against second opinion oncologist for recommending unnecessary chemotherapy.
Question 1: What went wrong in this case?
Misinterpretation of genomic classifier results and failure to apply high-quality clinical trial evidence.
Root causes:
1. Misunderstanding of Oncotype DX Recurrence Score interpretation - Recurrence Score is not treatment recommendation but prognostic information - TAILORx trial provides treatment guidance: Score + age determines chemotherapy benefit - Second oncologist either: - Did not know TAILORx results - Misapplied them (patient age 52, not ≤50) - Ignored them in favor of “reflexive” chemo for intermediate scores
2. “Intermediate risk” label confusion - Lab report language (“consider chemotherapy, benefit uncertain”) ambiguous - Patient interpreted “consider chemotherapy” as recommendation - Second oncologist may have anchored on “intermediate risk” without applying TAILORx
3. Communication failure - First oncologist (you) correctly applied evidence but patient sought second opinion - Second oncologist contradicted first opinion without clear discussion of evidence differences - Patient caught in conflicting recommendations without framework to evaluate
4. Cognitive biases (second oncologist) - Availability bias: “My patients with intermediate scores get chemo” - Omission bias: Fear of not treating > fear of overtreating - Anchoring: Treating a recurrence-score label as a treatment command without integrating age and trial evidence
Question 2: What documentation and review questions would matter if harm led to a complaint?
Jurisdiction-specific analysis:
Medical negligence, informed consent, causation, and damages are defined by applicable law and case-specific expert evidence. A teaching chapter cannot predict liability. The following questions organize review without declaring a verdict.
Evidence and clinical reasoning:
- What did the clinician understand about TAILORx and the patient’s age, nodal status, tumor biology, and score?
- Was the recommendation consistent with current guideline language and the evidence available at that time?
- Were reasonable alternatives, uncertainty, expected benefit, and material risks explained in terms the patient could understand?
- Did the record distinguish a prognostic score from evidence that chemotherapy changes outcomes?
- Were patient preferences elicited after, rather than before, the evidence and uncertainty were explained?
Complaint-side questions:
- Did the recommendation depart from the randomized evidence for a woman older than 50 with a score of 24?
- Did the consent discussion overstate chemotherapy benefit or understate febrile neutropenia, neuropathy, and financial burden?
- Is there a supported causal link between the disputed recommendation and the claimed injury?
Response-side questions:
- What contemporaneous clinical features or guidelines supported the recommendation?
- What did the patient choose after receiving a balanced explanation?
- Were the complications recognized risks that were discussed and managed appropriately?
- Would another qualified oncologist describe the recommendation as a permissible clinical judgment in that jurisdiction and at that time?
What TAILORx contributes:
TAILORx randomized women with hormone-receptor-positive, HER2-negative, axillary-node-negative disease and recurrence scores of 11–25 to endocrine therapy alone or chemoendocrine therapy. The overall randomized comparison established noninferiority of endocrine therapy alone, while exploratory age-score interactions suggested some benefit among certain women 50 or younger (Sparano et al., 2018). The trial did not randomize scores of 26–100 to the same comparison.
For this fictional 52-year-old with a score of 24, the randomized evidence strongly informs the discussion. It does not by itself decide a legal standard, which may incorporate guidelines, patient-specific facts, expert testimony, and local law.
Informed-consent documentation:
- Record the evidence discussed, including uncertainty and the limits of the assay.
- Use absolute benefit and harm estimates only when they are sourced and applicable to the individual.
- Document alternatives, questions, the patient’s goals, and the reason for the final plan.
- Avoid writing that the patient “wanted everything” as a substitute for evidence-based counseling.
Institutional learning review:
- Compare the report template with current evidence and guidelines.
- Review whether ambiguous laboratory language encouraged an unsupported treatment inference.
- Provide a pathway for reconciling materially different second opinions.
- Audit similar cases using a transparent rule, then review each flagged case rather than automatically labeling it inappropriate.
Do not infer a verdict, settlement value, or professional-discipline outcome from a fictional fact pattern. The durable lesson is that evidence interpretation, informed consent, documentation, and causal analysis must remain separate.
Question 3: How should genomic classifiers be used to avoid overtreatment?
Best practices for Oncotype DX and similar assays:
1. Understand What Genomic Classifiers Provide
Oncotype DX Recurrence Score: - Prognostic: Estimates recurrence risk with endocrine therapy alone - Predictive (for chemotherapy benefit): Requires integration with clinical trial data (TAILORx)
NOT a treatment decision in isolation. Must apply evidence:
| Recurrence Score | Age ≤50 | Age >50 |
|---|---|---|
| 0-10 (Low) | Endocrine alone | Endocrine alone |
| 11-25 (Randomized group) | Possible benefit in selected younger subgroups; discuss ovarian-suppression and chemotherapy interpretations | Endocrine therapy alone was noninferior in the randomized evidence |
| 26-100 (High) | Chemo recommended | Chemo recommended |
2. Shared Decision-Making Framework
When discussing intermediate Recurrence Scores (11-25) with patients:
For patients >50 years old:
“Your Recurrence Score is [X], which is intermediate risk. A large clinical trial called TAILORx studied over 10,000 women with breast cancer like yours. They found that for women over 50 with intermediate scores, adding chemotherapy to hormonal therapy did NOT improve survival compared to hormonal therapy alone.
My recommendation: Hormonal therapy alone (e.g., tamoxifen or aromatase inhibitor for 5-10 years).
Chemotherapy would: - NOT improve your survival based on trial evidence - Cause side effects (fatigue, nausea, hair loss, neuropathy risk, infection risk) - Cost $100,000–$150,000 - Take 3-4 months
Do you have questions about why I’m recommending against chemotherapy?“
For patients ≤50 years old with scores 16-25:
“Your Recurrence Score is [X]. The TAILORx trial showed that for women 50 and younger with intermediate scores, chemotherapy provided a small survival benefit, about 2-3% absolute reduction in recurrence risk at 9 years.
We should discuss: - Your personal risk tolerance - The small benefit (97% do fine with hormonal therapy alone, chemo helps 2-3% avoid recurrence) - Side effects of chemotherapy - Your preferences about treatment intensity
This is a close call where either hormonal therapy alone or chemotherapy + hormonal therapy is reasonable.”
3. Combat Cognitive Biases
Oncologist biases to recognize:
Omission bias: “I’d rather overtreat than undertreat” - Counter: Overtreatment causes real harm (neuropathy, infections, financial toxicity) - Chemotherapy without proven benefit is harm, not help
Availability bias: “I’ve always given chemo for intermediate scores” - Counter: TAILORx published 2018. Update practice based on evidence
Patient pressure: “Doctor, I want to do everything” - Counter: “Everything that HELPS, yes. Chemotherapy that doesn’t help is NOT doing everything. It’s exposing you to harm without benefit.”
4. Communicate Uncertainty Appropriately
When evidence is clear (TAILORx for age >50, score 11-25): - Do NOT say: “Chemotherapy is an option we could consider” - DO say: “The evidence shows chemotherapy does not improve survival for women your age with this score. I do not recommend it.”
When evidence is less clear (e.g., score 26-30, age >50): - DO say: “Your score is 26, which is just above the range where trials showed no chemotherapy benefit. We don’t have definitive evidence for scores 26-30 in women over 50. Some oncologists offer chemotherapy, others do not. Here are the considerations…”
5. Laboratory Reporting Improvements
Better Oncotype DX report format (include age and TAILORx guidance):
Recurrence Score: 26 (Intermediate Risk)
TREATMENT GUIDANCE (based on TAILORx trial):
- Patient age: 52 years old
- For women >50 with Recurrence Score 11-25: No chemotherapy benefit demonstrated
- For Recurrence Score 26-30: Limited data; chemotherapy benefit unclear for age >50
RECOMMENDATION: Discuss with oncologist. Endocrine therapy strongly recommended.
Chemotherapy benefit uncertain for this specific age/score combination.
6. Tumor Board and Second Opinion Processes
When second opinions diverge:
- Tumor board discussion to reconcile evidence interpretation
- Clear documentation of rationale for recommendations
- Patient provided with summary: “Dr. A recommends X because [evidence Y]. Dr. B recommends Z because [rationale W]. Here’s how they differ…”
7. Audit and Accountability
Institutional review: - Track Oncotype DX scores vs. chemotherapy administration - Flag cases where chemotherapy given despite TAILORx indicating no benefit - Peer review for appropriateness - Feedback to oncologists
Example audit trigger: - Patient age >50 + Recurrence Score 11-25 + chemotherapy given → Requires documented rationale
8. Patient Decision Aids
Provide visual tools. The figures below are fictional placeholders and must be replaced with patient-specific, sourced estimates before clinical use:
Your Situation:
- Age: 52
- Recurrence Score: 24
Without any treatment: 15% chance of recurrence in 9 years
With hormonal therapy alone: 7% chance of recurrence (85% effective)
With hormonal therapy + chemotherapy: 7% chance of recurrence (NO ADDITIONAL BENEFIT)
Chemotherapy WILL cause:
- Hair loss (100%)
- Fatigue (90%)
- Nausea (60%)
- Infection risk (20%)
- Permanent neuropathy (5-10%)
- Financial cost ($30,000 out-of-pocket)
Benefits of chemotherapy for you: NONE (based on TAILORx trial)
Lesson: Genomic classifiers provide prognostic information, but treatment decisions require integration with applicable clinical-trial evidence and patient factors. Clear communication should separate proven benefit, subgroup signals, and uncertainty. The fictional scenario does not determine a legal standard or outcome.
Questions About Oncology AI
What happened to IBM Watson for Oncology?
Published studies of Watson for Oncology primarily measured concordance with multidisciplinary recommendations and found variable results across cancers and settings. They did not establish improved patient outcomes. Investigative reporting also described unsafe or incorrect recommendations, illustrating why product-specific, outcome-matched evidence is required.
How accurate is AI for lung cancer screening CT?
There is no class-wide accuracy figure for lung-screening AI. Performance depends on the product, nodule definition, threshold, scanner, population, comparator, and workflow. Detection performance does not establish fewer invasive procedures or lower lung-cancer mortality.
Does mammography AI work?
In the MASAI randomized screening workflow, AI-supported reading had higher sensitivity than standard double reading at the same specificity. The initial analysis also reported lower screen-reading workload. Interval and aggressive cancer differences favored AI numerically but were not statistically significant.
Scenario 2: Lung Nodule AI False Negative and Delayed Diagnosis
In this fictional exercise, a radiologist at a community hospital uses a generic, FDA-authorized adjunctive chest-radiograph AI system. The product name, output language, performance, and workflow are invented for teaching and do not describe a specific authorized device.
System description: - Analyzes chest X-rays for lung nodules, infiltrates, pneumothorax, other findings - Flags abnormalities for radiologist review - Deployed as “concurrent reader” (AI analysis available while you read)
Your experience: 6 months deployed, generally helpful for detecting subtle findings, occasional false positives (you override)
Case presentation: 67-year-old man - Indication: Routine chest X-ray before elective hip replacement surgery - History: 40 pack-year smoking history (quit 5 years ago), no respiratory symptoms - Patient: Feels well, no cough, no weight loss, no hemoptysis
Chest X-ray performed: PA and lateral views
Fictional AI output: “No significant abnormality detected. Low suspicion for nodule. No urgent findings.”
Your interpretation: - Lungs clear - No infiltrate, no effusion - Heart size normal - Impression: “No acute cardiopulmonary process”
You sign report without further workup recommendations
6 months later: Patient develops persistent cough
Returns for chest X-ray:
Radiologist review (different radiologist): - Right upper lobe mass, 4.5 cm - Comparison to prior 6 months ago: “In retrospect, subtle 1.2 cm nodule visible in right apex on prior X-ray, now grown to 4.5 cm”
CT chest ordered: - Right upper lobe mass, 4.5 cm, spiculated - Mediastinal lymphadenopathy - No distant metastases
Biopsy: Adenocarcinoma of lung
Staging: Stage IIIA (T3N2M0) - locally advanced, not surgical candidate
Treatment: Concurrent chemoradiation (definitive intent, not curative)
Retrospective review of original chest X-ray:
Radiologist panel reviews (3 thoracic radiologists independently): - All 3 identify subtle 1.2 cm nodule in right apex on original X-ray - Consensus: “Nodule was visible but subtle. Missed by both AI and original radiologist. Would recommend CT follow-up for nodule in smoker.”
Patient outcome: - Completes chemoradiation - 18-month follow-up: Local progression, develops brain metastases - Stage IV disease, palliative intent treatment - Prognosis: 12-18 months median survival
If nodule detected at 1.2 cm 6 months earlier: - Likely Stage IA (T1bN0M0) - Surgical resection (lobectomy) - 5-year survival: 70-80% (vs. <15% with Stage IV)
Patient files malpractice lawsuit against you and hospital:
Allegations: - Missed lung nodule on chest X-ray - Failure to recommend follow-up imaging (CT) for smoker with nodule - AI system failure contributed to miss - 6-month delay resulted in progression from curable (Stage I) to incurable (Stage III→IV) disease
Question 1: What went wrong: AI false negative, radiologist error, or both?
Root causes:
1. AI false negative - Lunit INSIGHT CXR failed to detect 1.2 cm nodule in right apex - Retrospective review: 3/3 expert radiologists identified nodule → nodule was visible - AI limitation: Apical nodules challenging (rib overlap, clavicle overlap, lower contrast)
Why did AI miss it? - Training data may have underrepresented subtle apical nodules - Algorithm threshold set for higher specificity (reduce false positives) at cost of sensitivity - 1.2 cm nodule near detection limit for chest X-ray AI
2. Radiologist error (independent of AI) - You also missed the nodule on independent read - Common miss: Apical nodules are “satisfaction of search” blind spot - No AI alert may have contributed to complacency: “AI says no nodule, must be fine”
3. Automation bias - Cognitive bias: Trusting AI negative result without independent thorough search - “AI didn’t flag anything → I can read quickly” - Parallel reading (AI + radiologist simultaneous) may paradoxically reduce sensitivity if radiologist defers to AI
4. System implementation failure - Was AI sensitivity/specificity validated on apical lung nodules specifically? - Was there training for radiologists on known AI limitations? - Was there protocol for high-risk patients (smokers) to have lower threshold for CT?
Question 2: What facts would a clinical and legal review need to examine?
Jurisdiction-specific review:
The fictional facts cannot determine negligence, product liability, institutional responsibility, causation, or damages. Those questions depend on applicable law, contemporaneous standards, contracts, product labeling, expert evidence, and the exact clinical record.
Clinical interpretation questions: - Radiologists must systematically review all lung zones including apices - Was the suspected opacity reasonably visible on the original radiograph without hindsight? - Did the report and follow-up recommendation conform to accepted chest-radiograph practice at the time? - Did the radiologist maintain an independent search pattern, or did the negative AI output alter attention? - If a nodule was suspected on radiography, was diagnostic chest CT recommended before applying CT-based nodule-management guidance?
Institutional questions:
- What evidence supported procurement and the selected workflow?
- Was the exact product version validated locally for the intended population and image acquisition conditions?
- Were users trained on limitations, known failure modes, and disagreement handling?
- Did postdeployment monitoring detect false negatives, subgroup variation, or automation bias?
- Were incidents reported through the institutional and vendor pathways required by policy and regulation?
Product questions:
- What did the FDA labeling identify as the intended use, user, population, inputs, limitations, and required review?
- Did promotional statements match the authorization and supporting evidence?
- Were material limitations and version changes disclosed?
- Was the output functioning as designed, or was there a software, integration, image-quality, or configuration failure?
Causation questions:
- What was the cancer’s biology and stage at each time point?
- Would earlier CT, biopsy, and treatment more likely than not have changed the outcome under the jurisdiction’s causation standard?
- How much of the apparent difference reflects hindsight, uncertain growth history, or assumptions about stage at the earlier date?
- Were later treatment and outcomes affected by factors independent of the disputed read?
Competing interpretations:
- A claimant could argue that a visible finding was missed, the negative output contributed to automation bias, and the delay caused harm.
- A radiologist could argue that the finding was subtle, retrospective review is affected by hindsight, the device was adjunctive, and causation remains uncertain.
- An institution could point to procurement, training, monitoring, and workflow controls, while reviewers would assess whether those controls were adequate in practice.
- A manufacturer could rely on accurate labeling and intended-use limits, while reviewers would examine design, warnings, integration, versioning, and marketing evidence.
FDA authorization does not determine civil liability, establish perfect sensitivity, or transfer the interpreting clinician’s duties to the manufacturer. It also does not immunize an institution or clinician from review of a particular workflow.
No responsible analysis can assign percentages of fault, predict settlement, or declare a likely verdict from this fictional scenario. The exercise is designed to expose the evidence and workflow questions that a real review would need to answer.
Question 3: How should lung nodule AI be implemented to avoid false negatives and automation bias?
Best practices for chest X-ray AI deployment:
1. Understand AI Limitations Before Deployment
Performance validation (REQUIRED):
Ask vendor: - “What is sensitivity for nodules by location?” (apical vs. peripheral vs. central) - “What is sensitivity for nodules by size?” (≤5mm, 6-10mm, 11-20mm, >20mm) - “What is false negative rate in high-risk population (smokers)?” - “Provide external validation data (not just internal development set)”
Site-specific validation: - Run AI on historical cases with known nodules - Calculate sensitivity/specificity on YOUR scanner, YOUR population - Identify failure modes: What types of nodules does AI miss?
Illustrative local-validation findings, not measured values from a cited product study: - Sensitivity 85% for peripheral nodules, only 60% for apical nodules - Sensitivity 90% for nodules >10mm, only 70% for nodules 6-10mm
If sensitivity unacceptable for high-risk population: Do not deploy, or adjust protocol
2. Training Radiologists on AI Limitations
Mandatory education:
- AI is adjunct, not replacement
- Known failure modes (apical nodules, small nodules, ground-glass opacities)
- Automation bias risk: “AI says negative → I read carelessly”
- Systematic search pattern INDEPENDENT of AI
- High-risk patients (smokers) require extra scrutiny regardless of AI
Key message: “Read every X-ray as if the AI doesn’t exist. Then use AI as second opinion, not first opinion.”
3. Workflow Design: Independent Read THEN AI Review
Problem with concurrent reading: Radiologist sees AI output while reading → automation bias
Better workflow:
STEP 1: Radiologist reads X-ray independently (AI output hidden)
↓
STEP 2: Radiologist documents preliminary interpretation
↓
STEP 3: AI output revealed
↓
STEP 4: If AI finds something radiologist missed → re-review
↓
STEP 5: Final interpretation (integrate AI + radiologist findings)
Benefit: Reduces automation bias by forcing independent interpretation first
4. Clinical Decision Support for High-Risk Patients
EHR integration:
When a chest radiograph is ordered for a patient with a smoking history: - Alert radiologist: “Patient is high-risk (smoker). Consider CT if any suspicious finding.” - Support a recommendation for diagnostic chest CT when a suspicious radiographic nodule is identified
Separate radiographic detection from CT nodule management:
The Fleischner Society recommendations apply to incidentally detected nodules characterized on CT, not to an unconfirmed opacity measured on a chest radiograph. The first step after a suspicious radiographic finding is usually diagnostic CT characterization. Once CT establishes size, attenuation, multiplicity, and morphology, the clinician can apply the appropriate current nodule-management guideline to the actual patient and context.
Safer recommendation support: The system may surface a configurable, nonbinding prompt such as, “Suspicious pulmonary opacity detected. Consider diagnostic chest CT if clinically appropriate.” The radiologist remains responsible for the final recommendation and should not apply a CT surveillance interval to an uncharacterized radiographic opacity.
5. Second Read Protocol for High-Risk Cases
Risk-proportionate review policy: - Define which examinations receive second review from local risk, capacity, and measured performance - Do not assume that AI plus one reader is equivalent to two independent radiologists - Audit whether the selected policy improves clinically important detection without unacceptable false positives or delay
AI as second reader: - If AI flags nodule, radiologist MUST review and document agreement or disagreement - If radiologist disagrees with AI finding, must document reason
6. Audit and Feedback
Systematic review: - Track chest X-rays with subsequent CT showing nodule - Retrospective review: Was nodule visible on X-ray? - If yes → miss event → root cause analysis
Radiologist-specific feedback: - “You missed 3 apical nodules in past 6 months. Review these cases for learning” - Targeted education on personal blind spots
AI performance monitoring: - Track false negatives (nodule on subsequent CT, not flagged by AI) - Predefine review triggers from local baseline performance, clinical consequences, and statistical uncertainty rather than an arbitrary universal percentage
7. Informed Consent for AI-Assisted Imaging
Should patients be told AI is used?
Arguments for disclosure: - Transparency - Some patients may want to know - Medicolegal protection (“We use AI to improve detection”)
Arguments against: - May confuse patients - AI is internal radiologist tool, not separate test
Compromise: General disclosure without detail - Hospital website: “We use advanced computer analysis to assist radiologists in detecting abnormalities” - Radiology report footer: “AI-assisted interpretation”
8. Vendor Accountability
Questions for lung nodule AI vendors:
- “What is false negative rate for apical lung nodules specifically?”
- “Provide data on nodules missed by AI that were visible to radiologists”
- “What post-market surveillance do you conduct to track false negatives?”
- “What contractual, reporting, and incident-investigation obligations apply if a false negative contributes to delayed diagnosis?”
- “Will you indemnify radiologists for cases where AI false negative was contributing factor?”
RED FLAGS: - Vendor cannot provide location-specific or size-specific sensitivity data - Vendor claims “95% sensitivity” without subgroup analysis - Vendor resists post-market surveillance or false negative tracking - No mechanism for reporting AI failures
Lesson: Lung-nodule AI can support detection, but false negatives and automation bias remain possible. Radiologists need an independent search pattern, product-specific knowledge, and a monitored disagreement workflow. A suspicious finding on chest radiography should be characterized with diagnostic CT before CT-based Fleischner recommendations are applied. Legal responsibility is fact- and jurisdiction-specific.
Scenario 3: Liquid Biopsy MRD-Directed Treatment and Overtreatment
In this fictional exercise, a gastrointestinal oncology program offers an institutional tumor-informed ctDNA assay for molecular residual disease after colorectal-cancer surgery. The assay name, output, clinical pathway, and costs are invented. They do not describe Guardant Reveal, which uses a tissue-free approach rather than a patient-specific tumor-informed design.
Test: Fictional institutional tumor-informed ctDNA assay - Analyzes patient-specific tumor mutations - Detects residual ctDNA after curative-intent surgery - Predicts recurrence risk
Patient: 58-year-old man with colon cancer - Stage: IIA (T3N0M0) - tumor invaded through muscularis propria, no lymph node involvement - Surgery: Right hemicolectomy, R0 resection (clear margins) - Pathology: Moderately differentiated adenocarcinoma, 0/18 lymph nodes positive, no high-risk features - Microsatellite status: Microsatellite stable (MSS)
Illustrative baseline plan for this average-risk stage IIA case: - Surveillance after surgery - Discuss the limited expected benefit and known toxicities of adjuvant chemotherapy using current guideline and patient-specific evidence - Reconsider the plan if additional high-risk clinical or pathologic features are identified
Your initial plan: Surveillance (no chemotherapy)
Patient enrolled in institutional ctDNA study: Testing at 4 weeks after surgery
ctDNA result at 4 weeks post-op: POSITIVE (MRD detected) - 2 tumor-specific mutations detected in plasma at low levels - Interpretation: “MRD detected. High risk of recurrence.”
You review CIRCULATE-Japan trial data (Kotani et al., 2023): - ctDNA-positive patients across the stage II-IV GALAXY cohort: 61.4% disease recurrence at median 16.7 months follow-up, with 18-month DFS of only 38.4% - ctDNA-negative patients: 9.5% recurrence rate (18-month DFS 90.5%) - ctDNA very strong prognostic biomarker (HR 10.0 for recurrence)
Your assessment: The positive result indicates higher recurrence risk, but the cited cohort does not provide a precise 3-year probability for this individual or prove that surveillance necessarily means incurable metastatic detection.
You discuss with patient:
“Your blood test detected tumor-associated DNA after surgery. In research cohorts, that has been associated with a substantially higher risk of recurrence. It does not tell us your exact probability or prove that chemotherapy will change that risk. We should review the trial evidence, treatment harms, surveillance options, and available clinical trials.”
Patient: “Will chemotherapy prevent recurrence?”
You: “We do not have definitive proof that adding chemotherapy solely because of this result improves outcomes in a patient with your baseline features. Biological plausibility is not enough. The options include a clinical trial, treatment after a transparent uncertainty discussion, or a defined surveillance strategy.”
Patient: “If there’s an 80% chance of recurrence, I want to do everything possible. Let’s do chemo.”
Treatment: FOLFOX chemotherapy (oxaliplatin + 5-FU + leucovorin) × 6 months (12 cycles)
Chemotherapy course: - Cycle 3: Oxaliplatin dose reduction for developing peripheral neuropathy - Cycle 6: Grade 2 neuropathy, cold sensitivity severe - Cycle 9: Grade 3 neuropathy (difficulty buttoning shirts, dropping objects), oxaliplatin discontinued - Completes 5-FU/leucovorin only for cycles 10-12
Post-chemotherapy: - ctDNA testing at 3 months post-chemo: NEGATIVE (MRD cleared) - Surveillance imaging: No evidence of disease
2-year follow-up: No recurrence, disease-free
3-year follow-up: Still no recurrence
Patient outcome: - Cancer-free (excellent) - Permanent grade 2 peripheral neuropathy (affects hands and feet, permanent disability) - Financial toxicity: substantial fictional case costs, not a general estimate - Quality of life: Neuropathy limits hobbies (guitar playing, woodworking), early retirement due to disability
Patient reflects: - Joined online support group for colon cancer survivors - Learns that randomized evidence is strategy- and cancer-specific, and does not prove that chemotherapy helps every ctDNA-positive stage II patient - Learns that cohort recurrence estimates are not an individual counterfactual and do not show how chemotherapy changes that risk - Wonders: “Did I really need chemotherapy, or would I have been one of the 20% who never recurred?”
Patient feels uncertain about whether he was helped or harmed by MRD-directed treatment.
Question 1: Was MRD-directed chemotherapy appropriate for this patient?
Evidence review:
What the cited evidence supports: - Postoperative ctDNA status is associated with recurrence risk in defined cohorts - Absolute risks vary by cancer stage, assay, sampling schedule, treatment, follow-up, and analysis population - A cohort association does not reveal what would happen to the same patient under chemotherapy versus observation
What remains strategy-specific: - Whether a particular chemotherapy regimen improves outcomes for this exact ctDNA-positive population - Whether treatment effect differs according to this assay result - Whether serial clearance is a treatment mediator, a prognostic marker, or both
Critical distinction: Prognostic ≠ Predictive
Prognostic: Tells you risk (ctDNA does this well) Predictive: Tells you if treatment will help (ctDNA has NOT been proven to do this)
Randomized trial evidence now available: - DYNAMIC: The stage II management strategy reduced chemotherapy use and was noninferior for 5-year recurrence-free survival in the studied population (Tie et al., 2025). - COBRA: The trial stopped after a prespecified interim analysis because chemotherapy did not improve ctDNA clearance in the ctDNA-positive subgroup (Morris et al., 2024, meeting abstract). - DYNAMIC-III: ctDNA-guided de-escalation did not establish noninferiority for 3-year recurrence-free survival (Tie et al., 2025).
Current evidence: DYNAMIC supports a defined stage II management strategy with less chemotherapy and noninferior recurrence-free survival. It does not prove that chemotherapy specifically benefits every ctDNA-positive patient. COBRA and DYNAMIC-III reinforce that an intuitive biomarker-guided escalation or de-escalation rule can fail.
Question 2: What should have been communicated to the patient?
Honest informed consent conversation:
“Your ctDNA test is positive, which means the cited cohorts place you in a higher-risk group. They do not give an exact individual probability or tell us how much chemotherapy changes your risk. The question is whether a treatment strategy tested in a sufficiently similar population applies to you.
What we know: - You’re high-risk (ctDNA tells us that)
What we DON’T know: - Whether chemotherapy will help you - The absolute treatment effect for this exact assay, stage, regimen, and patient - Whether an applicable randomized strategy supports escalation, de-escalation, or neither
Two options:
Option 1: Chemotherapy now - Potential benefit: Might reduce recurrence risk (but unproven) - Known harms: Neuropathy and other toxicities, time burden, and financial toxicity, quantified from the applicable regimen evidence rather than this fictional example - We’re treating based on biological reasoning, not evidence
Option 2: Surveillance with serial ctDNA - Monitor ctDNA every 3 months - If ctDNA rises or imaging shows recurrence: Treat at that time - Risk: May detect recurrence when metastatic (harder to cure) - Benefit: Avoid chemotherapy toxicity if you’re in the 20% who do not recur
Decision process: Review current guidelines and trials, assess whether the patient matches the studied population, seek multidisciplinary input, and document why the selected option is reasonable. Shared decision-making does not make every unsupported option equivalent.”
Key elements: - Distinguish prognostic vs. predictive - Acknowledge lack of evidence for treatment benefit - Present both options as reasonable - Emphasize shared decision-making
Question 3: How should MRD testing be used to avoid premature or harmful interventions?
Best practices for ctDNA/MRD-directed therapy:
1. Distinguish Prognostic from Predictive Biomarkers
Prognostic biomarker: - Tells you risk - Example: ctDNA positivity → 80% recurrence risk - Does NOT tell you if treatment helps
Predictive biomarker: - Tells you if treatment will benefit - Example: KRAS wild-type in CRC → anti-EGFR therapy helps - Requires clinical trial validation
ctDNA is prognostic in multiple adjuvant settings, while predictive and clinical-utility claims remain assay-, cancer-, treatment-, and strategy-specific.
2. Evidence Requirements Before MRD-Directed Treatment
Level 1 evidence needed (randomized controlled trial): - Question: Does chemotherapy in ctDNA-positive patients improve outcomes compared to observation? - Trial design: ctDNA-positive patients randomized to chemotherapy vs. observation - Endpoint: Recurrence-free survival, overall survival
Until an applicable comparative strategy supports the decision: treatment based solely on a ctDNA result remains investigational for this fictional use case
3. Informed Consent for MRD Testing
Pre-test counseling (REQUIRED):
“We’re offering a blood test that detects microscopic cancer DNA. This test tells us your recurrence risk but does NOT tell us if chemotherapy will help you.
If test is positive: - The result may place the patient in a higher-risk group, but the magnitude is study-specific - The team will review whether an applicable randomized strategy supports treatment - The result may create a difficult decision if clinical utility has not been established
If test is negative: - Residual risk remains, and a negative result does not rule out recurrence - Surveillance should follow the applicable evidence and guideline, not the test alone
Do you want this test, understanding it may create uncertainty about treatment decisions?”
Some patients may decline MRD testing to avoid treatment dilemmas
4. When MRD-Directed Treatment May Be Reasonable
Scenarios where unproven therapy is more justifiable:
Disease with an established chemotherapy indication: - Treat according to current standard evidence - Use ctDNA to change duration or intensity only when an applicable trial or protocol supports that decision - Do not infer benefit from persistent positivity alone
Higher-uncertainty use: - Add treatment solely because of ctDNA positivity when the baseline indication is weak or absent - Extrapolate across assays, stages, tumor types, or regimens - Treat a prognostic association as proof of treatment benefit
5. Alternative Strategies to Avoid Overtreatment
Surveillance-based approach:
Post-op MRD testing:
If MRD-NEGATIVE → Routine surveillance
If MRD-POSITIVE →
Option A: Discuss chemotherapy (unproven) vs. close surveillance
Option B: Serial ctDNA monitoring every 8-12 weeks
- If ctDNA clears spontaneously → Continue surveillance
- If ctDNA persists or rises → Intensify surveillance, consider treatment
- If imaging detects recurrence → Treat at that time
Rationale: Serial testing may add information, but spontaneous clearance, assay variability, and lead-time bias must be considered. Monitoring should not become an improvised treatment trigger without evidence.
6. Clinical Trial Enrollment
Preferred evidence-generating option when available and appropriate: - Enroll in an active trial that matches the cancer, stage, assay, and treatment question - Contribute to evidence base - Avoid premature standard-of-care designation
If no trial available: - Discuss uncertainty transparently - Document shared decision-making - Consider tumor board review
7. Manage Patient Expectations
Avoid creating panic:
Poor communication: “Your test is positive. We need to start chemo immediately or cancer will come back!”
Better communication: “Your test shows higher risk, which is concerning. We need to discuss whether chemotherapy is right for you, given that we don’t yet know if it helps in this situation. Let’s review your options carefully.”
8. Financial Counseling
MRD testing can create financial burden: - Current assay prices and insurance coverage must be verified at the time of care - Patient responsibility can vary by laboratory, indication, plan, and assistance program - Testing can generate downstream imaging, visits, and treatment costs
Chemotherapy based on MRD: - Treatment costs vary across regimen, setting, insurance, supportive care, and complications - Coverage may be uncertain when the indication is investigational - Financial toxicity is a clinical outcome and should be discussed before testing and treatment
Discuss costs upfront: Verify a patient-specific estimate rather than quoting a fixed fictional price. Explain that a positive result may create additional tests or treatment decisions with uncertain coverage.
9. Audit and Accountability
Institutional review: - Track MRD testing → chemotherapy decisions - Peer review: Was chemotherapy appropriate given evidence? - Outcomes: Did MRD-directed chemotherapy improve outcomes vs. historical controls?
Fictional audit example, included only to illustrate why uncontrolled comparisons cannot establish benefit: - 40 MRD-positive Stage II patients - 35 received chemotherapy (88%) - Recurrence rate: 15% (vs. predicted 80% with observation) - Questions: - Did chemotherapy reduce recurrence, or would 15% have been the actual rate anyway? - Did 25 patients get unnecessary chemotherapy?
Without RCT, cannot answer these questions definitively
10. Ethical Considerations
Primum non nocere (first, do no harm): - Chemotherapy causes real harm (neuropathy, infections, financial toxicity) - Giving unproven treatment is harm without proven benefit - Burden of proof: Treatment should be proven beneficial before standard use
Equipoise: - If truly uncertain whether chemotherapy helps, randomized trial is ethical and necessary - Offering unproven chemotherapy outside trial may deprive science of answer
Patient autonomy: - Some patients prefer aggressive treatment despite uncertainty - Informed consent allows this - BUT must be truly informed (not coerced by incomplete information)
Lesson: ctDNA can provide strong prognostic information, but predictive and clinical-utility claims are specific to the assay, cancer, treatment, and tested decision strategy. DYNAMIC, DYNAMIC-III, COBRA, and IMvigor011 produced different results because they asked different questions. A biomarker-guided strategy must be validated as a strategy before it becomes a treatment rule.