Surgical Subspecialties
Colonoscopy AI has the largest randomized-trial evidence base among surgical AI applications, but effect sizes vary across settings and endpoints. Gastroenterology AI now extends to liver disease screening, celiac disease pathology, gastric cancer risk assessment, capsule reading, and reinforcement-learning anesthesia control. Urologic products support distinct pathology, prognosis, imaging, and planning tasks that must not be treated as interchangeable. Otolaryngology illustrates the same boundary: a prognostic model may be AI, while upgradeable firmware or conventional navigation is not automatically AI. The clinical question is never simply whether a specialty “has AI,” but which model performs which task, for whom, under what evidence and regulatory record.
After reading this chapter, you will be able to:
- Evaluate colonoscopy AI with RCT evidence and understand the implementation gap
- Assess gastroenterology AI beyond colonoscopy (liver cirrhosis screening, celiac disease pathology, gastric cancer detection)
- Identify FDA-authorized prostate cancer AI tools and their clinical evidence
- Evaluate bladder cancer, kidney stone, and robotic surgery AI applications
- Understand otolaryngology AI (cochlear implants, head and neck cancer, surgical navigation safety)
- Apply evidence-based frameworks for surgical subspecialty AI adoption
Introduction
Colonoscopy AI has something most surgical AI lacks: multiple randomized controlled trials measuring a recognized quality indicator. The pooled evidence shows higher adenoma detection, but newer pragmatic evidence demonstrates that the effect is not guaranteed in every practice. Urologic AI has advanced through product-specific authorization and multicenter diagnostic research, while otolaryngology includes both prognostic research and device functions frequently mislabeled as AI. A performance study asks whether a model predicts or detects; an implementation study asks what happens when clinicians use it; neither alone proves that patients benefit.
Part 1: Gastroenterology AI
Colonoscopy AI: The Evidence Leader
Clinical Need
Colonoscopy quality varies substantially by endoscopist. Adenoma detection rate (ADR), the proportion of screening colonoscopies detecting at least one adenoma, correlates with interval colorectal cancer risk. Each 1% increase in ADR is associated with 3% reduced interval cancer risk (Corley et al., 2014).
Computer-aided detection (CADe) systems aim to reduce missed polyps by providing real-time alerts during colonoscopy.
RCT Evidence
Meta-analysis findings (2024):
The largest meta-analysis of AI-assisted colonoscopy included 44 RCTs (Soleymanjahi et al., 2024):
- ADR increased from 36.7% to 44.7% (RR 1.21, 95% CI 1.15–1.28)
- Consistent benefit across CADe platforms
- Improved detection of sessile serrated lesions
The same analysis found similar average advanced colorectal neoplasia per colonoscopy, 0.16 with CADe versus 0.15 without it, while advanced-neoplasia detection rate increased modestly from 11.5% to 12.7%. CADe led to almost two additional non-neoplastic resections per ten colonoscopies and increased total withdrawal time by about half a minute (Soleymanjahi et al., 2024). More adenomas detected is not the same endpoint as fewer interval cancers or fewer colorectal cancer deaths.
A network meta-analysis of 13 RCTs (4,156 patients, 8 AI colonoscopy systems) found the pooled detection gain concentrated in diminutive polyps (SMD 0.21, 95% CI 0.07–0.35) with minimal effects on small and large polyps, and GRADE certainty low to very low (Ba et al., 2026). Size-stratified nuance on ADR claims, not evidence that CADe lacks any clinical value.
GI Genius specific evidence:
| Study | Design | ADR Effect |
|---|---|---|
| Repici et al., 2020 | RCT | ADR 54.8% vs 40.4% (RR 1.30) |
| COLO-DETECT, 2024 | Pragmatic RCT | ADR 56.6% vs 48.4% (adjusted OR 1.47, 95% CI 1.21-1.78) |
| Meta-analysis GI Genius studies | Multiple | Variable, I²=64% |
Real-World vs. RCT Performance
RCT vs. real-world discordance:
Real-world implementation studies show smaller or absent benefits compared to RCTs (Wei et al., 2024):
- Overall real-world ADR: 36.3% with CADe versus 35.8% without (RR 1.13, 95% CI 1.01–1.28), a statistically significant but small pooled difference that was no longer significant after excluding two abstracts
- GI Genius specifically: No significant difference (RR 0.96, 95% CI 0.85-1.07)
Why the gap?
- RCT conditions: Protocol adherence, selected endoscopists, controlled environments
- Real-world conditions: Variable technique, alert fatigue, workflow integration challenges
- Ceiling effect: High-performing endoscopists may not benefit from AI
Clinical implication: CADe is a supplement to rigorous colonoscopy technique. Baseline ADR may modify the effect, but no universal threshold identifies who will benefit.
2026 pragmatic evidence: A randomized trial in 12 German institutions included 1,627 outpatients and found no significant CADe effect on examiner-adjusted ADR, 40.6% versus 38.3%, while examiner identity had a large association with ADR (Zimmermann-Fraedrich et al., 2026). A separate 2026 private-practice RCT included 914 participants in the intention-to-treat analysis and likewise found no significant ADR difference, 34.5% with computer assistance versus 32.9% with traditional colonoscopy (Lux et al., 2026).
Because the two reports describe distinct trial populations and publication records, local evidence review must match each statistic to its exact study. The convergent practical message is narrower: positive pooled trial evidence does not guarantee a measurable benefit in every high-performing or community practice.
Living Guideline Update
The 2025 BMJ living clinical practice guideline for computer-aided detection and diagnosis of colonoscopy polyps drew on 44 RCTs with more than 30,000 participants plus microsimulation modeling. It found low-certainty evidence that CADe may increase positive endoscopy findings, but no direct evidence for colorectal cancer incidence or post-colonoscopy cancer incidence, and modeling suggested little to no 10-year effect on CRC incidence, CRC mortality, perforation, or bleeding (Foroutan et al., 2025).
Practical meaning: CADe can improve detection metrics, especially adenoma detection, but the patient-level case is still incomplete. Endoscopy units should track advanced adenomas, sessile serrated lesions, non-neoplastic resections, withdrawal time, downstream colonoscopy burden, and interval cancer rates rather than treating ADR alone as proof of benefit.
False Positive Burden
CADe increases detection of non-neoplastic polyps: - Hyperplastic polyps (not requiring removal if <5mm in rectosigmoid) - Artifacts, stool, mucosal folds triggering false alerts
Consequence: Increased polypectomy of benign lesions represents unnecessary intervention and procedural risk.
Serrated Lesion Detection
Sessile serrated lesions (SSLs) are precursors to interval cancers and historically difficult to detect. A seven-trial GI Genius-specific meta-analysis reported higher SSL detection (RR 1.27, 95% CI 1.11–1.47), but that product-specific result should not be transferred to every CADe system (Sattar et al., 2025).
The same meta-analysis found no significant advanced-adenoma detection difference (RR 1.01, 95% CI 0.90–1.13). Its product scope, seven-trial sample, and journal context make it complementary to, not a replacement for, the larger 44-trial analysis.
GI AI Beyond Colonoscopy
AI-ECG liver cirrhosis detection:
A pragmatic cluster-randomized clinical trial (98 primary care teams, 15,596 adults) tested whether an AI-enabled ECG model could identify undiagnosed advanced chronic liver disease. In the intervention group, new diagnoses of advanced liver disease doubled compared with usual care (1.0% vs. 0.5%, OR 2.09, p=0.007). Among AI-positive patients, detection was four-fold higher (4.4% vs. 1.1%, OR 4.37, p<0.001). The model detects cardiac electrical changes associated with liver cirrhosis from a routine 12-lead ECG (Simonetto et al., 2025).
Celiac disease pathology AI:
A machine learning model trained on 3,383 whole slide images of duodenal biopsies from four hospitals achieved pathologist-level diagnostic performance for celiac disease, with accuracy, sensitivity, and specificity exceeding 95% and AUC exceeding 99%. Inter-observer agreement between pathologist and model was statistically indistinguishable from pathologist-pathologist agreement (p>0.96), suggesting AI could address diagnostic bottlenecks in settings with pathologist shortages (Jaeckle et al., 2025).
GRAPE gastric cancer screening:
The GRAPE (Gastric Cancer Risk Assessment Procedure with AI) model uses noncontrast CT and deep learning to identify people at high risk for gastric cancer. Developed with 6,720 cases from two centers and validated across 16 independent centers with 18,160 cases, the model achieved an AUC of 0.927. In a reader study, assistance increased mean radiologist sensitivity by 6.6 percentage points and specificity by 13.3 percentage points. Retrospective opportunistic-screening cohorts included 78,593 consecutive noncontrast CT scans; among model-flagged people with verification, the reported cancer detection rate was 17.7–24.5%, and about 38–40% of confirmed cases in two regional hospitals lacked abdominal symptoms (Hu et al., 2025).
These cohorts evaluated diagnostic discrimination and reader assistance. They did not establish that population screening with noncontrast CT improves gastric cancer mortality or has a favorable benefit-to-harm ratio. A screening model needs evidence on downstream confirmation, incidental findings, false positives, treatment, and patient outcomes before population deployment.
Reinforcement-learning anesthesia control:
A 2026 multicenter randomized trial evaluated an automated reinforcement-learning controller for ciprofol infusion during gastrointestinal endoscopy. Among 420 randomized adults aged 18–65 years with ASA physical status I–II, 418 were analyzed after two postrandomization withdrawals. Hypoxemia occurred in 14.42% with automated control and 14.29% with clinician-managed anesthesia (OR 1.01, 95% CI 0.59–1.75), with no significant secondary safety differences (Bing et al., 2026). The result supports comparable measured safety in a selected low-risk population; it does not establish superiority, broader anesthesia autonomy, or safety in higher-risk patients.
AI-assisted capsule endoscopy:
NaviCam ProScan (AnX Robotica) received FDA De Novo authorization (DEN230027, 2024) as an adjunctive AI-assisted reading tool for small-bowel video capsule endoscopy. In FDA’s retrospective 87-patient evaluation, mean reading time decreased from 58.10 to 21.47 minutes (p<0.0001). A separate prospective seven-center ARTIC cohort evaluated diagnostic yield and noninferiority. The time result and diagnostic-yield result should remain attached to their respective designs (FDA DEN230027 decision summary). The labeling requires clinician review of the video and states that standalone sensitivity and specificity were not established.
Part 2: Urology AI
Urologic AI spans prostate pathology, prognostication, MRI, endoscopy, stone classification, and procedure planning. These are separate clinical tasks with different inputs, reference standards, and regulatory records.
Prostate Cancer AI
Prostate products span pathology detection, second review, prognostication, and planning for known disease. Evidence from one product cannot be transferred to another merely because both concern prostate cancer.
Gleason Grading AI
GleasonXAI, an explainable AI for Gleason pattern grading, was trained on 1,015 tissue microarray core images annotated by 54 international pathologists from 10 countries. Unlike conventional black-box models, GleasonXAI uses pathologist-defined terminology and “soft labels” reflecting inter-pathologist variability to provide transparent pattern-level explanations. The model achieved equivalent or better accuracy than conventional approaches while offering interpretability (Mittmann et al., 2025). The team also released the largest freely available dataset with explanatory Gleason pattern annotations.
PI-RADS and AI Integration
Multiparametric MRI (mpMRI) with PI-RADS (Prostate Imaging Reporting and Data System) scoring guides targeted biopsy decisions. AI aims to:
- Detect suspicious lesions on MRI
- Assign PI-RADS-equivalent scores
- Reduce inter-reader variability
- Improve clinically significant prostate cancer (csPCa) detection
PI-RADS Steering Committee Standards
The PI-RADS Steering Committee described considerations for developing and reporting AI in prostate MRI (Turkbey et al., 2025). The document is a methodological framework, not an FDA authorization or a universal clinical performance threshold.
Performance reporting:
- Define the intended task, clinically significant cancer endpoint, and decision threshold
- Compare with an appropriate radiologist or clinical reference when that comparator matches the intended use
- Report discrimination, precision-recall behavior, calibration, and threshold-specific performance as appropriate
Reporting requirements:
- Training data composition and demographics
- Biopsy correlation methodology
- External validation in independent populations
- Specific failure mode analysis
Clinical context: - State whether the population is biopsy-naive, previously biopsied, under surveillance, or otherwise selected - Define clinically significant cancer and the histopathologic reference standard - Describe how the output would alter biopsy, surveillance, or review decisions
Current AI Performance
Published performance varies with the model, MRI sequence, threshold, biopsy selection, and reference standard. In one biparametric MRI study, an AI system reported 88.4% sensitivity for clinically significant prostate cancer, compared with 89.5% for radiologists (Belue et al., 2025). That comparison does not establish equivalence, superior biopsy decisions, or improved outcomes outside the evaluated cohort.
Other research reports have evaluated PI-RADS classification and lesion-level mpMRI discrimination, but those figures should not be pooled into a generic “prostate MRI AI” accuracy. The relevant unit may be a patient, lesion, prostate sector, image series, or biopsy decision, and those units are not interchangeable.
External Validation Challenges
External performance can differ from development-site performance. A defensible evaluation reports both results with uncertainty and does not imply that every tool degrades by the same amount.
Factors affecting performance: - MRI quality (especially diffusion-weighted imaging) - Scanner differences - Protocol variations - Case selection and prevalence - Annotation and biopsy-reference differences
Clinical Implementation
Prostate MRI AI is best positioned for: - Second-read quality assurance - Lesion detection in high-volume practices - Training and education
Not ready for: - Autonomous PI-RADS scoring without radiologist review - Replacement of urologic clinical judgment
Bladder Cancer AI
Bladder cancer AI has progressed to multicenter diagnostic validation. One company has reported FDA Breakthrough Device Designation for its urine-test program, but a designation is not marketing authorization.
Cystoscopy AI:
A multicenter diagnostic study of the CAIDS system used 69,204 cystoscopy images from 10,729 patients across six hospitals and reported 93.9% accuracy and 95.4% sensitivity in a defined validation analysis (Wu et al., 2022). The study evaluated image-based diagnosis, not prospective patient outcomes. An AUA 2025 conference report described improved detection of small, flat, and carcinoma-in-situ lesions with a lightweight model, but conference reporting should be labeled preliminary until a full peer-reviewed record supports the exact design and endpoints (EUS/AUA 2025).
AI versus clinician performance: A 2025 meta-analysis reported pooled AI sensitivity and specificity of 83% and an AUC of 0.89, compared with pooled clinician sensitivity and specificity of 78% and AUC of 0.81 in the included comparisons (World Journal of Urology, 2025). Pooled retrospective diagnostic performance does not establish autonomous use or improved recurrence, progression, or survival.
Risk stratification: A 2025 conference report described the PROGRxN-BCa model in a non-muscle-invasive bladder cancer cohort of 12,659 patients and reported higher c-index values than established risk models, particularly for high-grade Ta disease (AUA 2025 report). The claim remains conference-level evidence until the full methods, temporal validation, censoring, calibration, and peer-reviewed results can be examined.
TOBY urine test: A June 2025 company release reported FDA Breakthrough Device Designation and internal AUC above 0.9 for a volatile-organic-compound urine test (TOBY company release, 2025). The statement is company-reported because a public primary FDA designation record was not identified. The test is not FDA-authorized for routine clinical use on the evidence cited here.
Kidney Stone AI
Intraoperative stone detection (AiFURS):
The AiFURS system provides real-time detection, classification, and measurement of kidney stones during flexible ureteroscopy. Clinical validation (100 in vivo cases, 80 external validation cases) demonstrated diagnostic accuracy of 92.2–95.3% in vivo and 86.8–92.2% on external validation, outperforming expert surgeons in patient-level stone type prediction (npj Digital Medicine, 2025).
CT-based stone composition: Research models use CT features to estimate stone composition. The earlier 80–90% range combined heterogeneous studies and should not be used as a universal accuracy claim. Any treatment-selection claim requires external validation, calibration, and evidence that using the estimate improves the choice among observation, medical therapy, shock-wave lithotripsy, ureteroscopy, and percutaneous procedures.
Simulation and robotics: A Johnson & Johnson announcement described collaboration with NVIDIA on simulation and digital technologies for the Monarch platform (J&J, 2025). This is a development announcement, not peer-reviewed evidence that an AI component improves kidney-stone outcomes. A digital twin, simulator, teleoperated robot, and clinical machine-learning system are different categories.
Urologic Robotic Surgery
The urologic robotic-surgery landscape includes several platforms, but a robot is not an AI system by default.
Hugo RAS (Medtronic/Covidien): FDA cleared the Class II modular electromechanical system in December 2025 for specified adult minimally invasive urologic procedures. The FDA record describes a 137-subject U.S. study across six sites (FDA K250725). That clearance is not evidence of an AI-assisted clinical outcome.
Focal One HIFU: Trade coverage has described algorithmic tissue-ablation visualization and treatment evaluation. Before labeling the function AI, the clinical team should inspect the exact FDA record and labeling for the cleared version and identify the documented machine-learning component, if any.
Current status: Research systems evaluate tissue recognition, surgical-phase identification, and video-derived quality metrics. Teleoperation, image registration, and programmed energy delivery should not be labeled AI without a separately evaluated learning system. Autonomous AI surgery is not established for routine urologic care.
Part 3: Otolaryngology AI
### AAO-HNS Task Force Report (2025)
The American Academy of Otolaryngology-Head and Neck Surgery Task Force published a specialty report on AI integration in the February 2025 issue of Otolaryngology–Head and Neck Surgery (Ayoub et al., 2025). It provides consensus principles and professional context, not product-specific clinical evidence or a clinical practice guideline.
Identified applications:
- Precision medicine in head and neck cancer
- Clinical decision support
- Operational efficiency (scheduling, documentation)
- Research and education tools
Key challenges:
- Data quality and bias
- Health equity concerns
- Privacy and security
- Regulatory gaps
- Ethical considerations
Recommendations:
- Careful validation before clinical deployment
- Attention to health equity implications
- Transparency in AI decision-making
- Specialty-specific training data development
Head and Neck Cancer AI
Oropharyngeal cancer multimodal prognostication:
A multimodal fusion framework (SMuRF) integrating CT imaging of the primary tumor and lymph nodes with whole-slide pathology images predicted disease-free survival and tumor grade in HPV-associated oropharyngeal squamous cell carcinoma (n=277). The model achieved c-index of 0.81 (development) and 0.79 (test) for disease-free survival, functioning as an independent prognostic biomarker with a hazard ratio of 17 (95% CI 4.9–58, p<0.0001) after controlling for clinical variables (Song et al., 2025). This represents the first study combining radiology and pathology imaging for biomarker discovery in oropharyngeal cancer.
### Surgical Navigation and Safety Signals
The TruDi Navigation System is an image-guided ENT surgical-navigation device. FDA records describe electromagnetic tracking, patient registration, and display of instrument position relative to CT or MR anatomy, not a documented machine-learning function (FDA K231862). A 2023 Class II recall for an earlier software version and specified curette concerned a discrepancy between the actual curette-tip location and the displayed location (FDA recall Z-0127-2024).
Potential harms identified in the recall included:
- Delayed or prolonged surgery
- Cerebrospinal fluid leak
- Visual impairment
- Skull-base structural damage
Investigative reporting and litigation may contain additional event narratives, but a reported sequence of malfunctions does not by itself prove that an AI algorithm was introduced or caused the event. The safety review should ask:
- Which device and software version was used?
- Was the failure in tracking, registration, hardware, workflow, or a learning model?
- Does the report provide a denominator and an independent causal assessment?
- Were the event, recall, correction, and current labeling linked to the same configuration?
Clinical implication: Surgical navigation requires rigorous validation regardless of whether it uses AI. Do not convert a navigation-device event into an AI-causation claim without evidence connecting the component, failure mechanism, and event.
Hearing AI Applications
Smartphone-based audiometry:
Smartphone tools can extend hearing screening outside traditional audiology settings, but the underlying method may be conventional signal processing rather than AI: - Direct-to-consumer apps for hearing self-assessment - School and community screening programs - Remote monitoring for hearing aid users
Evidence:
- Some smartphone audiometry systems correlate with standard audiometry in controlled settings
- Real-world performance varies with ambient noise and user technique
- Does not replace comprehensive audiologic evaluation for diagnosis
Age-related hearing loss (ARHL):
The AAO-HNS Clinical Practice Guideline on ARHL (2024) provides context: - ARHL affects 1 in 3 adults age 65-74 - Associated with dementia, depression, falls - Validated remote screening could expand early detection
Upgradeable cochlear implant technology:
The Cochlear Nucleus Nexa System is marketed with upgradeable firmware and implant memory. Those features may matter for lifecycle management, but they do not establish that the implant is an AI system. The original July 2025 approval claim did not match the FDA supplement record located during review and has therefore been removed. Product evidence should cite the exact PMA supplement and approved change before making a regulatory statement.
AI-predicted cochlear implant outcomes:
A multicenter cohort study of 278 children across the United States, Australia, and Hong Kong used deep transfer learning on preimplantation MRI and clinical features to classify higher versus lower spoken-language improvement. The combined model reported 92.39% accuracy, 91.22% sensitivity, and 93.56% specificity (Wang et al., 2026). The binary labels used different language measures across centers, cross-center generalization was limited, and the study did not test model-guided therapy.
Hearing aid optimization:
AI powers automatic adjustment of hearing aids based on: - Acoustic environment detection - User preferences and listening patterns - Real-time speech enhancement
Voice Analysis AI
Applications:
- Vocal cord pathology detection
- Analysis of voice recordings for nodules, polyps, paralysis
- Screening for laryngeal cancer
- Speech therapy monitoring
- Objective voice quality measures
- Treatment response tracking
- Neurological voice changes
- Parkinson’s disease voice biomarkers
- Stroke-related dysarthria assessment
Status: These examples remain largely research-stage. An absolute claim that no FDA-authorized voice product exists can become stale and should be replaced with product-specific verification of the exact intended use and record.
Voice recordings can also function as biometric data. A model-development or documentation workflow should assess reidentification risk, consent, secondary use, storage, and whether the receiving vendor is approved for the data class. Privacy and documentation controls are reviewed in AI-Assisted Clinical Documentation.
Sleep Apnea Screening
AI tools analyze: - Snoring patterns from audio recordings - Movement data from wearables - Oximetry trends
Evidence boundary: Audio, wearable, and oximetry models evaluate different screening tasks, thresholds, and populations. The earlier 80–90% range was not adequately sourced as a class-wide estimate. Model-specific prospective evidence and a confirmatory pathway are required; a screening output is not a polysomnographic diagnosis.
Part 4: Professional Society Positions
Gastroenterology Societies
American Gastroenterological Association (AGA):
The AGA published its first guideline critically evaluating AI in GI care in 2025:
- AGA Living Clinical Practice Guideline on CADe-Assisted Colonoscopy (Sultan et al., 2025)
- Makes no recommendation for or against CADe-assisted colonoscopy due to very low certainty of evidence regarding cancer outcomes
- Acknowledges modest ADR improvement (44.8% vs. 37.4%) but questions clinical significance
- Raises concern about overdiagnosis: 635 additional surveillance colonoscopies per 10,000 patients
The AGA Clinical Practice Update on AI in Polyp Diagnosis (Samarasena et al., 2023) provides additional context on polyp characterization AI.
American College of Gastroenterology (ACG):
The official ACG guideline collection did not identify a standalone AI clinical guideline during this review. The 2024 ACG/ASGE Quality Indicators for Colonoscopy (Rex et al., 2024):
- Establishes ADR benchmarks for defined screening populations and quality programs
- Does not include AI or CADe as a quality indicator
- Focuses on technique-based metrics: withdrawal time (≥8 minutes), bowel preparation adequacy (≥90%)
American Society for Gastrointestinal Endoscopy (ASGE):
The ASGE Position Statement on Priorities for AI in GI Endoscopy is a direct society policy source. Educational sessions or task-force activity should not be described as clinical guidance without a linked final document.
American Urological Association (AUA)
The AUA has no standalone AI clinical practice guideline. Review of official AUA Policy and Position Statements confirms no AI-specific policy document. However, the AUA has increased AI-related advocacy and education:
- Advocacy position (2024): The AUA recognizes AI as “inevitably integral to health care” and has identified strategic initiatives including education on AI use/misuse, incorporation into AUA committee activities, and infrastructure assessment (AUA Advocacy, June 2024)
- Annual meeting programming: AUA 2025 featured dedicated AI courses including “Practical AI for Practicing Urologists” and multiple AI-focused abstract sessions (AUA 2025)
Prostate MRI guidance: The AUA/SUO Early Detection of Prostate Cancer Guideline (2023) mentions AI only in Future Directions: “evolving MRI protocols, such as biparametric MRI and use of artificial intelligence, requires further study.” AI is not recommended as a current clinical adjunct.
Provider sentiment: A 2025 survey of urology healthcare providers found 83.4% believed AI will improve efficiency, but 82% expressed concerns about technical reliability and 76% worried about diagnostic errors from generative AI (Healthcare, 2025).
Content restrictions: The AUA Privacy Policy explicitly prohibits uploading AUA content into AI systems for training purposes.
Surgical Subspecialty Society Positions
American Academy of Orthopaedic Surgeons (AAOS):
- Position Statement on Artificial Intelligence (Document #1193, February 2025)
- Addresses physician understanding of AI benefits, risks, and ethical considerations
- Emphasizes need for socio-economic awareness of AI integration
American Association of Neurological Surgeons (AANS)/Council of State Neurosurgical Societies (CSNS):
- Policy Statement on the Use of AI in Neurosurgery (2025)
- Five core domains: responsible use, privacy/security, transparency, academic integrity, FDA/IRB oversight
- Key position: AI should augment, not replace, human decision-making
Society of American Gastrointestinal and Endoscopic Surgeons (SAGES):
- Defining Digital Surgery White Paper (Ali et al., 2024)
- Consensus Recommendations on Surgical Video Data (2023) for AI research standards
- Active AI Task Force work on video recording for machine learning
Society of University Surgeons:
- Position Statement on AI in Surgical Training (Kewalramani et al., 2026)
- Addresses AI literacy and prompt engineering as foundational competencies
- Warns of cognitive off-loading and “deskilling” risks in surgical education
Societies Without Formal AI Position Statements
Several surgical societies have educational resources but no formal AI policy document identified in their public guideline or position-statement collections at the time of this review:
- American College of Surgeons (ACS): Informatics and AI Committee, educational programming only
- Society of Thoracic Surgeons (STS): Educational content on AI/ML, no formal position
- American Society of Colon and Rectal Surgeons (ASCRS): Educational webinars only
- American Society of Plastic Surgeons (ASPS): Educational articles, no formal position
Cross-Specialty Themes
Across surgical societies with formal positions, consistent themes emerge:
- AI as adjunct, not replacement, for surgical judgment
- Specialty-specific validation required before deployment
- Human oversight of AI recommendations mandatory
- Academic integrity standards for AI-generated content
- Concerns about training data bias and equity
Clinical Scenarios
These are hypothetical teaching scenarios. The patients, products, outputs, local economics, and decisions are illustrative. They are not reports of actual care, legal outcomes, or validated product behavior. Local policy, labeling, specialist interpretation, and patient preferences govern real decisions.
Case: During a screening colonoscopy with CADe, the AI system generates an alert highlighting a mucosal fold. The endoscopist examines the area and determines it is a false positive. This is the fourth false alert during this procedure.
Question: How should the endoscopist manage alert fatigue while maintaining detection quality?
Discussion
Understanding alert fatigue:
CADe systems have high sensitivity, meaning they detect most polyps but also generate false positives for: - Mucosal folds - Stool particles - Artifacts - Vascular patterns
Appropriate response:
- Follow the validated workflow: The system’s labeling, local protocol, and endoscopist assessment determine the response
- Document when locally required: Structured override reasons can support quality tracking if the institution has defined that workflow
- Maintain technique: CADe supplements but does not replace systematic inspection
- Provide feedback: Some systems allow false positive marking to improve algorithms
What not to do:
- Disable or ignore alerts without following the approved local protocol
- Rely solely on the model rather than systematic inspection
- Reduce withdrawal time because the system is active
Case: A 62-year-old man with elevated PSA undergoes multiparametric prostate MRI. The radiologist assigns PI-RADS 3 (equivocal). An AI second-read system identifies the same lesion and assigns it as high-risk (equivalent to PI-RADS 4).
Question: How should the urologist interpret this discordance?
Discussion
Understanding discordance:
PI-RADS 3 is an equivocal imaging category. The probability of clinically significant cancer varies by population, lesion, PSA density, MRI quality, biopsy strategy, and study design. The model may also use a threshold that does not map directly to the radiologist’s PI-RADS category.
Factors to consider:
- AI validation: Was this AI system validated on similar patient populations?
- Clinical context: PSA density, prior biopsy results, family history
- Lesion characteristics: Location, size, DWI signal
- Patient preferences: Risk tolerance for biopsy vs. active surveillance
Possible approaches:
- Discuss discordance with radiologist
- Consider targeted biopsy (MRI-TRUS fusion or cognitive targeting)
- Repeat MRI if quality concerns
- PSA density and other biomarkers for risk stratification
What not to do:
- Automatically defer to AI over radiologist
- Ignore AI finding without consideration
- Proceed to saturation biopsy without targeted approach
Case: A 68-year-old patient shows you results from a smartphone hearing screening app indicating moderate hearing loss. They ask if they need hearing aids.
Question: How should you counsel this patient about the app results?
Discussion
Smartphone audiometry limitations:
- Ambient noise affects results
- Headphone quality varies
- Calibration may not match clinical audiometers
- Cannot assess word recognition, speech-in-noise, or middle ear function
Appropriate response:
- Validate concern: The app results suggest possible hearing loss worth evaluating
- Recommend formal testing: Refer to audiology for comprehensive evaluation
- Discuss ARHL: Age-related hearing loss is common and treatable
- Manage expectations: App results may overestimate or underestimate actual loss
Audiologic evaluation includes:
- Pure tone audiometry in sound-treated booth
- Speech recognition testing
- Tympanometry for middle ear function
- Hearing aid candidacy assessment
When apps are valuable:
- Motivating patients to seek evaluation
- Monitoring known hearing loss between visits
- Screening in resource-limited settings
Case: You are a GI division chief evaluating whether to purchase a CADe colonoscopy system. The sales representative presents RCT data showing 15% relative improvement in ADR. Your division’s current mean ADR is 45%.
Question: What factors should inform this decision?
Discussion
Evaluating the evidence:
The RCT data is promising, but consider:
Baseline ADR matters: A local mean ADR of 45% must be interpreted against the population, indications, sex distribution, exclusions, and current quality standard. Improvement may be smaller in high-performing settings.
Real-world versus RCT performance: A 2024 real-world meta-analysis found no significant ADR difference in the GI Genius subgroup (RR 0.96), while newer pragmatic randomized trials have also reported null effects. Implementation conditions differ from efficacy trials.
What improves: Primarily small adenomas and sessile serrated lesions. Advanced adenoma detection may not change.
What increases: Non-neoplastic polypectomy (false positives leading to unnecessary removal).
Cost-benefit analysis:
- Total local cost, including acquisition or lease, compatible hardware, service, cybersecurity, training, and any per-procedure charges
- Procedure time: May increase slightly
- Reimbursement and revenue assumptions verified with current payer policy rather than vendor projections
- Quality metrics: Potential ADR improvement affects reporting
Implementation requirements:
- Workflow integration
- Endoscopist training
- IT support
- Quality monitoring to verify benefit
Recommendation:
- Honest assessment of current quality gaps
- Pilot period with outcome tracking
- Focus on technique improvement alongside technology
- Consider centers with lower baseline ADR as priority
Does colonoscopy AI improve adenoma detection?
Across 44 randomized trials, computer-aided detection increased adenoma detection from 36.7% to 44.7%, but certainty about colorectal cancer outcomes remains low and a 2026 private-practice trial found no significant adenoma-detection difference.
What colonoscopy AI systems are FDA cleared?
Several product-specific systems have FDA records, including GI Genius, CAD EYE, EndoScreener, and ENDO-AID. The exact model, intended use, compatible hardware, and current authorization record should be verified before procurement.
How accurate is prostate MRI AI?
Some prostate MRI models have reported radiologist-comparable performance for a defined dataset and threshold, but results vary by scanner, protocol, population, and reference standard. Diagnostic comparisons do not establish better biopsy decisions or outcomes.
Why is there a gap between colonoscopy AI RCT results and real-world performance?
Differences in case mix, baseline adenoma detection, endoscopist behavior, protocol adherence, and workflow can change the measured effect. A local implementation should be evaluated against a prespecified comparator rather than assuming the pooled trial effect will transfer.
What AI tools are FDA-authorized for prostate cancer?
FDA records describe distinct tasks. Paige Prostate assists pathology detection, Galen Second Read flags initially benign cases for additional review, Avenda software supports evaluation and planning for known disease, and ArteraAI Prostate provides image-based 10-year prognostic risk estimates.
How accurate is AI for bladder cancer detection?
A 2025 meta-analysis reported pooled sensitivity and specificity of 83% for evaluated models, while a separate multicenter CAIDS study reported 93.9% accuracy and 95.4% sensitivity. These findings are study-specific and do not describe every bladder cancer AI system.
Can AI detect liver cirrhosis from a routine ECG?
A 2025 pragmatic cluster-randomized trial found more new advanced chronic liver disease diagnoses with AI-enabled ECG screening than usual care, 1.0% versus 0.5%. It tested a screening and follow-up workflow, not ECG diagnosis alone.
Does upgradeable cochlear implant firmware make a device an AI system?
No. Upgradeable firmware and internal memory are device features, not proof that a machine-learning component performs a clinical task. AI claims require a defined model, intended use, evidence, and applicable regulatory record.
Key Takeaways
Gastroenterology AI:
- Colonoscopy CADe has the strongest randomized evidence base, but pragmatic effects vary and direct colorectal cancer outcomes remain untested
- AI-enabled ECG screening increased new advanced chronic liver disease diagnoses in a pragmatic cluster-randomized trial
- Celiac pathology and GRAPE gastric cancer models have multicenter diagnostic evidence, not patient-outcome trials
- NaviCam ProScan has product-specific De Novo authorization for an adjunctive capsule-reading workflow; standalone sensitivity and specificity were not established
- Reinforcement-learning anesthesia control produced similar hypoxemia rates to clinician-managed care in a selected low-risk 2026 trial
- AGA makes no recommendation for or against CADe colonoscopy because certainty about cancer outcomes is very low
Urology AI:
- Four authorized prostate products have distinct roles: Paige detection, Galen second review, Avenda evaluation and planning for known disease, and Artera image-based prognosis
- Artera used randomized-trial cohorts for development and validation, but AI-assisted management was not itself tested in a randomized trial
- Bladder diagnostic models have promising pooled and multicenter accuracy; patient-benefit evidence is not established
- TOBY designation and performance claims are company-reported, not marketing authorization or peer-reviewed effectiveness evidence
- AiFURS results apply to the evaluated intraoperative stone-classification task
- Hugo RAS is an FDA-cleared robotic platform, not evidence of an AI-assisted outcome
- Transportability must be assessed for each model, protocol, scanner, and clinical population
Otolaryngology AI:
- The 2025 AAO-HNS Task Force report provides specialty principles, not a product-selection guideline
- Upgradeable firmware and internal memory do not by themselves make a cochlear implant an AI system
- A 2026 multicenter prognostic study reported 92.39% accuracy for classifying higher versus lower language improvement; no model-guided therapy was tested
- Oropharyngeal cancer multimodal AI reported a c-index of 0.81 in development and 0.79 in testing for survival prediction
- TruDi is a navigation system with an official device-version recall; the located FDA records do not establish that a machine-learning component caused reported failures
- Voice analysis and sleep-apnea screening models remain research-stage examples requiring task-specific validation
Implementation principles:
- Real-world performance may differ from RCT results in either direction
- AI supplements but does not replace surgical skill and judgment
- Do not label conventional navigation, teleoperation, or upgradeable firmware as AI without a documented learning component
- Specialty-specific validation is essential
- Monitor alert burden, override behavior, downstream intervention, and patient-relevant outcomes