Surgery, Anesthesiology, and Perioperative Care
Surgery combines technical skill, anatomy, physiology, team coordination, and time-sensitive decisions in settings where some actions are irreversible. AI applications span perioperative risk estimation, planning, video analysis, robotic-platform analytics, selected intraoperative imaging, and postoperative monitoring. Evidence must distinguish prediction from intervention, navigation from AI, teleoperation from autonomy, and device authorization from outcome benefit. An accurate retrospective model is not yet a safe surgical intervention.
After reading this chapter, readers will be able to:
- Evaluate AI systems for surgical risk prediction and optimization
- Understand computer vision applications in robotic and minimally invasive surgery
- Create evaluation plans for robotic surgery platforms and surgical video AI
- Assess AI tools for surgical phase recognition and workflow analysis
- Navigate AI-assisted surgical planning and simulation
- Identify postoperative complication prediction systems
- Recognize limitations and failure modes of surgical AI
- Balance AI augmentation with surgical judgment and technical skill
Introduction
Surgery stands apart from other medical specialties in its immediacy, irreversibility, and technical demands. While radiologists can analyze images over minutes, surgeons make split-second decisions with scalpel in hand. While internists can adjust management based on patient response, surgical decisions, once made, cannot be easily undone.
This unique context shapes how AI can and cannot help surgeons. The most promising applications assist with the cognitive work surrounding surgery (risk assessment, planning, outcome prediction) rather than replacing the surgeon’s hands or judgment during the operation itself.
The sections that follow cover surgical AI applications across the perioperative spectrum, from preoperative optimization through postoperative care.
Preoperative AI Applications
Surgical Risk Prediction
The Clinical Problem:
Surgeons face a fundamental question before every operation: Will this patient tolerate this procedure? Traditional risk assessment relies on clinical judgment supplemented by scoring systems (ASA classification, NSQIP risk calculator, RCRI for cardiac risk in non-cardiac surgery). These tools have limitations:
- Incorporate limited variables (20-30 factors)
- Use linear models that miss complex interactions
- Provide population-level estimates, not personalized predictions
- Updated infrequently as new evidence emerges
Machine Learning Solutions:
Modern ML approaches improve risk prediction by:
- Analyzing larger feature sets: 100+ variables from EHR, imaging, labs, medications, vital signs, social determinants
- Capturing nonlinear relationships: Age × frailty × procedure complexity interactions
- Continuous learning: Models updated with new outcome data
- Personalized predictions: Patient-specific risk estimates rather than population averages
Evidence:
The original MySurgeryRisk study developed and validated models in a large single-center surgical cohort (Bihorac et al., 2019):
- Mortality prediction: AUC 0.77–0.83 across the evaluated mortality horizons
- Major complications: AUC 0.82–0.94 across eight complication types
- Classification boundary: low-prevalence outcomes had lower positive predictive values despite high negative predictive values
A 2026 retrospective, longitudinal multicenter cohort analysis then applied the framework to 508,097 major inpatient surgical encounters from 366,875 patients at 14 institutions. In the validation data, AUROC was 0.93 for ICU admission, 0.94 for postoperative mechanical ventilation, 0.92 for acute kidney injury, and 0.95 for in-hospital mortality. Positive predictive value was substantially lower, including 0.35, 0.25, 0.25, and 0.04 for those respective outcomes. The models were developed and validated inside the OneFlorida+ network, not prospectively deployed to guide care, and the authors identified data drift and subgroup blind spots (Ren et al., 2026). High AUROC can coexist with low positive predictive value and no evidence that displaying the score improves patient outcomes.
Clinical Applications:
Preoperative Optimization:
- Identify modifiable risk factors (anemia, hyperglycemia, nutritional deficits)
- Triage patients for preoperative clinic vs. day-of-surgery admission
- Guide prehabilitation referrals
Shared Decision-Making:
- Provide personalized risk estimates during surgical consults
- Facilitate discussions about alternative treatments
- Support goals-of-care conversations for high-risk patients
Resource Allocation:
- Predict ICU vs. floor bed requirements
- Identify patients needing enhanced postoperative monitoring
- Optimize OR scheduling based on predicted case duration
Quality Improvement:
- Risk-adjust outcome comparisons between surgeons/hospitals
- Identify outliers for focused improvement efforts
- Benchmark performance against predicted outcomes
Critical Limitations:
Risk calculators should inform, not dictate, surgical decisions:
- Algorithms miss important factors: patient goals, functional trajectory, social support, frailty nuances
- High-risk patients may still benefit from surgery if alternative is certain poor outcome
- Low-risk predictions do not guarantee good outcomes
- Models trained on one population may not generalize to different populations
Clinical Bottom Line: Use risk prediction AI to enhance shared decision-making and optimize preoperative preparation. Do not deny surgery based solely on algorithmic risk scores.
Randomized Model-Guided Surgical Decision Evidence
The Risk-Guided Temporary Ileostomy Decision system provides a stronger test than retrospective discrimination alone. In a randomized trial of 872 patients with stage I–III rectal cancer undergoing anterior resection, 750 remained in the final analysis after 122 post-randomization exclusions for postponement, cancellation, or intraoperative protocol deviations. Temporary diverting ileostomy was performed in 18.6% of the model-guided group and 40.5% of the surgeon-discretion group (P < .001). The authors classified 17.7% versus 41.3% as unnecessary stoma formation under their decision framework. Anastomotic leakage occurred in 2.4% versus 2.7% (P = .753) (Shao et al., 2026).
The safety interpretation requires restraint. The study was underpowered to establish noninferiority for anastomotic leakage, only 19 leaks occurred, the model’s overall positive predictive value for leakage was 7.6%, and the post-randomization exclusions complicate the intention-to-treat interpretation. The trial supports reduced stoma use in its protocol, not proof that model guidance is equally safe across institutions or procedures. External replication should preserve surgeon override, examine leak severity and reoperation, and report the consequences of false-low and false-high risk classifications.
Preoperative Planning and Simulation
AI-Assisted Anatomical Segmentation:
Surgical planning for complex cases (oncologic resections, liver surgery, orthopedic reconstructions) traditionally requires manual analysis of CT/MRI to identify anatomy, plan approaches, and anticipate challenges. AI automates and enhances this process:
Applications:
Oncologic Surgery:
- Tumor segmentation and volumetry
- Relationship to critical structures (vessels, bile ducts, nerves)
- Predicted resection margins
- Assessment of resectability
Liver Surgery:
- Vascular and biliary anatomy mapping
- Liver volumetry for donation or resection planning
- Future liver remnant calculation
- Virtual hepatectomy simulation
Orthopedic Surgery:
- Joint replacement planning (alignment, component sizing)
- Osteotomy planning for deformity correction
- Fracture reduction simulation
- Bone tumor resection planning
Neurosurgery:
- Brain tumor segmentation and eloquent cortex mapping
- Surgical approach trajectory planning
- Vascular anatomy for aneurysm clipping
- Epilepsy focus localization
Evidence boundary:
Segmentation research demonstrates task performance in selected datasets, but the cited surgical-AI review does not establish a class-wide 60% to 80% reduction in planning time, expert-equivalent reliability, or improved outcomes (Hashimoto et al., 2018). A segmentation metric is not evidence that a surgical plan is safe, feasible, or outcome-improving. Planning-time, margin, operative-time, complication, and patient-understanding claims require their own comparative evidence.
Orthognathic soft-tissue outcome prediction
Orthognathic jaw repositioning requires a postoperative facial soft-tissue estimate, not only bony segmentation. PhysSFI-Net, a physics-informed geometric deep-learning model trained on 135 patients and externally validated on 33, reported internal global shape error of 1.070 ± 0.088 mm, surface deviation of 1.296 ± 0.349 mm, and landmark error of 2.445 ± 1.326 mm, with external global shape error of 1.431 ± 0.087 mm, outperforming ACMT-Net and other baselines on those geometric metrics (Bao et al., 2026). These values are retrospective millimeter-scale shape errors against postoperative ground truth, not planning-time, occlusion, complication, or patient-reported outcome evidence, and they do not justify autonomous surgical planning.
Limitations:
- Segmentation errors can propagate to surgical plans (always verify)
- Quality depends on input imaging (motion artifacts, contrast timing)
- Does not account for intraoperative findings (adhesions, variant anatomy)
- Most effective for anatomy-driven procedures with good imaging
3D Printing and Surgical Models:
AI-segmented anatomy can be converted to 3D-printed models for:
- Pre-surgical rehearsal of complex cases
- Patient education and consent
- Trainee education
- Custom surgical guides and implants
Clinical Impact: Mixed. Some studies show reduced operative time and improved outcomes for complex cases; others show no benefit beyond surgeon confidence. Cost and workflow integration remain barriers to widespread adoption.
Intraoperative AI Applications
Computer Vision in Minimally Invasive Surgery
The laparoscope and robotic camera create continuous video streams, ideal data for computer vision AI. Applications range from documentation to real-time guidance, with varying degrees of validation and clinical readiness.
Surgical Phase Recognition:
What it does: AI analyzes surgical video and identifies current phase (e.g., “dissection of gallbladder from liver bed” in laparoscopic cholecystectomy)
How it works: Deep learning models trained on annotated surgical videos learn to recognize instrument configurations, anatomical landmarks, and surgeon actions characteristic of each phase.
Performance:
- Accuracy approximately 82% for laparoscopic cholecystectomy phase recognition (Twinanda et al., 2017)
- Works across multiple procedures (bariatric, colorectal, gynecologic)
- Real-time capability (15-30 frames/second)
Potential applications:
- Context-aware instrument tracking
- Automated surgical documentation
- OR efficiency analysis
- Surgical skill assessment
- Adverse event detection
Current status: Primarily research tool. Limited clinical deployment because phase recognition alone does not provide actionable guidance. Surgeons already know which phase they’re in.
Future potential: Phase recognition is foundational for more advanced applications (predictive alerts, context-aware instrument suggestions).
Anatomical Structure Recognition:
The promise: Computer vision identifies critical anatomy (bile ducts, ureters, vessels) to prevent surgical injury.
The reality: This is extraordinarily difficult and not yet clinically reliable.
Why it’s hard:
- Visual variability: Blood, smoke, retraction, lighting changes, cautery artifacts
- Anatomical variants: Textbook anatomy is the exception, not the rule
- Dynamic deformation: Tissue moves, stretches, changes appearance continuously
- Occlusion: Critical structures often partially hidden
- Context-dependence: What looks like ureter may be vessel or adhesion band
Current evidence:
Research systems report structure-recognition performance in selected video datasets. Blood, smoke, occlusion, inflammation, tissue deformation, off-axis views, and unfamiliar anatomy can change the input distribution. No source cited here supports a universal 70% to 85% accuracy range or an “acceptable” class-wide error rate. Dataset accuracy does not authorize an irreversible intraoperative action.
Critical safety concern:
Surgeons cannot rely on AI to definitively identify critical structures. Visual confirmation, tactile feedback, anatomical knowledge, and methodical dissection remain essential. AI suggesting “safe to divide this structure” is not acceptable with current technology.
More promising near-term application:
Warning systems: AI detecting absence of expected structures (“ureter not identified in expected location, double-check before dividing anything”) may be safer than positive identification. Alert surgeons to uncertainty rather than provide false confidence.
AI in Robotic Surgery
Robotic surgery belongs inside surgical AI evaluation, not as a separate specialty. It is a platform category, a surgical training problem, a data-capture layer for surgical AI, and a possible future path for supervised autonomy. Evaluation must separate the robot, the surgeon, the procedure, the AI layer, and the local health-system context.
Current state: teleoperation, not autonomy
Current clinical soft-tissue robotic platforms are teleoperated minimally invasive systems. The surgeon controls instrument motion from a console. The platform may provide three-dimensional visualization, wristed instruments, tremor filtering, motion scaling, ergonomic benefits, and integrated data capture, but those features do not equal autonomous surgical judgment.
The installed base and procedure volume are now large enough that robotic surgery requires routine service-line governance rather than innovation-lab treatment. Intuitive Surgical reported approximately 3.153 million da Vinci procedures in 2025, compared with approximately 2.683 million in 2024, and reported a system installed base of more than 12,100 by year-end 2025 (Intuitive Surgical, 2026 annual report).
FDA status is platform-specific
FDA clearance or authorization applies to specific devices, indications, and intended uses. It does not validate every hospital’s use case, every surgeon’s learning curve, or every AI claim layered on top of the robotic platform.
| Platform | FDA status | Current relevance |
|---|---|---|
| da Vinci 5 | 510(k) clearance, K232610, March 2024 | Fifth-generation multiport da Vinci system. Clearance supports the platform’s intended use, not autonomous surgery (FDA K232610). |
| Versius Surgical System | De Novo authorization, DEN230078, October 2024 | Class II modular electromechanical surgical system, initially indicated in the United States for adult cholecystectomy (FDA DEN230078). |
| Hugo RAS System | 510(k) clearance, K250725, December 2025 | Class II modular electromechanical surgical system for adult minimally invasive urologic surgical procedures (FDA K250725). |
FDA status should be verified from FDA records before procurement, credentialing, patient-facing materials, or handbook updates. Vendor announcements can describe launch strategy and training ecosystems, but FDA records define the cleared indication.
Commercial landscape and evidence status
Commercial activity now falls into distinct evidence categories. FDA-authorized teleoperated platforms, including da Vinci, Versius, and Hugo, should be evaluated by device indication and procedure-specific outcomes. Surgical video and operating-room analytics companies, including OR Black Box, Touch Surgery, Theator, and Caresyntax, should be evaluated as data-capture, documentation, video-review, or quality-measurement systems; peer-reviewed evidence supports feasibility and implementation analysis, but not automatic outcome improvement (Thornton et al., 2026; Aklilu et al., 2024). Autonomy-first startups, including Aleph Surgical, signal where the market is looking next, but Aleph’s March 2026 research preview is not peer-reviewed clinical evidence and no FDA authorization or clinical outcomes publication was located as of June 18, 2026. Treat these companies as systems to track, not as evidence of clinical readiness.
Recent clinical evidence
Robotic surgery does not have one evidence grade. Each procedure has its own learning curve, comparator, outcomes, and cost structure.
| Evidence question | Recent finding | Evidence quality |
|---|---|---|
| Middle and low rectal cancer | The REAL randomized trial included 1171 patients with median 43-month follow-up and found lower 3-year locoregional recurrence with robotic surgery than laparoscopy (1.6% vs. 4.0%) and higher disease-free survival (87.2% vs. 83.4%), with similar overall survival (Feng et al., 2025). | Moderate to high: randomized trial, but high-volume expert centers in China limit generalization. |
| Acute-care cholecystectomy | A propensity-matched JAMA Surgery cohort found similar bile duct injury rates for robotic-assisted and laparoscopic cholecystectomy, but higher major postoperative complications with robotic-assisted surgery (8.37% vs. 5.50%), more drain use, and longer length of stay (Woldehana et al., 2025). A newer JAMA Surgery analysis examined comparative safety in contemporary practice, underscoring that cholecystectomy safety should be monitored with current local outcomes rather than platform assumptions (Mullens et al., 2026). | Low to moderate: large observational cohorts, residual confounding remains possible. |
| Adoption drivers | A 2026 JAMA Network Open cohort of 20,313 surgeons found receipt of direct industry payment was associated with increased proportional use of robotic-assisted surgery, with a dose-response pattern (San Loh et al., 2026). | Low to moderate: observational policy evidence, useful for governance rather than causal proof of patient benefit. |
| Credentialing and privileging | A 2026 JAMA Surgery Viewpoint proposed competency-based, vendor-neutral privileging for robotic surgery rather than platform-specific exposure or case volume alone (Ashar & Selber, 2026). | Expert framework: useful for governance, not patient-outcome evidence. |
| Intraoperative AI implementation | A multisite qualitative study of AI-based Operating Room Black Box implementation found gaps between expectations and delivery: additional AI training needs, difficult data access, limited postoperative complication prediction, and limited academic deliverables (Thornton et al., 2026). | Low: qualitative implementation evidence, high value for deployment planning. |
| Surgical video AI | NEJM AI published a computer-vision study that used laparoscopic cholecystectomy video to identify surgical actions associated with blood loss and surgical experience (Aklilu et al., 2024). A 2026 NEJM AI appendectomy cohort extended this line of work by using swarm learning to train patient-level surgical-video disease-staging models across institutions without centralizing video data (Saldanha et al., 2026). | Low to moderate: promising analytic methods for local and privacy-preserving multicenter validation, not clinical intervention trials. |
The practical lesson is not that robotic surgery is good or bad. The lesson is that the unit of evidence is the procedure-institution-surgeon combination. A claim from rectal cancer surgery at expert centers should not be transferred to acute-care cholecystectomy, ventral hernia repair, hysterectomy, or community urology without local evidence.
Creating evaluations for robotic surgery
Every robotic surgery evaluation should begin with a falsifiable claim:
- Clinical outcome claim: robotic surgery reduces complications, conversion, recurrence, readmissions, pain, length of stay, or reoperation.
- Operational claim: robotic surgery improves OR throughput, case scheduling, turnover, staffing efficiency, or surgeon ergonomics.
- Training claim: robotic analytics improve skill acquisition, feedback quality, or credentialing reliability.
- AI claim: a model detects phases, instruments, anatomy, errors, risk states, or quality signals accurately enough to change decisions.
- Autonomy claim: a system performs a bounded physical subtask safely under defined oversight and abort conditions.
Do not evaluate “robotic surgery” as a generic intervention. Evaluate one claim at a time.
| Layer | Core question | Minimum evidence |
|---|---|---|
| Platform safety | Does the robot perform as intended under the cleared indication? | FDA record, device training requirements, malfunction reporting plan, local incident tracking. |
| Surgeon performance | Are operators past the learning curve for the procedure? | Case logs, simulation results, proctored cases, conversion and complication monitoring by surgeon. |
| Procedure outcomes | Does the robotic approach improve outcomes over the local comparator? | Procedure-specific outcomes, risk adjustment, comparable surgeon experience, 30-day and long-term outcomes when relevant. |
| AI perception | Does the model correctly detect instruments, anatomy, phase, smoke, bleeding, or errors? | Sensitivity, specificity, false alarms per hour, time-to-detection, external validation across surgeons and video systems. |
| Human factors | Do users understand when to trust, ignore, or override the system? | Simulation, silent-mode pilots, override audits, alert fatigue monitoring, qualitative workflow assessment. |
| Physical autonomy | Can the system act safely when tissue, lighting, bleeding, and anatomy vary? | Bench, simulation, ex vivo, animal, and eventually prospective clinical testing with predefined abort conditions. |
Match metrics to risk
Low-risk analytics, such as video indexing or case length prediction:
- Annotation accuracy
- Time saved in review or scheduling
- Inter-rater agreement with expert reviewers
- External validation across services
- User adoption and correction burden
Moderate-risk decision support, such as phase recognition or complication prediction:
- Sensitivity, specificity, PPV, NPV at local prevalence
- False alerts per case and per hour
- Time-to-detection before human recognition
- Calibration by procedure and patient subgroup
- Silent-mode performance before clinical use
High-risk intraoperative guidance, such as anatomy labeling or “do not cut” warnings:
- False negative rate for critical structures
- False positive rate causing unnecessary dissection delay
- Performance under blood, smoke, glare, lens fog, obesity, inflammation, adhesions, and variant anatomy
- Confidence display and uncertainty calibration
- Surgeon override and verification behavior
Physical action, including supervised autonomy:
- Task success rate and safety-margin violations
- Tissue trauma, force, thermal spread, bleeding, and clip or suture placement accuracy
- Recovery from near-miss states
- Human takeover latency
- Abort reliability
- Failure mode severity under worst-case scenarios
For irreversible actions, the evaluation threshold must be higher than diagnostic AI. A missed pulmonary nodule can still be reviewed. A divided bile duct cannot be undivided.
Use staged evidence, not one benchmark
The IDEAL framework for surgical robotics emphasizes development, comparative evaluation, and long-term monitoring rather than treating a single trial as the endpoint (Marcus et al., 2024). A practical hospital sequence is:
- Technical verification: confirm FDA indication, service contracts, instrument compatibility, downtime plan, and MAUDE reporting workflow.
- Simulation and dry-lab testing: test surgeon setup, docking, instrument exchange, emergency undocking, and AI display failure.
- Proctored clinical introduction: restrict to selected surgeons and procedures with clear exclusion criteria.
- Silent AI trial: run AI video analytics without clinical display, compare against expert annotation and outcomes.
- Limited visible pilot: display AI outputs to trained users, require manual verification, and audit overrides.
- Service-line deployment: monitor outcomes, costs, case mix, surgeon learning curves, and patient-reported outcomes.
- Post-deployment surveillance: review complications, conversions, reoperations, device malfunctions, video-model drift, and alert burden.
Skipping stages is most dangerous when AI moves from retrospective analytics to real-time intraoperative guidance.
Evals by use case
Surgical video analytics: Evaluate video analytics as a measurement system before treating them as a quality or safety intervention. Minimum evidence includes public and local test sets separated by surgeon, site, patient factors, and video system; inter-rater agreement for ground-truth labels; performance under blood, smoke, glare, lens cleaning, and off-axis camera views; error taxonomy; and prospective silent-mode validation before clinical display.
Robotic skill assessment: Automated performance metrics can reduce subjectivity in surgical education, but they must not collapse skill into speed or motion economy alone. A systematic review in the British Journal of Surgery found heterogeneous tools for robotic technical skills assessment and emphasized validity and reliability as central requirements (Boal et al., 2024). Minimum evidence includes correlation with blinded expert ratings, predictive validity for patient outcomes or supervised entrustment decisions, fairness across training level and prior robotic exposure, separation of technical execution from case complexity, and feedback that identifies remediable behaviors.
Anatomy labeling and warning systems: Anatomy labeling is high risk because confident false labels can create false reassurance. The safer near-term design is an uncertainty-aware warning system rather than an authoritative “this is the duct” display. Minimum evidence includes critical-structure false negative rates under worst-case visual conditions, confidence calibration, “unknown” states, stress tests with inflammation and variant anatomy, and explicit prohibition on irreversible action based on AI label alone.
Autonomous and semi-autonomous subtasks: Research is moving quickly. SRT-H used language-conditioned imitation learning for autonomous ex vivo cholecystectomy steps and achieved 100% success across 8 unseen pig gallbladders (Kim et al., 2025). A separate Science Robotics study introduced a surgical embodied intelligence simulator and demonstrated task autonomy across simulated, ex vivo, and in vivo animal settings (Long et al., 2025). These studies show meaningful progress in perception, planning, recovery, and sim-to-real transfer. They do not establish clinical readiness. Human surgery adds live bleeding, patient motion, anesthetic constraints, instrument failures, legal accountability, and rare anatomy that small experimental samples cannot resolve.
Procurement and governance
Robotic surgery programs should be governed like service-line investments. The evaluation should compare robotic surgery against the institution’s current best alternative, not against a theoretical average laparoscopic program.
Core local metrics:
- Case volume by procedure and surgeon
- Conversion to open surgery
- Intraoperative injury and bleeding
- Operative time, docking time, turnover time, and late-day delays
- 30-day complications, readmissions, emergency department returns, and reoperations
- Cancer-specific outcomes where relevant
- Patient-reported pain, function, urinary, sexual, and quality-of-life outcomes when relevant
- Direct and total costs, including instruments, disposable supplies, service contracts, staffing, depreciation, and OR time
- Training throughput and effects on laparoscopic competency
A 2025 systematic review of cost analyses in randomized trials found that robotic-surgery cost analyses are often incomplete, which means hospitals should not accept generic cost-effectiveness claims without local accounting (Bosscha et al., 2025).
Robotic surgery adoption can be shaped by marketing, patient demand, hospital competition, and industry relationships. The 2026 JAMA Network Open study linking industry payments to increased robotic-assisted surgery use does not prove inappropriate care, but it does justify governance around disclosure, credentialing, and value review (San Loh et al., 2026).
Robotic surgery committees should include surgery, anesthesia, nursing, sterile processing, biomedical engineering, finance, compliance, patient safety, and informatics. AI-enabled modules add model governance, data governance, and cybersecurity requirements.
Red flags:
- FDA status is unclear, misrepresented, or outside the intended use
- The vendor describes teleoperation as autonomy
- Outcomes are reported without comparator, case mix, surgeon experience, or learning-curve context
- AI analytics are trained on one institution’s videos and deployed elsewhere without external validation
- Anatomy labels are displayed without uncertainty or “unknown” states
- The system cannot export error logs, model version, video timestamp, and user override records
- Case volume is too low to maintain proficiency or amortize cost
- Training emphasizes platform operation but not failure recognition and emergency undocking
- Patient-facing marketing implies superior outcomes without procedure-specific evidence
No robotic-surgery AI should be trusted for irreversible intraoperative action without independent surgeon verification.
Postoperative AI Applications
Complication Prediction
Surgical Site Infection (SSI) Prediction:
ML models predict SSI risk using:
- Patient factors (diabetes, obesity, smoking, immunosuppression)
- Operative characteristics (duration, complexity, contamination class)
- Intraoperative variables (glucose control, normothermia, antibiotic timing)
- Postoperative factors (drain output, pain scores)
Evidence boundary: Prediction performance varies by outcome, institution, time horizon, reference standard, and comparator. No source cited here supports a class-wide AUC advantage over clinical judgment.
Limitations:
- Low prevalence can produce low positive predictive value even when discrimination appears strong
- A risk score alone does not establish that changing prophylaxis prevents infection
- A proposed surveillance use requires a tested response pathway and alert-burden analysis
Postoperative Delirium:
Prediction models incorporating preoperative cognitive assessment, anesthesia factors, and postoperative medications identify high-risk patients for:
- Non-pharmacologic prevention (reorientation, sleep hygiene, family presence)
- Avoidance of deliriogenic medications
- Enhanced monitoring
Evidence boundary: A model may stratify risk without outperforming clinician judgment or improving delirium outcomes. A linked prevention bundle requires separate evidence.
Anastomotic Leak Prediction:
ML models analyzing postoperative labs (CRP trajectory), vital signs, and clinical notes can identify leak risk earlier than clinical suspicion alone.
Challenge: Low-prevalence outcomes can produce imprecise estimates and low positive predictive value. Prevalence and performance should be reported for the actual surgical population.
Deterioration Monitoring
AI systems can combine continuous vital signs, laboratory trends, nursing documentation, and medication data into deterioration-risk estimates. Prediction horizon and comparative performance are model- and setting-specific; no universal 6–12 hour advantage is established.
Applications:
- Postoperative hemorrhage
- Respiratory failure
- Sepsis
- Cardiac events
Randomized deployment evidence: A pragmatic cluster-randomized trial placed a passive display of continuous risk trajectories on an 85-bed cardiology and cardiac-surgery ward. Across 10,422 inpatient visits, the display did not improve the primary outcome of hours free of clinical deterioration. Clinicians transferred 801 visits between display-on and display-off beds, preferentially moving sicker patients toward display beds and weakening the randomized comparison (Keim-Malpass et al., 2026). A visible risk score without a defined response pathway did not improve the trial’s primary patient outcome.
Wong and colleagues’ Epic Sepsis Model study was an external validation, not evidence that a postoperative alert improves outcomes (Wong et al., 2021). Safe implementation should therefore test who receives the signal, what assessment follows, how competing alerts are prioritized, and whether the complete response pathway improves prespecified outcomes.
Surgical Quality and Education
Video-Based Surgical Assessment
AI analysis of surgical videos enables objective skill assessment and quality improvement.
Applications:
Skill Scoring:
- Objective assessment of technical performance
- Identifies specific errors (tissue trauma, bleeding, inefficiency)
- Provides quantitative feedback for training
Evidence: Lavanchy and colleagues developed a video-based workflow-recognition and skill-assessment approach in a selected laparoscopic dataset. Its relationship to expert ratings and outcomes does not establish a class-wide training benefit (Lavanchy et al., 2021). A 2026 JAMA Surgery quality-improvement study evaluated multi-instrument tracking and spatiotemporal kinematic features for granular assessment of laparoscopic cholecystectomy skill, adding a current example of measurement research rather than evidence that automated feedback improves patient outcomes (Zeng et al., 2026).
Benefits for surgical education:
- Objective feedback supplements subjective faculty evaluation
- Tracks skill progression over time
- Identifies specific areas needing improvement
- Benchmarks against peer performance
Quality Improvement:
- Retrospective review of complications to identify technical factors
- Process improvement for OR efficiency
- Standardization of surgical techniques
Challenges:
- Privacy and medicolegal concerns about routine recording
- Surgeon resistance to surveillance
- Does not capture decision-making quality (only technical execution)
- Storage and analysis infrastructure requirements
Natural Language Processing for Operative Notes
AI extraction of structured data from operative notes enables:
Quality Metrics:
- Automated calculation of process measures (antibiotic timing, VTE prophylaxis)
- Complication detection from dictated notes
- Adherence to surgical best practices
Registry Auto-Population:
- Reduces manual data entry burden for NSQIP, VASQIP, other registries
- Improves data completeness and accuracy
Clinical Decision Support:
- Extraction of critical operative details for downstream care (mesh type in hernia repair, prosthesis in joint replacement)
Evidence boundary: Extraction performance varies by data element, institution, note template, reference standard, and the treatment of missing or ambiguous language. No source cited here supports a class-wide accuracy above 95%. Nuanced findings, laterality, implants, complications, and judgment-based assessments require element-level validation and review.
Specialty-Specific Applications
Different surgical specialties face unique challenges and opportunities for AI integration:
General Surgery
- Hernia recurrence risk prediction
- Cholecystectomy difficulty scoring
- Bile duct injury prevention (research phase)
Orthopedic Surgery
- Fracture detection AI (high accuracy for simple fractures)
- Joint replacement planning and component sizing
- Spinal navigation systems (FDA-cleared)
- Ligament injury diagnosis from MRI
Neurosurgery
- Brain tumor segmentation for resection planning
- Epilepsy focus localization
- Surgical navigation systems
- Intraoperative tumor margin assessment (research)
Cardiac Surgery
- Surgical risk models (STS score enhanced with ML)
- Intraoperative echocardiography interpretation
- ICU outcome prediction
Perioperative hypotension prediction: The HYPE-2 randomized trial tested a machine-learning-derived Hypotension Prediction Index with diagnostic guidance during elective on-pump cardiac surgery and ICU care. Among 130 patients included in the primary analysis, the intervention reduced the median time-weighted average of MAP below 65 mm Hg by 63% and reduced time spent in hypotension by a median 28 minutes versus standard care (Schuurmans et al., 2025). The study supports protocolized hemodynamic decision support in cardiac anesthesia, but it was single-center and measured hypotension burden, not downstream complications or mortality.
Thoracic Surgery
- Lung nodule characterization from CT
- Surgical approach selection (VATS vs. thoracotomy)
- Lymph node metastasis prediction
Vascular Surgery
- AAA rupture risk prediction
- Vascular anatomy segmentation
- Endovascular procedure planning
Plastic Surgery
- Breast reconstruction outcome prediction
- Aesthetic outcome simulation
- Flap viability monitoring (research)
Breast Surgery
Claire OCT System (Perimeter Medical Imaging AI), PMA P250008:
FDA approved the Claire OCT System on March 3, 2026, as an adjunctive three-dimensional imaging tool for excised lumpectomy margins in patients with biopsy-confirmed breast cancer. The system combines wide-field optical coherence tomography with an AI computer-aided detection algorithm that marks focal areas suspicious for cancer and is used concurrently with physician interpretation. It is intended for use with other standard margin-evaluation methods, not as a replacement (FDA PMA P250008; FDA Summary of Safety and Effectiveness).
- The pivotal study enrolled 206 participants. After standard care, 35 had an unaddressed positive margin; after device-aided assessment, 28 did, an absolute reduction of 3.4 percentage points and relative reduction of 20% that met the prespecified performance goal (P = .0050).
- Clinicians using the system identified actionable residual disease in 14 of the 35 participants who had residual disease after standard care.
- The 88.1% margin-level clinical-decision accuracy reported in the labeling was a post hoc measure of the clinician-device workflow, not standalone algorithm accuracy (FDA user manual P250008).
- FDA approved a predetermined change control plan specifying permitted modifications and validation controls. It is not unrestricted permission to change the model without regulatory constraints.
Clinical significance: Claire does not replace standard histopathology or make an autonomous excision decision. The FDA evidence supports an adjunctive intraoperative workflow and a margin endpoint. Longer-term re-excision, recurrence, patient-reported, and system-level outcomes remain separate claims.
Critical Limitations and Risks
Immediacy of Harm: Unlike diagnostic errors that can be caught through physician review, intraoperative AI errors cause immediate, potentially irreversible patient harm.
Complexity of Surgical Judgment: Surgery requires integration of visual, tactile, and proprioceptive information with anatomical knowledge, pattern recognition from thousands of prior cases, and real-time adaptation to unexpected findings. AI does not replicate this.
Medicolegal Implications: Responsibility depends on the facts, applicable law, institutional policy, device labeling, product design and warnings, professional conduct, and causation. Following or overriding an AI output does not create a categorical liability result. A foreseeable risk is defensive overreliance when the governance program treats every alert as a legal command rather than a clinical signal requiring interpretation.
Technology Failure Modes: Computer vision fails with blood, smoke, optical artifacts. ML models fail with out-of-distribution inputs (unusual anatomy, rare findings). Risk models fail when patient circumstances differ from training data.
Trust Calibration: Surgeons must neither over-trust (following AI suggestions without verification) nor under-trust (ignoring useful AI alerts). Achieving appropriate calibration is difficult (Char et al., 2018).
Regulatory and Medicolegal Considerations
FDA Regulation of Surgical AI
FDA pathway and device class depend on the exact product’s intended use, risk, technological characteristics, and applicable statutory pathway. Labels such as planning software, navigation system, risk calculator, robot, or autonomous feature do not by themselves determine class or oversight. Verify the product’s FDA record, indication, inputs, users, and required supervision before clinical use.
Medicolegal Principles
No universal rule assigns every AI-related loss to the surgeon. Clinicians, institutions, and manufacturers may have different duties depending on the facts and jurisdiction. Key documentation practices include:
- Determine with legal, ethics, clinical, and patient input when AI use is material to a patient’s decision and what disclosure is required
- Preserve product, version, output, review, disagreement, and action when they are clinically or operationally material
- Do not treat a documentation template as proof that independent verification occurred
The Liability Dilemma
- Following an incorrect output: may support an allegation of overreliance, but liability still requires duty, breach, causation, and damages under applicable law
- Overriding a correct output: may be examined against the information available at the time, the product’s role, the clinical reasoning, and the applicable standard
- Governance response: define verification, escalation, override, logging, and incident-review processes before deployment; obtain jurisdiction-specific advice rather than predicting outcomes categorically
Evidence-Based Guidelines for Surgical AI Adoption
Before Adopting Any Surgical AI:
- Demand evidence: Prospective validation studies in diverse populations, not just retrospective accuracy metrics (Nagendran et al., 2020)
- Understand training data: Was the model trained on cases like yours? (Procedure types, patient populations, institutional practices) (Beam & Kohane, 2018)
- Know the failure modes: How does the system fail? What are the error rates? What happens with unusual cases? (Vabalas et al., 2019)
- Assess workflow integration: Does this fit your existing workflow or require disruptive changes?
- Clarify liability: What does your malpractice carrier say about using this AI? What does hospital legal counsel advise?
- Verify regulatory status: Is this FDA-cleared? For what specific indication?
- Evaluate cost-effectiveness: Does the benefit justify the cost (both financial and cognitive/workflow burden)?
Safe Implementation Practices:
- Pilot testing: Start with low-stakes applications, expand carefully based on performance
- Parallel validation: Run AI alongside current practice, compare results before replacing current approach
- Defined oversight: Clear protocols for who reviews AI outputs and how discrepancies are resolved
- Incident reporting: Systems to capture AI errors or near-misses
- Ongoing validation: Monitor real-world performance, do not assume initial validation persists indefinitely
- User training: Ensure all users understand AI capabilities, limitations, and appropriate use
- Informed consent: Discuss AI use with patients when material to their decision-making
Red Flags (Avoid These AI Systems):
- Claims of autonomous surgical decision-making
- Black-box models with no explanation of predictions
- Lack of prospective validation studies
- Vendors unwilling to disclose training data characteristics
- No mechanism for reporting errors or failures
- Regulatory status unclear or misrepresented
- Pressure to adopt without adequate evaluation period
Professional Society Guidelines on AI in Surgery
The American College of Surgeons maintains an official Artificial Intelligence in Surgery resource page and an online course, Artificial Intelligence and Machine Learning: Transforming Surgical Practice and Education. The course covers fundamentals, selected surgical applications, video analysis, explainability, and ethics. ACS also announced expansion of its AI and machine-learning education in its 2025 education strategy.
These are educational and professional resources. They are not a product-specific clinical practice guideline, FDA authorization, or evidence that a particular surgical AI system improves outcomes. The distinction matters when a chapter turns professional activity into an attributed society recommendation.
AI Applications Recognized by ACS
ACS educational materials discuss several categories relevant to surgical practice, including:
- Ambient AI: Automated documentation of surgical encounters and procedures
- Prediction tools: Perioperative risk assessment and outcome prediction
- Research and writing solutions: Literature review, manuscript preparation assistance
NSQIP and Risk Prediction
The ACS National Surgical Quality Improvement Program Surgical Risk Calculator is an established statistical decision aid. It should be described as AI-adjacent rather than automatically classified as AI:
- It estimates procedure-specific risks from registry-derived models
- It can support shared decision-making when its assumptions and uncertainty are explained
- Its current model, coverage, and documentation should be verified from official ACS materials before stating dataset size or update frequency
- Its existence does not validate a separate machine-learning product
SAGES Guidelines
SAGES educational and research activity has addressed computer vision, video analysis, and emerging digital-surgery methods. That activity should not be restated as a formal product guideline unless a direct current guideline is linked. Areas of research interest include:
- Computer vision for surgical field analysis
- Real-time anatomical structure identification during laparoscopic procedures
- Surgical video analysis for quality improvement and training
Implementation Note: No SAGES product-selection clinical practice guideline was identified for this review. The chapter’s deployment framework is therefore an evidence-based handbook recommendation, not a quoted SAGES rule. Educational programming, society discussion, and clinical practice guidelines are different evidence categories.
Future Directions
Demonstrated Capabilities
- Retrospective and multicenter development of perioperative risk models
- Procedure- and platform-specific robotic-surgery comparative evidence
- Video-based phase, workflow, and skill measurement in selected datasets
- Randomized testing of model-guided ileostomy decisions and hypotension decision support
- FDA-approved adjunctive AI-assisted OCT margin imaging for a bounded breast-surgery workflow
- Autonomous execution of selected subtasks in simulation, ex vivo tissue, and animal research
Under Evaluation
- Real-time anatomical recognition with explicit uncertainty and independent surgeon verification
- Context-aware intraoperative warnings that improve decisions without increasing distraction or false reassurance
- Video-based feedback that improves training, entrustment decisions, or patient outcomes rather than merely correlating with expert scores
- Postoperative prediction connected to a response protocol that improves prespecified clinical outcomes
- Supervised robotic assistance for narrowly defined subtasks with reliable abort and recovery behavior
Beyond Current Clinical Evidence
- General autonomous soft-tissue surgery in patients without continuous surgeon control
- AI replacement of surgical judgment across unexpected anatomy, bleeding, equipment failure, and competing goals
- Real-time tissue characterization that is interchangeable with pathology across procedures
- Prediction of rare complications with near-perfect accuracy and negligible alert burden
- Elimination of surgical complications through AI
A capability demonstrated in simulation, ex vivo tissue, or an animal model remains preclinical until human clinical evidence supports the intended use.
Conclusion
Surgery is fundamentally a human activity requiring manual skill, real-time judgment, and adaptation to unique patient circumstances. AI can enhance the cognitive work surrounding surgery (risk assessment, planning, quality improvement) and may eventually provide useful intraoperative information. But the surgeon’s hands, eyes, judgment, and responsibility remain central.
The most successful surgical AI applications will be those that respect the complexity of surgery, acknowledge uncertainty transparently, augment rather than replace expertise, and prioritize patient safety over technological impressiveness.
Surgeons should embrace AI as a powerful adjunct while maintaining the healthy skepticism, independent verification, and personal accountability that define good surgical practice.
Check Your Understanding
The following cases are hypothetical teaching exercises. Patients, products, model outputs, complications, costs, legal arguments, fault allocation, verdicts, and institutional responses are illustrative unless a source is linked in the same sentence. They are not reports of actual events or predictions of legal outcomes.
Scenario 1: Hypothetical Risk Calculator Overestimates Surgical Risk
You’re a colorectal surgeon evaluating an 82-year-old woman with Stage III colon cancer. She’s otherwise healthy: active, independent ADLs, no major comorbidities, ECOG 0.
A fictional AI surgical risk calculator estimates:
- 30-day mortality risk: 18%
- Major complication risk: 45%
- Recommendation: “High risk - consider non-operative management”
Traditional ACS NSQIP calculator estimates:
- 30-day mortality: 3.2%
- Major complication: 12%
Patient’s oncologist refers to you stating “AI says surgery too risky. Recommend palliative chemo only.”
Your clinical assessment: Patient is good surgical candidate. Age alone should not preclude curative surgery. Frailty assessment normal. Cardiopulmonary exam reassuring.
Decision point: Do you:
- Follow AI recommendation, refer to medical oncology for palliative chemotherapy
- Override AI, recommend surgery based on your clinical judgment
Answer 1: What explains the discrepancy between AI and traditional calculators?
The discrepancy cannot be attributed to age without examining the model. Plausible questions include whether the fictional model represented:
- Functional status: Patient is ECOG 0, independent, not frail
- Comorbidity burden: Minimal comorbidities despite age 82
- Fitness indicators: Normal cardiopulmonary reserve
Potential model and data problems to test:
- Whether age acted as a proxy for unmeasured frailty or comorbidity
- Whether functional status was missing, miscoded, or outside the training distribution
- Whether the requested procedure and indication were mapped correctly
- Whether the interface showed a calibrated absolute risk for this population
Comparison calculator:
- Uses validated risk factors (ASA class, functional status, comorbidities)
- May produce a different estimate because it uses a different cohort, variable set, endpoint definition, and calibration method. The lower number is not automatically the correct number.
Answer 2: What are the liability implications of each choice?
Choice A (Follow AI, decline surgery):
Plaintiff argument (if patient dies from untreated cancer):
- Surgeon inappropriately deferred to AI algorithm
- Failed to exercise independent clinical judgment
- Denied patient potentially curative treatment based on flawed AI estimate
- The clinical decision was not supported by an individualized assessment or a documented shared-decision process
No cited legal authority supports a categorical precedent that following a conflicting decision-support output establishes liability.
Choice B (Override AI, proceed with surgery):
If patient has major complication or dies:
Plaintiff argument:
- Surgeon ignored AI warning of 18% mortality risk
- Proceeded with high-risk surgery against AI recommendation
- The discrepancy was not investigated or communicated before proceeding
Defense argument:
- AI is decision support tool, not substitute for clinical judgment
- The full clinical assessment incorporated function, frailty, comorbidity, goals, operative alternatives, and anesthesia evaluation
- The model’s inputs, endpoint, and applicability were examined rather than accepted or rejected reflexively
- Traditional NSQIP calculator (validated, widely used) supported decision
- Patient underwent informed consent understanding risks
Outcome boundary: The vignette cannot predict a verdict. Liability would depend on jurisdiction, the patient’s course, expert evidence, the product and workflow, the decision process, causation, and the complete record. Documentation is important evidence, but it does not transform a poor decision into a reasonable one.
Answer 3: How should you handle this AI-clinical judgment conflict?
Appropriate approach:
- Investigate AI discrepancy
- Review AI inputs: What features drove high-risk estimate?
- Compare with an appropriate validated risk tool, such as the relevant ACS NSQIP estimate when applicable
- Consult surgical colleagues: Would they operate on this patient?
- Comprehensive clinical assessment
- Gait speed, grip strength (frailty markers)
- Cardiopulmonary exercise testing if available
- Geriatric assessment
- Functional status (independent vs. dependent ADLs)
- Multidisciplinary discussion
- Present case at tumor board
- Geriatric surgery consult if available
- Anesthesia risk assessment
- Transparent informed consent
- Discuss both AI and traditional risk estimates with patient
- Explain why estimates differ (age vs. physiologic status)
- Present alternatives (surgery, chemotherapy alone, observation)
- Document: “AI calculator estimated 18% mortality; however, clinical assessment suggests patient is physiologically fit. Traditional NSQIP calculator estimates 3.2% mortality. Discussed both estimates with patient. Patient understands risks, chooses surgery.”
- Document clinical reasoning
- “AI risk calculator estimates high risk primarily based on age 82. However, patient demonstrates excellent functional status (ECOG 0, independent ADLs, normal gait speed), minimal comorbidities, normal cardiopulmonary reserve. Traditional NSQIP calculator estimates mortality 3.2%. Clinical judgment: patient is appropriate surgical candidate. AI estimate likely over-weighted chronologic age without adequate consideration of physiologic fitness.”
Lesson: Discordance is a reason to investigate inputs, endpoints, calibration, and patient context, not a reason to select whichever number supports the preferred plan. The final recommendation should integrate the clinical assessment, appropriate comparison tools, alternatives, patient goals, and uncertainty.
Scenario 2: Hypothetical Intraoperative AI Misidentifies Critical Anatomy
A surgeon is performing robotic-assisted partial nephrectomy for a small renal mass using a fictional teleoperated platform with an experimental anatomy-labeling overlay. This is not a description of a clinically available da Vinci Xi feature or an FDA-authorized autonomous anatomy-labeling function.
AI system features:
- Real-time anatomical labeling (kidney, renal artery, renal vein, ureter, tumor)
- Proximity alerts when instruments near critical structures
- Augmented reality overlay on surgical view
Intraoperative event:
During hilar dissection, AI labels renal artery and renal vein on display. You prepare to clamp renal artery for tumor excision.
Your visual assessment: Structure labeled “renal artery” appears larger than expected, bluish tint, pulsations not prominent.
Uncertainty: Is this truly renal artery or is AI mislabeling renal vein as artery?
Decision point: Do you:
- Trust AI label, clamp structure labeled “renal artery”
- Pause, verify anatomy manually before clamping
You choose: Option B (pause and verify)
Manual verification: In this fictional event, Doppler ultrasound indicates that the structure labeled “renal artery” is actually the renal vein. The true renal artery lies posteriorly and was not labeled.
If you had clamped based on AI label: Would have clamped renal vein, not artery → inadequate ischemic control → bleeding during tumor excision, potential need for total nephrectomy.
Answer 1: Why did the AI mislabel critical anatomy?
AI computer vision failure modes:
Anatomical variation: This patient had variant renal vascular anatomy (early branching, aberrant vessel course)
- AI trained on typical anatomy
- The development data may not adequately represent the relevant vascular variation; no prevalence or training-set composition is assumed in this fictional case
Tissue appearance similarity: Renal artery and vein can appear similar on video (both red/pink, both tubular)
- AI relies on position, caliber, pulsatility
- In variant anatomy, typical positional relationships disrupted
Partial occlusion: Surgical manipulation may have partially occluded artery → reduced pulsations → AI misidentified as vein
Display and abstention design: The interface may have displayed a label without calibrated uncertainty, an “unknown” state, or a clear boundary on how the output could be used. The fictional 60% to 70% confidence range previously assigned here was not measured.
Answer 2: What are the liability implications if you had clamped the wrong vessel?
If you clamped renal vein instead of artery:
Immediate consequences:
- Inadequate tumor ischemia → bleeding during excision
- Potential renal vein thrombosis
- May require total nephrectomy instead of partial
- Patient loses kidney function unnecessarily
Malpractice analysis:
Plaintiff argument:
- Surgeon blindly followed AI labeling without manual verification
- Failed to exercise fundamental surgical principle: verify anatomy before clamping/cutting
- Fell below standard of care by deferring anatomical judgment to AI
- Patient lost kidney due to surgeon’s inappropriate reliance on technology
Defense argument:
- AI was marketed as “surgical intelligence” system
- The institution represented the system as suitable for the workflow and trained the surgeon to use the overlay
- Anatomical variation not surgeon’s fault
- Product design, labeling, training, configuration, and warnings should be examined along with the surgeon’s conduct
Outcome boundary: No verdict follows automatically. A case would examine the product’s actual FDA status and intended use, the training and warnings, the surgeon’s contemporaneous observations, ordinary verification steps for the procedure, institutional governance, causation, and damages. Smith v. Hospital was a fabricated precedent in the restored text and is not a legal authority. A fictional case must not be presented as precedent, even when its lesson sounds plausible.
Answer 3: What are the appropriate use principles for intraoperative AI?
Surgical AI as a fallible information source:
- AI suggestions are hypotheses, not facts
- AI labels = “This might be renal artery”
- Surgeon verifies = “I confirm this is renal artery”
- Verify before irreversible action
- Before clamping, cutting, coagulating: manual confirmation
- Use additional tools: Doppler, manual palpation, ICG angiography, direct visualization
- Heightened skepticism in variant anatomy
- If anatomical landmarks do not match expected positions
- If AI labels conflict with visual assessment
- If patient has known anatomical variants (duplicated vessels, horseshoe kidney)
- Evaluate uncertainty and abstention
- Determine whether displayed confidence is calibrated for the procedure and visual conditions
- Require an “unknown” or abstain state where the model lacks adequate evidence
- Do not convert a high confidence display into permission for an irreversible action
- Continuous cross-checking
- Compare AI labels with your visual assessment at each step
- If discrepancy, investigate before proceeding
Institutional safeguards:
- Training requirements
- A local credentialing and training program can cover:
- AI failure modes
- When to trust vs. verify AI
- Manual verification techniques
- A local credentialing and training program can cover:
- Quality assurance
- Review cases where AI labeling was incorrect
- Share at M&M conferences
- Track AI error rates by anatomy type, procedure
- Documentation
- When AI labeling conflicts with surgeon assessment, document:
- “AI labeled [structure] as [label]; however, manual verification with [Doppler/ICG/palpation] confirmed [correct identity]”
- When AI labeling conflicts with surgeon assessment, document:
Lesson: Intraoperative AI is assistive, not authoritative. Critical anatomy should be confirmed through the independent checks appropriate to the procedure before an irreversible action. “Verify independently, then use AI as additional information” is safer than treating the display as the starting truth.
Scenario 3: Hypothetical Postoperative AI Alert Fatigue
A surgical quality director is implementing a fictional AI-based early-warning system for postoperative deterioration. The metrics and patient event are illustrative and are not performance claims about the Rothman Index or another commercial product.
AI system: Analyzes vital signs, lab values, nursing assessments every 15 minutes. Generates alert when patient predicted to be at increased risk for:
- Sepsis
- Respiratory failure
- Acute kidney injury
- Need for ICU transfer
Month 1 performance:
- Alerts generated: 847 alerts across 320 postoperative patients (2.6 alerts per patient)
- True-positive alerts under the fictional adjudication rule: 23
- Alerts not linked to an adjudicated complication: 824
- False-discovery proportion among alerts: 97.3%
- Alert-level positive predictive value: 2.7%
These alert-level quantities are not the same as patient-level sensitivity, specificity, or false-positive rate. Repeated alerts per patient require an explicitly defined unit of analysis and adjudication window.
Clinical impact:
- Nursing staff overwhelmed by alerts
- Most alerts dismissed as “AI crying wolf”
- Alert fatigue setting in (nurses ignoring alerts)
Week 4 critical event:
- 62-year-old man, post-colectomy day 2
- AI generates alert at 2 AM: “High risk for sepsis - recommend immediate evaluation”
- Night nurse dismisses alert (patient appears stable, vital signs acceptable)
- No physician notification
- 6 AM: Patient found hypotensive (BP 82/45), tachycardic (HR 128), altered mental status
- Diagnosis: Anastomotic leak with peritonitis and sepsis
- Patient requires emergent return to OR, ICU care
- Prolonged hospital stay, family files complaint: “Why wasn’t the AI alert acted on?”
Answer 1: What caused the alert fatigue?
High nonactionable-alert burden may be driven by:
- Low disease prevalence
- In the fictional month, 25 of 320 patients experienced the defined complication outcome and 23 received at least one prior alert
- Low prevalence can depress positive predictive value even when patient-level sensitivity appears high
- Alert-level PPV cannot be reconstructed from patient-level sensitivity and specificity when patients can generate repeated alerts
- Threshold calibration
- A low threshold is one possible cause, but investigators should also examine duplicated alerts, prediction horizon, refractory periods, outcome definition, data delay, and calibration
- Lack of clinical context
- AI analyzes physiologic data only
- Does not know: patient just returned from 2-hour physical therapy session (explains elevated HR), patient received fluid bolus (explains improved BP trends), patient had expected postoperative fever
- Poor alarm design
- All alerts same priority level
- No distinction between “mild concern” vs. “urgent evaluation needed”
- No incorporation of clinical trajectories (improving vs. worsening trends)
Alert fatigue:
- Repeated nonactionable alerts can teach staff that the display carries little information
- In this fictional stream, 97.3% of alerts were not linked to the adjudicated outcome, not necessarily physiologically “false”
- Clinically important signals can become difficult to distinguish from repeated low-value alerts
Answer 2: Who is liable for the missed anastomotic leak?
The record raises questions about both institutional design and individual response, but liability cannot be allocated from the vignette.
Hospital institutional liability:
Plaintiff argument:
- Hospital deployed a system whose alert stream had a 97.3% false-discovery proportion under its own adjudication rule
- Created alert fatigue environment where critical alerts ignored
- Failed to calibrate system before clinical deployment
- Should have monitored alert fatigue, intervened when nurses began dismissing alerts
Nursing liability:
Plaintiff argument:
- Nurse dismissed AI alert without evaluating patient
- Failed to notify physician of high-risk alert
- Did not document why alert was dismissed
- The response did not follow the locally defined escalation and documentation protocol
Defense argument (nursing):
- The alert stream had low positive predictive value and high repeated burden
- Nurse made reasonable judgment based on clinical assessment (patient appeared stable)
- Hospital created untenable alert burden
- Individual nurse cannot be expected to thoroughly evaluate 2.6 alerts per patient per shift
Outcome boundary: A review would examine the protocol, training, staffing, alert design, other concurrent alarms, the nurse’s assessment, system logs, causal timing, and the jurisdiction’s law. No percentage or priority of fault can be assigned responsibly from these hypothetical facts.
Answer 3: How should AI early warning systems be implemented safely?
System calibration:
- Prespecified operating point
- Choose sensitivity, positive predictive value, alert rate, lead time, and repeat-alert behavior from the clinical consequence and response capacity
- Report patient-level and alert-level performance separately
- Do not import a universal minimum PPV or acceptable sensitivity loss from another deployment
- Tiered alert system
- Low priority (informational): “Monitor patient closely”
- Medium priority (nursing assessment): “Evaluate patient within 1 hour”
- High priority (physician notification): “Urgent evaluation needed, notify MD immediately”
- Reserve interruptive alerts for conditions whose urgency, evidence, and response pathway justify interruption; no universal 30% PPV threshold is established
- Clinical context integration
- Test whether context-aware suppression during expected postoperative physiology reduces burden without delaying recognition
- Incorporate clinical context (patient just ambulated, received fluid bolus, normal post-op fever)
- Trend analysis (worsening vs. stable vs. improving)
Workflow integration:
- Alert response protocol
- Define a response time and escalation path from the clinical condition, staffing, and local protocol rather than using an unsupported universal 15-minute rule
- Document: “AI alert reviewed. Patient assessed. Findings: [stable vs. concerning]. Action: [continued monitoring vs. physician notified].”
- Feedback loop
- Track AI alert accuracy
- Review how many alerts were actionable, how many were repeated, which outcomes occurred without a signal, and whether the response changed care
- Adjust thresholds based on performance
- Human oversight
- Assign clear ownership for review and action
- Direct paging, routed review, and dashboard display each require human-factors testing; a human gatekeeper is not automatically safer
Quality monitoring:
- Track alert fatigue
- Monitor alert dismissal rates
- Monitor dismissal without assessment, response time, duplicate alerts, interruptions, and staff-reported burden
- Prespecify action thresholds from the local baseline and clinical risk rather than an unsupported universal 80% cutoff
- Audit missed complications
- For every complication, determine: Did AI alert? Was alert acted on?
- If multiple complications missed due to dismissed alerts → pause system, recalibrate
- Continuous improvement
- Vendor partnership: Provide feedback on false positives
- Request threshold adjustment or better risk stratification
Lesson: AI early-warning systems should be evaluated as a complete sociotechnical intervention. Discrimination is insufficient without calibration, alert-rate analysis, human-factors testing, a response pathway, and prospective outcome evaluation. High sensitivity does not compensate for an alert stream that the care team cannot interpret or act on reliably.