Appendix H — Clinical AI Vendor Evaluation Toolkit

TL;DR

Six review domains: Clinical validation, patient safety, fairness and access, privacy and security, workflow and change control, and business continuity.

Decision boundary: The scorecards are discussion aids, not validated procurement instruments. A high average cannot offset a material safety, privacy, regulatory, or evidence failure.

Critical questions: Does the cited evidence match the exact product version and claim? What data leave the institution? Who acts on the output? How are changes, incidents, monitoring, and exit handled?

Purpose

This appendix provides a structured discussion framework for evaluating commercial AI products in clinical practice. It is not a validated scoring instrument, procurement standard, or substitute for clinical, privacy, security, regulatory, legal, and contracting review.

Who should use this: - Hospital administrators evaluating AI vendors - Department chairs considering AI tools for their specialty - Chief Medical Information Officers (CMIOs) assessing AI products - Physician practice leaders making technology decisions - IRB/ethics committees reviewing AI deployments

What you’ll find: - Structured evaluation scorecards - Red flag identification guides - Sample questions to ask vendors - Decision frameworks for Go/No-Go - Contract discussion prompts for institutional counsel and procurement teams


Introduction: Why Vendor Evidence Needs Structured Review

The Clinical AI Morgue shows that public evidence does not support one aggregate cost or harm total across clinical AI failures. It does show recurring, detectable problems: claim-evidence mismatch, biased targets, version drift, weak transportability, workflow burden, and missing accountability.

Vendor due diligence should identify what is known, what remains uncertain, and what would block use.

Algorithm Transparency in Certified Health IT

For systems embedded in certified health IT, vendor due diligence should include applicable HTI-1 algorithm-transparency documentation. The rule established transparency requirements for predictive decision support interventions within its certified-health-IT scope (ONC, HTI-1 Final Rule). It is not a universal disclosure law for every healthcare AI product.

Procurement teams should ask whether the product is part of an ONC-certified health IT module, whether it qualifies as a decision support intervention or predictive model under the certification program, and what baseline information the vendor can provide to assess fairness, appropriateness, validity, effectiveness, and safety. HTI-1 transparency does not prove clinical benefit. It creates a minimum documentation floor that should be combined with local validation, subgroup performance review, security review, and post-deployment monitoring.

The Vendor-Physician Information Asymmetry

Vendor knows: - Where the model was trained - What its clinical limitations are - Where external validation failed - What patient populations it doesn’t work for - What the false alarm rate is in real-world use - Which hospitals abandoned the system after pilot

The purchasing institution may initially know: - What the vendor provides in sales materials - What can be verified in publications, regulatory records, reference checks, and technical documentation

The review process should close that information gap before a consequential decision.


Quick Reference: The Six-Domain Evaluation Framework

Use this framework to systematically evaluate any clinical AI vendor:

Domain Key Questions Red Flags
1. Clinical Validation External validation? Peer-reviewed publication? Improved outcomes? Internal validation only; no publications; only AUC reported
2. Patient Safety Safety testing? Adverse events tracked? Failure mode analysis? No safety data; no mention of harms; “100% accurate” claims
3. Fairness & Equity Performance across demographics? Bias audits? No fairness testing; “we don’t use race so it’s fair”
4. Privacy & Security HIPAA roles? BAA when required? Data flows? Security controls? Vague compliance claims; unknown subprocessors, retention, or training use
5. Workflow Integration End-user testing? EHR integration? Training provided? No physician research; “plug and play”; minimal training
6. Business Viability Continuity, support, portability, exit, and exact regulatory status? Unclear roadmap, dependency risk, no exit or continuity plan

Domain 1: Clinical Validation

The Questions to Ask

Critical Validation Questions
  1. Training Data
  2. Validation Studies
  3. Clinical Performance Metrics
  4. Clinical Outcomes (MOST IMPORTANT)

Red Flags

Proceed with extreme caution if:

  • “Validated on 100,000 patients” - But all from the same institution (not external validation)
  • “95% accuracy” - On cherry-picked test set; no external validation
  • “AUC 0.92” - But no data on whether clinical outcomes improved
  • “Deployed in 150+ hospitals” - Deployment ≠ Effectiveness; no outcome data
  • “Proprietary validation” - No peer-reviewed publications; “trust us”
  • Internal validation only - This does not answer whether performance transports to the proposed site.
  • Vendor-funded validation without transparent methods or independent replication - Funding is not automatic disqualification, but it affects the independence evidence needed.
  • Only technical metrics reported - AUC, accuracy without clinical outcomes
  • “FDA-cleared” without recall history or postmarket performance data - FDA authorization is not the end of evaluation. A 2025 cross-sectional study of 950 FDA-cleared AI-enabled devices found recalls uncommon (60 devices, 6.3%) but concentrated early, with 43.4% of recall events occurring within 12 months of clearance, and lack of clinical validation independently associated with recall (odds ratio 2.8; 95% CI, 1.6-4.7) (Lee et al., 2025)

Lessons from the Epic Sepsis Evidence

One external evaluation of the Epic Sepsis Model found hospitalization-level AUROC 0.63, sensitivity 33%, and positive predictive value 12% at the evaluated threshold. The reported 7% was a process-of-care subgroup, not a before-onset detection rate, and the study did not establish that alerts caused fatigue or patient harm (Wong et al., 2021). Later prospective version 2 evidence showed stronger discrimination with site variability and substantial alert burden (Wong et al., 2026).

Lesson: Preserve version, threshold, site, recipients, workflow, and endpoint before comparing results.

Scoring Rubric

The following 0–2 scale can structure discussion. It has not been validated to predict product safety or clinical utility.

Criterion 0 Points 1 Point 2 Points
External Validation None or internal only Some separate evaluation Evidence fit to the intended population and claim
Peer-Reviewed Publication No appraisable report Limited methods available Full appraisable peer-reviewed report
Independent Researchers Vendor employees only Partial independence Fully independent validation
Clinical Outcomes None reported Retrospective outcomes Prospective RCT or quasi-experimental
Generalizability Evidence No evidence Similar populations Evaluated in the intended local population and workflow

Scoring: - Higher discussion score: More evidence is available for appraisal. - Lower discussion score: Named evidence gaps remain. - Any material failure: Do not let an average score override a safety, regulatory, privacy, or evidence blocker.


Domain 2: Patient Safety

The Questions to Ask

Critical Safety Questions
  1. Safety Testing
  2. Clinical Outcomes
  3. Alert Burden
  4. Human Factors

Red Flags

Proceed with extreme caution if:

  • “No reported adverse events” - Likely means no monitoring system, not that it’s safe
  • “Clinicians love it” - No quantitative data on alert fatigue or response rates
  • “Seamless integration” - No workflow analysis or physician research
  • False-positive and alert burden not tied to the use case - No universal percentage defines acceptable burden.
  • No outcome data - Only technical performance (AUC) reported
  • “100% accurate” - Overconfident claims; every system has failure modes
  • No failure mode analysis - Every AI system can fail; what happens when it does?

Lessons from Watson for Oncology

Published Watson for Oncology studies mainly measured concordance with local multidisciplinary recommendations. Concordance varied by cancer and setting, and the designs did not establish patient-outcome benefit or one universal rate of unsafe recommendations (Jie et al., 2021; Zhou et al., 2019).

Lesson: Recommendation agreement, safety, and patient benefit are distinct endpoints. Request evidence for the exact claim rather than repeating unsourced incident narratives.

Scoring Rubric

The following 0–2 scale is an unvalidated discussion aid. It does not establish a deployment threshold.

Criterion 0 Points 1 Point 2 Points
Safety Testing None mentioned Some testing Comprehensive failure mode analysis
Outcome Evidence None Retrospective analysis Prospective RCT or quasi-experimental
Alert Burden Unknown Reported without action context Quantified and acceptable for the prespecified workflow
Human Factors No testing Limited usability testing Comprehensive workflow analysis with physicians
Adverse Event Monitoring No system Passive reporting Active surveillance system (like drug safety)

Scoring: - Higher discussion score: More safety evidence is available. - Lower discussion score: Named safety gaps remain. - Any material safety failure: Pause regardless of the total.


Domain 3: Fairness and Equity

The Questions to Ask

Critical Fairness Questions
  1. Bias Testing
  2. Training Data Representativeness
  3. Proxy Variables
  4. Health Equity Impact

Red Flags

Proceed with extreme caution if:

  • “We don’t use race as a feature, so it’s fair” - Fairness through unawareness doesn’t work; race correlated with many features
  • “Our algorithm is objective” - Algorithms encode human biases in historical data
  • “High accuracy means fair” - Accuracy ≠ Fairness (see OPTUM case)
  • No fairness testing - Bias is default; fairness must be tested, not assumed
  • Using costs as proxy for health needs - See OPTUM case: costs reflect access barriers, not just illness severity
  • Trained on non-representative data - Academic medical centers only, commercially insured only, etc.

Lessons from Population-Health Algorithmic Bias

The OPTUM algorithm for care management: - Accurately predicted healthcare costs - High technical performance - But costs ≠ health needs, especially for Black patients - Systematically underestimated Black patients’ health needs - Result: Black patients less likely to receive needed care coordination - Replacing cost with health need increased the proportion of Black patients selected for additional help from 17.7% to 46.5% (Obermeyer et al., 2019).

Lesson: The result was a change in the proportion selected, not “46.5% more”. Test the target, model performance, and downstream allocation across clinically relevant groups.

Scoring Rubric

The following 0–2 scale is an unvalidated discussion aid. It cannot define fairness by itself.

Criterion 0 Points 1 Point 2 Points
Fairness Audit None conducted Internal audit Independent external audit
Subgroup Analysis No reporting Some subgroups Comprehensive (race, ethnicity, age, sex, insurance)
Training Data Diversity Homogeneous (AMCs only) Somewhat diverse Highly representative of your population
Proxy Variable Assessment No assessment Acknowledged Validated against direct clinical outcomes
Equity Impact Plan No plan Monitoring planned Active mitigation strategies for disparities

Scoring: - Higher discussion score: More equity evidence is available. - Lower discussion score: Named target, subgroup, access, or allocation gaps remain. - Any material equity failure: Resolve it before scaling.


Domain 4: Privacy and Security

The Questions to Ask

Critical Privacy and Security Questions
  1. Regulatory Compliance
  2. Data Handling
  3. Security Measures
  4. Privacy by Design
  5. Transparency & Accountability
  6. Data Rights & Ownership

Why this matters: Contract language can allocate rights in derivative works, improvements, feedback, or training artifacts. The institution should inspect the actual terms rather than assume who owns clinician corrections.

Red Flags

Proceed with extreme caution if:

  • Refuses a required BAA - A blocker when the vendor is acting as a HIPAA business associate.
  • Defers HIPAA role analysis - Resolve roles and required agreements before the proposed data flow begins.
  • Vague about data storage location - “Cloud” is not specific enough; which cloud? Which region?
  • Data sent to foreign servers - Compliance and privacy risks
  • “We anonymize data so HIPAA doesn’t apply” - Verify the de-identification method, residual risk, contractual use, and which laws still apply.
  • No SOC 2 or security certification - Unvetted security practices
  • “Trust us with your data” - Trust requires verification
  • Unclear data retention/deletion - Your patients’ data may persist indefinitely
  • Vendor claims rights to “derivative data” or “improvements” - Your corrections become their IP
  • No data export provision - Vendor lock-in; can’t leave without losing your data

Lessons from Streams

Streams used the NHS acute-kidney-injury detection algorithm and a digitally enabled response pathway, not an AI model. The Royal Free data-sharing arrangement was investigated by the UK Information Commissioner’s Office. A separate service evaluation found improved recognition and some process measures, but no significant step change in the primary renal-recovery outcome (Connell et al., 2019).

The case demonstrates that data governance, technical category, workflow performance, and clinical outcomes are separate review domains.

Lesson: Privacy promises should be translated into applicable legal terms, technical controls, and auditable operations. Data minimization should follow the task, applicable law, and institutional policy.

Scoring Rubric

The following 0–2 scale is an unvalidated discussion aid. Certifications and agreements do not by themselves establish security or privacy adequacy.

Criterion 0 Points 1 Point 2 Points
HIPAA Compliance No BAA or refuses Will sign BAA BAA + SOC 2 + HITRUST
Data Minimization Collects everything Some minimization Strict minimization; edge deployment option
Security Certifications None SOC 2 Type I SOC 2 Type II + penetration testing
Transparency Vague policies Clear policies Detailed + third-party audit
Data Control Vendor retains indefinitely Retention period defined You control data; deletion guaranteed

Scoring: - Higher discussion score: More privacy and security evidence is available. - Lower discussion score: Named data-flow, control, or accountability gaps remain. - Any material privacy or security failure: Pause regardless of the total.


Domain 5: Workflow Integration

The Questions to Ask

Critical Workflow Integration Questions
  1. Physician Research
  2. EHR Integration
  3. Training & Support
  4. Customization
  5. Monitoring & Feedback

Red Flags

Proceed with extreme caution if:

  • “Plug and play” - Clinical medicine is complex; no system is truly plug-and-play
  • “Works with all EHRs” - Each EHR integration is custom; this claim is implausible
  • “No training needed” - Training needs depend on intended use and workflow, but consequential tools require users to understand limitations, escalation, and change.
  • “One-size-fits-all” - Different hospitals have different workflows and patient populations
  • An implementation timeline without local discovery - Duration depends on interfaces, governance, validation, training, workflow, and risk.
  • No physician research - Designed in isolation from actual clinical workflows
  • Minimal support - Email-only support; no phone; no dedicated account manager
  • Black box, no customization - Can’t adjust thresholds or workflows to fit your practice

Lessons from Retinal AI Deployment in Thailand

A human-centered study across 11 Thai clinics documented image-quality, connectivity, workflow, and patient-expectation problems. It did not support one universal field-accuracy or ungradable-image percentage, did not study India, and did not report that the pilot was abandoned.

Lesson: Field performance includes acquisition, connectivity, staff behavior, referral capacity, and patient flow (Beede et al., 2020).

Scoring Rubric

The following 0–2 scale is an unvalidated discussion aid.

Criterion 0 Points 1 Point 2 Points
Physician Research None Some physician testing Extensive ethnographic research with your specialty
EHR Integration No integration or manual entry Some EHR support Verified integration with the exact local configuration
Training Program Risks and limitations omitted Basic role training Role-appropriate training with reinforcement and escalation
Customization Black box, no customization Limited adjustments Highly customizable to your workflows
Support Quality Email only Email + phone Dedicated account manager + on-site clinical support

Scoring: - Higher discussion score: More workflow evidence and support are available. - Lower discussion score: Named workflow gaps remain. - Any material failure: Pause until the workflow and safety gap is addressed.


Domain 6: Business Viability

The Questions to Ask

Critical Business Viability Questions
  1. Company Stability
  2. Customer Base
  3. Product Maturity
  4. Regulatory Status
  5. Pricing & Contracts

Red Flags

Proceed with extreme caution if:

  • No credible continuity plan - Company size alone does not determine continuity risk.
  • Cannot provide relevant references or explain discontinued deployments - Limits independent operational verification.
  • Version and change history are unclear - Version number alone does not establish maturity.
  • Vague about pricing - “It depends”; no transparency; potential for unexpected costs
  • Long-term contract with no exit clause - You’re locked in even if it doesn’t work
  • Required regulatory status cannot be verified - The proposed use may fall outside the authorized intended use.
  • Leadership with no healthcare experience - Tech team with no clinical domain expertise
  • Recent layoffs or leadership turnover - Financial instability

Lessons from Watson Procurement

The MD Anderson Oncology Expert Advisor project illustrates that a major institution, major vendor, and substantial investment do not guarantee successful integration or clinical deployment. A JNCI report stated that the project cost $62 million and ended before use on actual patients (Schmidt, 2017). This project should not be conflated with every Watson for Oncology deployment.

Lesson: Company scale is not product evidence. Procurement should examine deliverables, integration, evidence, continuity, and exit.

Scoring Rubric

The following 0–2 scale is an unvalidated discussion aid.

Criterion 0 Points 1 Point 2 Points
Company Stability Continuity unknown Partial continuity evidence Tested continuity, support, and transition plan
Customer Base No relevant references Some comparable references Multiple appraisable references and discontinued-use history
Product Maturity Version unclear Version identified Versioned evidence, change history, and support record
Regulatory Required status unverifiable Status under review Exact record and intended-use match verified
Pricing Transparency Vague or hidden Somewhat clear Fully transparent, fair terms

Scoring: - Higher discussion score: More continuity and procurement evidence is available. - Lower discussion score: Named dependency or exit gaps remain. - Any material failure: Resolve the continuity, regulatory, or contract blocker before proceeding.


Putting It All Together: The Overall Evaluation Matrix

Use this matrix to display discussion scores, not to automate a decision. The weights are illustrative and unvalidated.

Domain Illustrative Weight Discussion Score (0–10) Weighted Display
1. Clinical Validation 30% _____ _____
2. Patient Safety 25% _____ _____
3. Fairness & Equity 20% _____ _____
4. Privacy & Security 15% _____ _____
5. Workflow Integration 5% _____ _____
6. Business Viability 5% _____ _____
TOTAL 100% _____ / 10

Decision Framework

Record one of four conclusions:

  • Proceed to bounded local evaluation: The known evidence and controls justify a prespecified evaluation, but do not yet prove local benefit.
  • Proceed only after named gaps close: Specify the missing evidence, control, owner, and reassessment trigger.
  • Decline: An unacceptable requirement remains unmet.
  • Defer: Information is insufficient for a defensible decision.

A weighted total must never conceal one unacceptable safety, privacy, regulatory, or evidence failure. For each conclusion, record the evidence, unresolved uncertainty, accountable owner, and change or reassessment trigger.


Sample Questions for Vendor Meetings

Use these scripts to extract critical information:

Clinical Validation Questions

Script: > “Can you provide the peer-reviewed publication of your external validation study? We’d like to see performance metrics at hospitals not involved in development, with results stratified by patient demographics and clinical outcomes data.”

Follow-ups if vendor hesitates: - “If there’s no peer-reviewed external validation, when do you plan to conduct one?” - “Can you share the names of hospitals where validation occurred so we can contact physician references?” - “What were the clinical outcomes (mortality, complications, readmissions) at hospitals using this system?”

Fairness Questions

Script: > “We’re committed to health equity. Can you show us the fairness audit results? Specifically, we need sensitivity, specificity, and PPV broken down by race, ethnicity, age, sex, and insurance type.”

Follow-ups: - “If no fairness audit has been done, why not?” - “What is your plan for ongoing bias monitoring after deployment?” - “What happens if we discover bias affecting our patient population after deployment?”

Privacy Questions

Script: > “Walk us through exactly what patient data leaves our hospital, where it goes, how it’s stored, and how we can verify this. Can we see the Business Associate Agreement and SOC 2 Type II report?”

Follow-ups: - “What specific PHI elements does your system need?” - “Can the system work on-premise without sending data to the cloud?” - “What happens to our patients’ data if we terminate the contract?” - “Have there been any data breaches or security incidents?”

Safety Questions

Script: > “What patient outcomes have improved at hospitals using your system? Can you provide data on mortality, complications, length of stay, readmissions, or quality of life from prospective studies?”

Follow-ups if vendor only cites AUC or accuracy: - “AUC is a technical metric. Has deployment demonstrably improved patient outcomes?” - “What is the false positive rate in real-world clinical use?” - “What are the failure modes? What happens when the model fails?” - “Have there been any adverse events or patient harm attributed to the system?” - “Has the product, a prior version, or a closely related predicate device had any FDA recalls, corrections, MAUDE reports, field safety notices, or urgent software patches?”

Workflow Integration Questions

Script: > “Has your system been tested with physicians and nurses at hospitals like ours? What did the usability testing reveal? How much time does it add or save per patient?”

Follow-ups: - “What is the typical implementation timeline from contract signing to go-live?” - “What ongoing clinical support do you provide?” - “Can we speak with 3 attending physician users at other hospitals about their experience?”


Procurement and Contract Discussion Prompts

These are issue-spotting prompts, not contract language or legal advice. Performance warranties, service levels, insurance, indemnification, liability caps, breach notice, data deletion, publication rights, and termination terms depend on the product, jurisdiction, risk, applicable law, and institutional bargaining position. Institutional counsel and procurement teams should draft the actual agreement.

1. Performance Definitions

Define the exact product version, intended use, population, operating point, comparator, reference standard, denominator, uncertainty, and measurement period. State which metrics are technical, workflow, safety, equity, or patient outcomes. The agreement should explain who measures performance, how data quality is handled, what counts as material degradation, and what correction, suspension, or exit options follow.

Avoid inserting sensitivity, specificity, positive predictive value, satisfaction, or outcome thresholds before the institution defines the clinical use and decision consequence. A threshold suitable for one task can be unsafe or infeasible for another.

2. Equity and Access

Identify clinically relevant populations, allocation decisions, access barriers, and subgroup analyses. Specify the data needed to estimate differences with adequate precision, how missing demographic information is handled, who reviews a signal, and what mitigation or pause process applies.

Do not use a universal 10% difference as the definition of fairness. The relevant standard depends on the endpoint, prevalence, intervention, uncertainty, and consequences of error.

3. Data Privacy and Security

Map covered data, HIPAA roles, business-associate obligations where applicable, subprocessors, storage locations, access controls, encryption, retention, deletion, model-training use, secondary use, audit logs, incident response, recovery, and verification rights. A BAA or SOC 2 report answers only part of the review.

Specify when data must be returned or deleted, what must be retained by law, how deletion is verified, and what happens to derived artifacts, fine-tuned components, embeddings, logs, and backups. Security controls and notice periods should follow the institution’s risk assessment and applicable requirements rather than copied fixed values.

4. Validation, Monitoring, and Change

Consider rights to conduct independent evaluation, access sufficient documentation and outputs, publish results subject to legitimate confidentiality constraints, receive performance and incident information, and inspect relevant quality records. Define notification and review for changes to models, prompts, source corpora, thresholds, interfaces, data pipelines, intended use, or regulatory status.

The agreement should distinguish ordinary maintenance from a material change that requires re-evaluation. It should also identify who can pause or roll back the system during investigation.

5. Responsibility, Insurance, and Indemnification

Allocation of responsibility is fact-, contract-, product-, and jurisdiction-specific. Counsel should examine vendor representations, product design and warnings, institutional configuration, clinician workflow, data quality, cybersecurity, third-party components, professional liability, product liability, insurance, indemnification, defense, exclusions, and caps.

It is not defensible to prescribe a universal insurance minimum or a liability cap equal to a fixed multiple of annual contract value. FDA authorization does not determine civil liability, and a contract cannot make every model-associated patient outcome attributable to one party.

6. Continuity and Termination

Define termination for cause, safety concern, regulatory change, material nonperformance, security incident, insolvency, convenience, and failed remediation. Address notice, refund or payment consequences, continued support during transition, data portability, documentation, interface shutdown, record retention, deletion verification, and replacement planning.

An exit right is useful only when the institution can actually retrieve data, stop the workflow safely, and transition care.


Bounded Local Evaluation Plan

The form and duration of local evaluation should follow clinical risk, event rate, sample-size needs, workflow cycles, and the decision to be made. A conventional pilot is often useful, but no universal three-phase calendar applies to every system.

Phase 1: Controlled Evaluation

Scope: - A bounded population and workflow selected from the intended use - Volume sufficient for the prespecified analysis - Intensive monitoring - Dedicated clinical champion

Prespecified decision criteria: - Technical: Metrics and uncertainty appropriate to the task and operating point - Clinical: A comparator and endpoint that match the benefit claim - Human factors: Use, response, workload, comprehension, and override measures - Safety: Defined potential-harm, near-miss, incident, and stopping criteria - Equity: Clinically meaningful subgroup and allocation analyses with uncertainty

Metrics to Track: - Technical performance (sensitivity, specificity, PPV, NPV, AUC) - Alert burden (alerts/day, false positive rate, response time) - Physician experience (satisfaction, time spent per alert, override rate) - Workflow impact (time added/saved per patient) - Clinical outcomes (compare to baseline: mortality, complications, length of stay) - Equity impact (outcomes by race, ethnicity, age, sex, insurance) - Adverse events (any patient harm attributed to AI)

Decision: Proceed, correct and re-evaluate, narrow use, pause, or terminate according to prespecified evidence and safety rules. “Most criteria met” is not adequate when the unmet criterion is a material blocker.

Phase 2: Expanded Evaluation

Scope: - Additional settings justified by the intended-use and transportability question - Continue intensive monitoring - Broader physician engagement

Objectives: - Validate Phase 1 results at larger scale - Test in diverse clinical settings (ICU, floor, ED, outpatient) - Identify implementation challenges - Refine workflows and alert thresholds

Phase 3: Scale Decision

Scope: - Hospital-wide or health system-wide

Requirements: - Evidence supports the intended claim at the proposed scale - Physician training completed for all users - Ongoing monitoring system in place - Risk-based monitoring and re-evaluation triggers are defined - Governance structure for AI oversight

Boundary: A local evaluation does not cure an unauthorized intended use or a missing safety prerequisite. Some low-risk tools may not need a traditional pilot, while higher-risk tools may need comparative clinical study beyond a local pilot.


Constructed Case Study: Using the Checklist

Example: Evaluating a Hypothetical Sepsis Prediction Tool

All vendors, claims, performance values, deployment counts, prices, scores, and institutional recommendations below are synthetic. They illustrate an appraisal process and are not evidence about a real product. The numerical score is a discussion display, not a validated decision threshold.

Vendor Claims: - “AI predicts sepsis 6 hours before clinical recognition” - “92% sensitivity, 87% specificity” - “Deployed in 200+ hospitals” - “$400K/year for hospital-wide license”

Illustrative institutional evaluation:

Domain 1: Clinical Validation (Score: 4/10)

  • Published in peer-reviewed journal
  • Internal validation only (same health system, 3 hospitals)
  • No independent external validation
  • 92% sensitivity in paper, but what about at external sites?
  • No prospective outcome data (mortality, length of stay)
  • Red flag: “Deployed in 200+ hospitals” ≠ Evidence of effectiveness (Epic sepsis model lesson!)

Domain 2: Patient Safety (Score: 3/10)

  • Safety mentioned in paper
  • No prospective outcome studies showing mortality benefit
  • No data on whether deployment reduced deaths or complications
  • False positive rate not clearly reported for real-world use
  • Alert burden unknown
  • Red flag: Only technical metrics (AUC, sensitivity), no patient outcomes

Domain 3: Fairness & Equity (Score: 2/10)

  • No fairness audit mentioned
  • No performance stratified by race/ethnicity
  • When asked, vendor says “we don’t use race as a feature, so it’s fair”
  • Red flag: Fairness through unawareness (doesn’t work!)

Domain 4: Privacy & Security (Score: 7/10)

  • Will sign BAA
  • SOC 2 Type II certified
  • Data encrypted at rest and in transit
  • Data stored in vendor cloud (no on-premise option)
  • No HITRUST certification

Domain 5: Workflow Integration (Score: 5/10)

  • Integrates with your EHR (Epic)
  • Implementation takes 3-6 months
  • Training: 2-hour online module (seems insufficient)
  • No evidence supporting the proposed local threshold
  • No usability testing data shared

Domain 6: Business Viability (Score: 8/10)

  • Established company, 7 years in business
  • 200 hospital customers (they claim)
  • Willing to provide 2 physician references
  • Pricing seems high ($400K/year)
  • Company reports FDA 510(k) clearance, but the team has not yet verified the exact record or intended-use match

Overall Weighted Score: 4.5 / 10

Illustrative decision: Do not deploy until named blockers close - Insufficient clinical validation (internal only, no external validation) - No outcome evidence (AUC/sensitivity are not enough; need mortality/LOS data) - No fairness testing (high bias risk) - Workflow concerns (alert burden unknown, may cause alert fatigue)

Recommendation to Hospital Leadership:

“The evaluation identified unresolved evidence and workflow blockers. The discussion score is not the decision rule.

Key concerns: - No external validation (validation only within vendor’s own health system) - No evidence of improved patient outcomes (only sensitivity/specificity reported, no mortality or LOS data) - No fairness audit (risk of bias affecting minority patients, similar to Epic sepsis model) - High cost ($400K/year) without demonstrated ROI

We recommend: 1. Request external evidence appropriate to the intended population and claim 2. Request fairness audit with performance by race/ethnicity/insurance 3. Request prospective outcome data (mortality, length of stay, time to antibiotics) 4. Verify the FDA record, intended use, model version, and postmarket record 5. Define the evidence that would trigger reassessment

Alternative: Continue the current sepsis pathway while comparing this proposal with other clinical and operational investments.”


Summary: Key Principles for Vendor Evaluation

  1. Match evidence to the claim: Transportability, workflow use, and patient benefit require different designs.
  2. Preserve the exact version and intended use: Evidence does not automatically transfer across changes.
  3. Evaluate the complete pathway: Include failed inputs, alert burden, human response, downstream action, and follow-up capacity.
  4. Examine equity and access: Test the target, performance, allocation, and intervention.
  5. Control change: Define monitoring, incident, re-evaluation, pause, and rollback.
  6. Use institution-specific contracts: Counsel should address evidence, data, security, responsibility, continuity, and exit.
  7. Decline unresolved material risk: A procurement decision should fail closed when intended use, evidence, data flow, accountability, or change control cannot be verified.

The central lesson: An AI label, prestigious vendor, large deployment count, or regulatory authorization does not substitute for evidence that matches the proposed use.


Additional Resources

Related Evaluation Frameworks

This toolkit complements established academic frameworks for evaluating clinical AI:

  • APPRAISE-AI Tool (Kwong et al., JAMA Network Open 2023): Quantitative tool for evaluating AI studies for clinical decision support. Provides structured scoring of study quality.
  • CLIX-M Checklist (NPJ Digital Medicine, 2025): Clinician-informed 14-item checklist for evaluating explainable AI in clinical decision support systems. Focuses on transparency and interpretability.
  • ROBUST-ML Checklist (European Heart Journal Digital Health, 2022): Checklist for ruling out bias using standard tools in machine learning studies. Helps identify methodological red flags.
  • APA AI Tool Evaluation Checklist (APA Practice Organization): Practical checklist for clinicians evaluating AI tools for practice integration.
  • CASoF (Clinical AI Sociotechnical Framework) (JMIRx Med, 2025): 35-item checklist addressing sociotechnical factors in AI implementation.

Regulatory and Professional Resources


What questions should hospitals ask AI vendors?

Critical questions cover the exact product and version, intended use, regulatory record, external validation, endpoint and comparator, subgroup performance, data flows, security controls, workflow testing, monitoring, change control, and termination rights.

What is external validation for medical AI?

External validation evaluates performance in data or settings separate from model development. It addresses transportability, but the number and type of sites should follow the intended-use claim; no universal three-site rule establishes clinical utility.

How do I know if an AI vendor is trustworthy?

Trust cannot be reduced to one credential. Verify the intended use, regulatory record, evidence, data flows, security controls, limitations, version history, incident process, pricing, monitoring access, and willingness to support independent evaluation.

What score should AI vendors achieve on evaluation frameworks?

The six-domain scorecards are author-created discussion aids, not validated decision instruments. Do not add domain scores into an automatic pass-fail total or allow a high average to offset a material safety, privacy, regulatory, or evidence failure.

What contract terms should hospitals require for AI purchases?

Contract terms depend on the product, jurisdiction, risk, and institution. Counsel and procurement teams should consider performance definitions, validation and audit rights, data use and deletion, security, change notice, incident response, allocation of responsibility, insurance, publication rights, continuity, and termination.

Bottom line: The preferred system is the one whose evidence, intended use, workflow, equity, privacy, security, monitoring, and continuity support the actual clinical decision.