Appendix H — Clinical AI Vendor Evaluation Toolkit
Six review domains: Clinical validation, patient safety, fairness and access, privacy and security, workflow and change control, and business continuity.
Decision boundary: The scorecards are discussion aids, not validated procurement instruments. A high average cannot offset a material safety, privacy, regulatory, or evidence failure.
Critical questions: Does the cited evidence match the exact product version and claim? What data leave the institution? Who acts on the output? How are changes, incidents, monitoring, and exit handled?
This appendix provides a structured discussion framework for evaluating commercial AI products in clinical practice. It is not a validated scoring instrument, procurement standard, or substitute for clinical, privacy, security, regulatory, legal, and contracting review.
Who should use this: - Hospital administrators evaluating AI vendors - Department chairs considering AI tools for their specialty - Chief Medical Information Officers (CMIOs) assessing AI products - Physician practice leaders making technology decisions - IRB/ethics committees reviewing AI deployments
What you’ll find: - Structured evaluation scorecards - Red flag identification guides - Sample questions to ask vendors - Decision frameworks for Go/No-Go - Contract discussion prompts for institutional counsel and procurement teams
Introduction: Why Vendor Evidence Needs Structured Review
The Clinical AI Morgue shows that public evidence does not support one aggregate cost or harm total across clinical AI failures. It does show recurring, detectable problems: claim-evidence mismatch, biased targets, version drift, weak transportability, workflow burden, and missing accountability.
Vendor due diligence should identify what is known, what remains uncertain, and what would block use.
Algorithm Transparency in Certified Health IT
For systems embedded in certified health IT, vendor due diligence should include applicable HTI-1 algorithm-transparency documentation. The rule established transparency requirements for predictive decision support interventions within its certified-health-IT scope (ONC, HTI-1 Final Rule). It is not a universal disclosure law for every healthcare AI product.
Procurement teams should ask whether the product is part of an ONC-certified health IT module, whether it qualifies as a decision support intervention or predictive model under the certification program, and what baseline information the vendor can provide to assess fairness, appropriateness, validity, effectiveness, and safety. HTI-1 transparency does not prove clinical benefit. It creates a minimum documentation floor that should be combined with local validation, subgroup performance review, security review, and post-deployment monitoring.
The Vendor-Physician Information Asymmetry
Vendor knows: - Where the model was trained - What its clinical limitations are - Where external validation failed - What patient populations it doesn’t work for - What the false alarm rate is in real-world use - Which hospitals abandoned the system after pilot
The purchasing institution may initially know: - What the vendor provides in sales materials - What can be verified in publications, regulatory records, reference checks, and technical documentation
The review process should close that information gap before a consequential decision.
Quick Reference: The Six-Domain Evaluation Framework
Use this framework to systematically evaluate any clinical AI vendor:
| Domain | Key Questions | Red Flags |
|---|---|---|
| 1. Clinical Validation | External validation? Peer-reviewed publication? Improved outcomes? | Internal validation only; no publications; only AUC reported |
| 2. Patient Safety | Safety testing? Adverse events tracked? Failure mode analysis? | No safety data; no mention of harms; “100% accurate” claims |
| 3. Fairness & Equity | Performance across demographics? Bias audits? | No fairness testing; “we don’t use race so it’s fair” |
| 4. Privacy & Security | HIPAA roles? BAA when required? Data flows? Security controls? | Vague compliance claims; unknown subprocessors, retention, or training use |
| 5. Workflow Integration | End-user testing? EHR integration? Training provided? | No physician research; “plug and play”; minimal training |
| 6. Business Viability | Continuity, support, portability, exit, and exact regulatory status? | Unclear roadmap, dependency risk, no exit or continuity plan |
Domain 1: Clinical Validation
The Questions to Ask
- Training Data
- Validation Studies
- Clinical Performance Metrics
- Clinical Outcomes (MOST IMPORTANT)
Red Flags
Proceed with extreme caution if:
- “Validated on 100,000 patients” - But all from the same institution (not external validation)
- “95% accuracy” - On cherry-picked test set; no external validation
- “AUC 0.92” - But no data on whether clinical outcomes improved
- “Deployed in 150+ hospitals” - Deployment ≠ Effectiveness; no outcome data
- “Proprietary validation” - No peer-reviewed publications; “trust us”
- Internal validation only - This does not answer whether performance transports to the proposed site.
- Vendor-funded validation without transparent methods or independent replication - Funding is not automatic disqualification, but it affects the independence evidence needed.
- Only technical metrics reported - AUC, accuracy without clinical outcomes
- “FDA-cleared” without recall history or postmarket performance data - FDA authorization is not the end of evaluation. A 2025 cross-sectional study of 950 FDA-cleared AI-enabled devices found recalls uncommon (60 devices, 6.3%) but concentrated early, with 43.4% of recall events occurring within 12 months of clearance, and lack of clinical validation independently associated with recall (odds ratio 2.8; 95% CI, 1.6-4.7) (Lee et al., 2025)
Lessons from the Epic Sepsis Evidence
One external evaluation of the Epic Sepsis Model found hospitalization-level AUROC 0.63, sensitivity 33%, and positive predictive value 12% at the evaluated threshold. The reported 7% was a process-of-care subgroup, not a before-onset detection rate, and the study did not establish that alerts caused fatigue or patient harm (Wong et al., 2021). Later prospective version 2 evidence showed stronger discrimination with site variability and substantial alert burden (Wong et al., 2026).
Lesson: Preserve version, threshold, site, recipients, workflow, and endpoint before comparing results.
Scoring Rubric
The following 0–2 scale can structure discussion. It has not been validated to predict product safety or clinical utility.
| Criterion | 0 Points | 1 Point | 2 Points |
|---|---|---|---|
| External Validation | None or internal only | Some separate evaluation | Evidence fit to the intended population and claim |
| Peer-Reviewed Publication | No appraisable report | Limited methods available | Full appraisable peer-reviewed report |
| Independent Researchers | Vendor employees only | Partial independence | Fully independent validation |
| Clinical Outcomes | None reported | Retrospective outcomes | Prospective RCT or quasi-experimental |
| Generalizability Evidence | No evidence | Similar populations | Evaluated in the intended local population and workflow |
Scoring: - Higher discussion score: More evidence is available for appraisal. - Lower discussion score: Named evidence gaps remain. - Any material failure: Do not let an average score override a safety, regulatory, privacy, or evidence blocker.
Domain 2: Patient Safety
The Questions to Ask
- Safety Testing
- Clinical Outcomes
- Alert Burden
- Human Factors
Red Flags
Proceed with extreme caution if:
- “No reported adverse events” - Likely means no monitoring system, not that it’s safe
- “Clinicians love it” - No quantitative data on alert fatigue or response rates
- “Seamless integration” - No workflow analysis or physician research
- False-positive and alert burden not tied to the use case - No universal percentage defines acceptable burden.
- No outcome data - Only technical performance (AUC) reported
- “100% accurate” - Overconfident claims; every system has failure modes
- No failure mode analysis - Every AI system can fail; what happens when it does?
Lessons from Watson for Oncology
Published Watson for Oncology studies mainly measured concordance with local multidisciplinary recommendations. Concordance varied by cancer and setting, and the designs did not establish patient-outcome benefit or one universal rate of unsafe recommendations (Jie et al., 2021; Zhou et al., 2019).
Lesson: Recommendation agreement, safety, and patient benefit are distinct endpoints. Request evidence for the exact claim rather than repeating unsourced incident narratives.
Scoring Rubric
The following 0–2 scale is an unvalidated discussion aid. It does not establish a deployment threshold.
| Criterion | 0 Points | 1 Point | 2 Points |
|---|---|---|---|
| Safety Testing | None mentioned | Some testing | Comprehensive failure mode analysis |
| Outcome Evidence | None | Retrospective analysis | Prospective RCT or quasi-experimental |
| Alert Burden | Unknown | Reported without action context | Quantified and acceptable for the prespecified workflow |
| Human Factors | No testing | Limited usability testing | Comprehensive workflow analysis with physicians |
| Adverse Event Monitoring | No system | Passive reporting | Active surveillance system (like drug safety) |
Scoring: - Higher discussion score: More safety evidence is available. - Lower discussion score: Named safety gaps remain. - Any material safety failure: Pause regardless of the total.
Domain 3: Fairness and Equity
The Questions to Ask
- Bias Testing
- Training Data Representativeness
- Proxy Variables
- Health Equity Impact
Red Flags
Proceed with extreme caution if:
- “We don’t use race as a feature, so it’s fair” - Fairness through unawareness doesn’t work; race correlated with many features
- “Our algorithm is objective” - Algorithms encode human biases in historical data
- “High accuracy means fair” - Accuracy ≠ Fairness (see OPTUM case)
- No fairness testing - Bias is default; fairness must be tested, not assumed
- Using costs as proxy for health needs - See OPTUM case: costs reflect access barriers, not just illness severity
- Trained on non-representative data - Academic medical centers only, commercially insured only, etc.
Lessons from Population-Health Algorithmic Bias
The OPTUM algorithm for care management: - Accurately predicted healthcare costs - High technical performance - But costs ≠ health needs, especially for Black patients - Systematically underestimated Black patients’ health needs - Result: Black patients less likely to receive needed care coordination - Replacing cost with health need increased the proportion of Black patients selected for additional help from 17.7% to 46.5% (Obermeyer et al., 2019).
Lesson: The result was a change in the proportion selected, not “46.5% more”. Test the target, model performance, and downstream allocation across clinically relevant groups.
Scoring Rubric
The following 0–2 scale is an unvalidated discussion aid. It cannot define fairness by itself.
| Criterion | 0 Points | 1 Point | 2 Points |
|---|---|---|---|
| Fairness Audit | None conducted | Internal audit | Independent external audit |
| Subgroup Analysis | No reporting | Some subgroups | Comprehensive (race, ethnicity, age, sex, insurance) |
| Training Data Diversity | Homogeneous (AMCs only) | Somewhat diverse | Highly representative of your population |
| Proxy Variable Assessment | No assessment | Acknowledged | Validated against direct clinical outcomes |
| Equity Impact Plan | No plan | Monitoring planned | Active mitigation strategies for disparities |
Scoring: - Higher discussion score: More equity evidence is available. - Lower discussion score: Named target, subgroup, access, or allocation gaps remain. - Any material equity failure: Resolve it before scaling.
Domain 4: Privacy and Security
The Questions to Ask
- Regulatory Compliance
- Data Handling
- Security Measures
- Privacy by Design
- Transparency & Accountability
- Data Rights & Ownership
Why this matters: Contract language can allocate rights in derivative works, improvements, feedback, or training artifacts. The institution should inspect the actual terms rather than assume who owns clinician corrections.
Red Flags
Proceed with extreme caution if:
- Refuses a required BAA - A blocker when the vendor is acting as a HIPAA business associate.
- Defers HIPAA role analysis - Resolve roles and required agreements before the proposed data flow begins.
- Vague about data storage location - “Cloud” is not specific enough; which cloud? Which region?
- Data sent to foreign servers - Compliance and privacy risks
- “We anonymize data so HIPAA doesn’t apply” - Verify the de-identification method, residual risk, contractual use, and which laws still apply.
- No SOC 2 or security certification - Unvetted security practices
- “Trust us with your data” - Trust requires verification
- Unclear data retention/deletion - Your patients’ data may persist indefinitely
- Vendor claims rights to “derivative data” or “improvements” - Your corrections become their IP
- No data export provision - Vendor lock-in; can’t leave without losing your data
Lessons from Streams
Streams used the NHS acute-kidney-injury detection algorithm and a digitally enabled response pathway, not an AI model. The Royal Free data-sharing arrangement was investigated by the UK Information Commissioner’s Office. A separate service evaluation found improved recognition and some process measures, but no significant step change in the primary renal-recovery outcome (Connell et al., 2019).
The case demonstrates that data governance, technical category, workflow performance, and clinical outcomes are separate review domains.
Lesson: Privacy promises should be translated into applicable legal terms, technical controls, and auditable operations. Data minimization should follow the task, applicable law, and institutional policy.
Scoring Rubric
The following 0–2 scale is an unvalidated discussion aid. Certifications and agreements do not by themselves establish security or privacy adequacy.
| Criterion | 0 Points | 1 Point | 2 Points |
|---|---|---|---|
| HIPAA Compliance | No BAA or refuses | Will sign BAA | BAA + SOC 2 + HITRUST |
| Data Minimization | Collects everything | Some minimization | Strict minimization; edge deployment option |
| Security Certifications | None | SOC 2 Type I | SOC 2 Type II + penetration testing |
| Transparency | Vague policies | Clear policies | Detailed + third-party audit |
| Data Control | Vendor retains indefinitely | Retention period defined | You control data; deletion guaranteed |
Scoring: - Higher discussion score: More privacy and security evidence is available. - Lower discussion score: Named data-flow, control, or accountability gaps remain. - Any material privacy or security failure: Pause regardless of the total.
Domain 5: Workflow Integration
The Questions to Ask
- Physician Research
- EHR Integration
- Training & Support
- Customization
- Monitoring & Feedback
Red Flags
Proceed with extreme caution if:
- “Plug and play” - Clinical medicine is complex; no system is truly plug-and-play
- “Works with all EHRs” - Each EHR integration is custom; this claim is implausible
- “No training needed” - Training needs depend on intended use and workflow, but consequential tools require users to understand limitations, escalation, and change.
- “One-size-fits-all” - Different hospitals have different workflows and patient populations
- An implementation timeline without local discovery - Duration depends on interfaces, governance, validation, training, workflow, and risk.
- No physician research - Designed in isolation from actual clinical workflows
- Minimal support - Email-only support; no phone; no dedicated account manager
- Black box, no customization - Can’t adjust thresholds or workflows to fit your practice
Lessons from Retinal AI Deployment in Thailand
A human-centered study across 11 Thai clinics documented image-quality, connectivity, workflow, and patient-expectation problems. It did not support one universal field-accuracy or ungradable-image percentage, did not study India, and did not report that the pilot was abandoned.
Lesson: Field performance includes acquisition, connectivity, staff behavior, referral capacity, and patient flow (Beede et al., 2020).
Scoring Rubric
The following 0–2 scale is an unvalidated discussion aid.
| Criterion | 0 Points | 1 Point | 2 Points |
|---|---|---|---|
| Physician Research | None | Some physician testing | Extensive ethnographic research with your specialty |
| EHR Integration | No integration or manual entry | Some EHR support | Verified integration with the exact local configuration |
| Training Program | Risks and limitations omitted | Basic role training | Role-appropriate training with reinforcement and escalation |
| Customization | Black box, no customization | Limited adjustments | Highly customizable to your workflows |
| Support Quality | Email only | Email + phone | Dedicated account manager + on-site clinical support |
Scoring: - Higher discussion score: More workflow evidence and support are available. - Lower discussion score: Named workflow gaps remain. - Any material failure: Pause until the workflow and safety gap is addressed.
Domain 6: Business Viability
The Questions to Ask
- Company Stability
- Customer Base
- Product Maturity
- Regulatory Status
- Pricing & Contracts
Red Flags
Proceed with extreme caution if:
- No credible continuity plan - Company size alone does not determine continuity risk.
- Cannot provide relevant references or explain discontinued deployments - Limits independent operational verification.
- Version and change history are unclear - Version number alone does not establish maturity.
- Vague about pricing - “It depends”; no transparency; potential for unexpected costs
- Long-term contract with no exit clause - You’re locked in even if it doesn’t work
- Required regulatory status cannot be verified - The proposed use may fall outside the authorized intended use.
- Leadership with no healthcare experience - Tech team with no clinical domain expertise
- Recent layoffs or leadership turnover - Financial instability
Lessons from Watson Procurement
The MD Anderson Oncology Expert Advisor project illustrates that a major institution, major vendor, and substantial investment do not guarantee successful integration or clinical deployment. A JNCI report stated that the project cost $62 million and ended before use on actual patients (Schmidt, 2017). This project should not be conflated with every Watson for Oncology deployment.
Lesson: Company scale is not product evidence. Procurement should examine deliverables, integration, evidence, continuity, and exit.
Scoring Rubric
The following 0–2 scale is an unvalidated discussion aid.
| Criterion | 0 Points | 1 Point | 2 Points |
|---|---|---|---|
| Company Stability | Continuity unknown | Partial continuity evidence | Tested continuity, support, and transition plan |
| Customer Base | No relevant references | Some comparable references | Multiple appraisable references and discontinued-use history |
| Product Maturity | Version unclear | Version identified | Versioned evidence, change history, and support record |
| Regulatory | Required status unverifiable | Status under review | Exact record and intended-use match verified |
| Pricing Transparency | Vague or hidden | Somewhat clear | Fully transparent, fair terms |
Scoring: - Higher discussion score: More continuity and procurement evidence is available. - Lower discussion score: Named dependency or exit gaps remain. - Any material failure: Resolve the continuity, regulatory, or contract blocker before proceeding.
Putting It All Together: The Overall Evaluation Matrix
Use this matrix to display discussion scores, not to automate a decision. The weights are illustrative and unvalidated.
| Domain | Illustrative Weight | Discussion Score (0–10) | Weighted Display |
|---|---|---|---|
| 1. Clinical Validation | 30% | _____ | _____ |
| 2. Patient Safety | 25% | _____ | _____ |
| 3. Fairness & Equity | 20% | _____ | _____ |
| 4. Privacy & Security | 15% | _____ | _____ |
| 5. Workflow Integration | 5% | _____ | _____ |
| 6. Business Viability | 5% | _____ | _____ |
| TOTAL | 100% | _____ / 10 |
Decision Framework
Record one of four conclusions:
- Proceed to bounded local evaluation: The known evidence and controls justify a prespecified evaluation, but do not yet prove local benefit.
- Proceed only after named gaps close: Specify the missing evidence, control, owner, and reassessment trigger.
- Decline: An unacceptable requirement remains unmet.
- Defer: Information is insufficient for a defensible decision.
A weighted total must never conceal one unacceptable safety, privacy, regulatory, or evidence failure. For each conclusion, record the evidence, unresolved uncertainty, accountable owner, and change or reassessment trigger.
Sample Questions for Vendor Meetings
Use these scripts to extract critical information:
Clinical Validation Questions
Script: > “Can you provide the peer-reviewed publication of your external validation study? We’d like to see performance metrics at hospitals not involved in development, with results stratified by patient demographics and clinical outcomes data.”
Follow-ups if vendor hesitates: - “If there’s no peer-reviewed external validation, when do you plan to conduct one?” - “Can you share the names of hospitals where validation occurred so we can contact physician references?” - “What were the clinical outcomes (mortality, complications, readmissions) at hospitals using this system?”
Fairness Questions
Script: > “We’re committed to health equity. Can you show us the fairness audit results? Specifically, we need sensitivity, specificity, and PPV broken down by race, ethnicity, age, sex, and insurance type.”
Follow-ups: - “If no fairness audit has been done, why not?” - “What is your plan for ongoing bias monitoring after deployment?” - “What happens if we discover bias affecting our patient population after deployment?”
Privacy Questions
Script: > “Walk us through exactly what patient data leaves our hospital, where it goes, how it’s stored, and how we can verify this. Can we see the Business Associate Agreement and SOC 2 Type II report?”
Follow-ups: - “What specific PHI elements does your system need?” - “Can the system work on-premise without sending data to the cloud?” - “What happens to our patients’ data if we terminate the contract?” - “Have there been any data breaches or security incidents?”
Safety Questions
Script: > “What patient outcomes have improved at hospitals using your system? Can you provide data on mortality, complications, length of stay, readmissions, or quality of life from prospective studies?”
Follow-ups if vendor only cites AUC or accuracy: - “AUC is a technical metric. Has deployment demonstrably improved patient outcomes?” - “What is the false positive rate in real-world clinical use?” - “What are the failure modes? What happens when the model fails?” - “Have there been any adverse events or patient harm attributed to the system?” - “Has the product, a prior version, or a closely related predicate device had any FDA recalls, corrections, MAUDE reports, field safety notices, or urgent software patches?”
Workflow Integration Questions
Script: > “Has your system been tested with physicians and nurses at hospitals like ours? What did the usability testing reveal? How much time does it add or save per patient?”
Follow-ups: - “What is the typical implementation timeline from contract signing to go-live?” - “What ongoing clinical support do you provide?” - “Can we speak with 3 attending physician users at other hospitals about their experience?”
Procurement and Contract Discussion Prompts
1. Performance Definitions
Define the exact product version, intended use, population, operating point, comparator, reference standard, denominator, uncertainty, and measurement period. State which metrics are technical, workflow, safety, equity, or patient outcomes. The agreement should explain who measures performance, how data quality is handled, what counts as material degradation, and what correction, suspension, or exit options follow.
Avoid inserting sensitivity, specificity, positive predictive value, satisfaction, or outcome thresholds before the institution defines the clinical use and decision consequence. A threshold suitable for one task can be unsafe or infeasible for another.
2. Equity and Access
Identify clinically relevant populations, allocation decisions, access barriers, and subgroup analyses. Specify the data needed to estimate differences with adequate precision, how missing demographic information is handled, who reviews a signal, and what mitigation or pause process applies.
Do not use a universal 10% difference as the definition of fairness. The relevant standard depends on the endpoint, prevalence, intervention, uncertainty, and consequences of error.
3. Data Privacy and Security
Map covered data, HIPAA roles, business-associate obligations where applicable, subprocessors, storage locations, access controls, encryption, retention, deletion, model-training use, secondary use, audit logs, incident response, recovery, and verification rights. A BAA or SOC 2 report answers only part of the review.
Specify when data must be returned or deleted, what must be retained by law, how deletion is verified, and what happens to derived artifacts, fine-tuned components, embeddings, logs, and backups. Security controls and notice periods should follow the institution’s risk assessment and applicable requirements rather than copied fixed values.
4. Validation, Monitoring, and Change
Consider rights to conduct independent evaluation, access sufficient documentation and outputs, publish results subject to legitimate confidentiality constraints, receive performance and incident information, and inspect relevant quality records. Define notification and review for changes to models, prompts, source corpora, thresholds, interfaces, data pipelines, intended use, or regulatory status.
The agreement should distinguish ordinary maintenance from a material change that requires re-evaluation. It should also identify who can pause or roll back the system during investigation.
5. Responsibility, Insurance, and Indemnification
Allocation of responsibility is fact-, contract-, product-, and jurisdiction-specific. Counsel should examine vendor representations, product design and warnings, institutional configuration, clinician workflow, data quality, cybersecurity, third-party components, professional liability, product liability, insurance, indemnification, defense, exclusions, and caps.
It is not defensible to prescribe a universal insurance minimum or a liability cap equal to a fixed multiple of annual contract value. FDA authorization does not determine civil liability, and a contract cannot make every model-associated patient outcome attributable to one party.
6. Continuity and Termination
Define termination for cause, safety concern, regulatory change, material nonperformance, security incident, insolvency, convenience, and failed remediation. Address notice, refund or payment consequences, continued support during transition, data portability, documentation, interface shutdown, record retention, deletion verification, and replacement planning.
An exit right is useful only when the institution can actually retrieve data, stop the workflow safely, and transition care.
Bounded Local Evaluation Plan
The form and duration of local evaluation should follow clinical risk, event rate, sample-size needs, workflow cycles, and the decision to be made. A conventional pilot is often useful, but no universal three-phase calendar applies to every system.
Phase 1: Controlled Evaluation
Scope: - A bounded population and workflow selected from the intended use - Volume sufficient for the prespecified analysis - Intensive monitoring - Dedicated clinical champion
Prespecified decision criteria: - Technical: Metrics and uncertainty appropriate to the task and operating point - Clinical: A comparator and endpoint that match the benefit claim - Human factors: Use, response, workload, comprehension, and override measures - Safety: Defined potential-harm, near-miss, incident, and stopping criteria - Equity: Clinically meaningful subgroup and allocation analyses with uncertainty
Metrics to Track: - Technical performance (sensitivity, specificity, PPV, NPV, AUC) - Alert burden (alerts/day, false positive rate, response time) - Physician experience (satisfaction, time spent per alert, override rate) - Workflow impact (time added/saved per patient) - Clinical outcomes (compare to baseline: mortality, complications, length of stay) - Equity impact (outcomes by race, ethnicity, age, sex, insurance) - Adverse events (any patient harm attributed to AI)
Decision: Proceed, correct and re-evaluate, narrow use, pause, or terminate according to prespecified evidence and safety rules. “Most criteria met” is not adequate when the unmet criterion is a material blocker.
Phase 2: Expanded Evaluation
Scope: - Additional settings justified by the intended-use and transportability question - Continue intensive monitoring - Broader physician engagement
Objectives: - Validate Phase 1 results at larger scale - Test in diverse clinical settings (ICU, floor, ED, outpatient) - Identify implementation challenges - Refine workflows and alert thresholds
Phase 3: Scale Decision
Scope: - Hospital-wide or health system-wide
Requirements: - Evidence supports the intended claim at the proposed scale - Physician training completed for all users - Ongoing monitoring system in place - Risk-based monitoring and re-evaluation triggers are defined - Governance structure for AI oversight
Boundary: A local evaluation does not cure an unauthorized intended use or a missing safety prerequisite. Some low-risk tools may not need a traditional pilot, while higher-risk tools may need comparative clinical study beyond a local pilot.
Constructed Case Study: Using the Checklist
Example: Evaluating a Hypothetical Sepsis Prediction Tool
Vendor Claims: - “AI predicts sepsis 6 hours before clinical recognition” - “92% sensitivity, 87% specificity” - “Deployed in 200+ hospitals” - “$400K/year for hospital-wide license”
Illustrative institutional evaluation:
Domain 1: Clinical Validation (Score: 4/10)
- Published in peer-reviewed journal
- Internal validation only (same health system, 3 hospitals)
- No independent external validation
- 92% sensitivity in paper, but what about at external sites?
- No prospective outcome data (mortality, length of stay)
- Red flag: “Deployed in 200+ hospitals” ≠ Evidence of effectiveness (Epic sepsis model lesson!)
Domain 2: Patient Safety (Score: 3/10)
- Safety mentioned in paper
- No prospective outcome studies showing mortality benefit
- No data on whether deployment reduced deaths or complications
- False positive rate not clearly reported for real-world use
- Alert burden unknown
- Red flag: Only technical metrics (AUC, sensitivity), no patient outcomes
Domain 3: Fairness & Equity (Score: 2/10)
- No fairness audit mentioned
- No performance stratified by race/ethnicity
- When asked, vendor says “we don’t use race as a feature, so it’s fair”
- Red flag: Fairness through unawareness (doesn’t work!)
Domain 4: Privacy & Security (Score: 7/10)
- Will sign BAA
- SOC 2 Type II certified
- Data encrypted at rest and in transit
- Data stored in vendor cloud (no on-premise option)
- No HITRUST certification
Domain 5: Workflow Integration (Score: 5/10)
- Integrates with your EHR (Epic)
- Implementation takes 3-6 months
- Training: 2-hour online module (seems insufficient)
- No evidence supporting the proposed local threshold
- No usability testing data shared
Domain 6: Business Viability (Score: 8/10)
- Established company, 7 years in business
- 200 hospital customers (they claim)
- Willing to provide 2 physician references
- Pricing seems high ($400K/year)
- Company reports FDA 510(k) clearance, but the team has not yet verified the exact record or intended-use match
Overall Weighted Score: 4.5 / 10
Illustrative decision: Do not deploy until named blockers close - Insufficient clinical validation (internal only, no external validation) - No outcome evidence (AUC/sensitivity are not enough; need mortality/LOS data) - No fairness testing (high bias risk) - Workflow concerns (alert burden unknown, may cause alert fatigue)
Recommendation to Hospital Leadership:
“The evaluation identified unresolved evidence and workflow blockers. The discussion score is not the decision rule.
Key concerns: - No external validation (validation only within vendor’s own health system) - No evidence of improved patient outcomes (only sensitivity/specificity reported, no mortality or LOS data) - No fairness audit (risk of bias affecting minority patients, similar to Epic sepsis model) - High cost ($400K/year) without demonstrated ROI
We recommend: 1. Request external evidence appropriate to the intended population and claim 2. Request fairness audit with performance by race/ethnicity/insurance 3. Request prospective outcome data (mortality, length of stay, time to antibiotics) 4. Verify the FDA record, intended use, model version, and postmarket record 5. Define the evidence that would trigger reassessment
Alternative: Continue the current sepsis pathway while comparing this proposal with other clinical and operational investments.”
Summary: Key Principles for Vendor Evaluation
- Match evidence to the claim: Transportability, workflow use, and patient benefit require different designs.
- Preserve the exact version and intended use: Evidence does not automatically transfer across changes.
- Evaluate the complete pathway: Include failed inputs, alert burden, human response, downstream action, and follow-up capacity.
- Examine equity and access: Test the target, performance, allocation, and intervention.
- Control change: Define monitoring, incident, re-evaluation, pause, and rollback.
- Use institution-specific contracts: Counsel should address evidence, data, security, responsibility, continuity, and exit.
- Decline unresolved material risk: A procurement decision should fail closed when intended use, evidence, data flow, accountability, or change control cannot be verified.
The central lesson: An AI label, prestigious vendor, large deployment count, or regulatory authorization does not substitute for evidence that matches the proposed use.
Additional Resources
Related Evaluation Frameworks
This toolkit complements established academic frameworks for evaluating clinical AI:
- APPRAISE-AI Tool (Kwong et al., JAMA Network Open 2023): Quantitative tool for evaluating AI studies for clinical decision support. Provides structured scoring of study quality.
- CLIX-M Checklist (NPJ Digital Medicine, 2025): Clinician-informed 14-item checklist for evaluating explainable AI in clinical decision support systems. Focuses on transparency and interpretability.
- ROBUST-ML Checklist (European Heart Journal Digital Health, 2022): Checklist for ruling out bias using standard tools in machine learning studies. Helps identify methodological red flags.
- APA AI Tool Evaluation Checklist (APA Practice Organization): Practical checklist for clinicians evaluating AI tools for practice integration.
- CASoF (Clinical AI Sociotechnical Framework) (JMIRx Med, 2025): 35-item checklist addressing sociotechnical factors in AI implementation.
Regulatory and Professional Resources
- Appendix E (The Clinical AI Morgue): Detailed failure case studies showing what goes wrong
- FDA AI/ML Medical Device Guidance: https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices
- Coalition for Health AI (CHAI): https://www.chai.org/
- AMIA Clinical Informatics: https://amia.org/education-events/clinical-informatics-board-review-course
- AMA AI Governance Guidance (8 steps for health system AI success): https://www.ama-assn.org/practice-management/digital-health/8-steps-position-your-health-system-ai-success
What questions should hospitals ask AI vendors?
Critical questions cover the exact product and version, intended use, regulatory record, external validation, endpoint and comparator, subgroup performance, data flows, security controls, workflow testing, monitoring, change control, and termination rights.
What is external validation for medical AI?
External validation evaluates performance in data or settings separate from model development. It addresses transportability, but the number and type of sites should follow the intended-use claim; no universal three-site rule establishes clinical utility.
How do I know if an AI vendor is trustworthy?
Trust cannot be reduced to one credential. Verify the intended use, regulatory record, evidence, data flows, security controls, limitations, version history, incident process, pricing, monitoring access, and willingness to support independent evaluation.
What score should AI vendors achieve on evaluation frameworks?
The six-domain scorecards are author-created discussion aids, not validated decision instruments. Do not add domain scores into an automatic pass-fail total or allow a high average to offset a material safety, privacy, regulatory, or evidence failure.
What contract terms should hospitals require for AI purchases?
Contract terms depend on the product, jurisdiction, risk, and institution. Counsel and procurement teams should consider performance definitions, validation and audit rights, data use and deletion, security, change notice, incident response, allocation of responsibility, insurance, publication rights, continuity, and termination.
Bottom line: The preferred system is the one whose evidence, intended use, workflow, equity, privacy, security, monitoring, and continuity support the actual clinical decision.