AI Tools Every Physician Should Know
FDA maintains a periodically updated AI-enabled medical-device list and cautions that it is not comprehensive (FDA AI-Enabled Medical Devices). The list is a discovery surface, not evidence that a device improves workflow, equity, or patient outcomes. Verify the exact authorization record, version, indication, and decision documents.
After completing this chapter, clinicians should be able to:
- Identify FDA-cleared and clinically validated AI tools by specialty
- Understand capabilities and limitations of each tool category
- Evaluate tools appropriate for your practice setting
- Navigate privacy, liability, and reimbursement considerations
- Distinguish evidence-based tools from marketing hype
- Access hands-on resources for learning AI tools
Introduction: Navigating the AI Tool Landscape
FDA’s periodically updated AI-enabled medical device list is large and growing, with hundreds more tools marketed without FDA oversight (clinical decision support, wellness applications). For physicians, the challenge is not finding AI tools, it’s identifying which ones actually work, have solid evidence, integrate into workflows, and provide clinical value (FDA AI-Enabled Medical Devices).
Verify FDA Status at the Device-Record Level
Company names and algorithm families are not FDA records. Each retained product claim should resolve to the exact device, version, pathway, intended use, and decision documents.
| Example record | Pathway | What the record establishes |
|---|---|---|
| IDx-DR, DEN180001 | De Novo | The authorized autonomous diabetic-retinopathy device, indication, classification, and decision documents |
| IDx-DR v2.3, K213037 | 510(k) | A later cleared version and its substantial-equivalence record |
| Paige Prostate, DEN200080 | De Novo | An adjunct prostate-cancer detection device, not an automated Gleason-grading authorization |
| InVision Precision LVEF, K232331 | 510(k) | The labeled echocardiography software and intended use |
FDA authorization does not establish superiority, local workflow benefit, equitable performance, reimbursement, or improved patient outcomes.
Category 1: Clinical Decision Support
Traditional CDS (Pre-AI Era)
UpToDate - Type: Evidence-based clinical reference - AI Features: The product has added AI-assisted retrieval and question-answering functions; exact availability depends on the current subscription and product configuration - Evidence: A historical observational study associated access to the underlying clinical reference with lower risk-adjusted mortality, length of stay, and cost, but it did not isolate a modern generative-AI feature (Isaac et al., 2012) - Cost: Institutional and individual prices change; verify the current official quote rather than relying on a historical list price - Strength: Curated clinical reference content with regular editorial updates - Limitation: Evidence for the reference product should not be transferred automatically to a newer AI feature
DynaMedex with Dyna AI - Type: Clinical decision support combining DynaMed (evidence-based clinical information) and Micromedex (drug information) with AI integration - AI Features: Dyna AI surfaces concise, evidence-based clinical information quickly from exclusively DynaMedex sources - Cost: Access terms depend on the current institutional or membership arrangement - Access: The product has been included in the ACP AI Resource Hub; current eligibility should be verified with ACP
OpenEvidence - Type: AI copilot for evidence-based clinical decision support at point of care - Function: Natural language clinical questions answered with evidence from peer-reviewed literature (NEJM, JAMA, Cochrane, specialty journals) - Adoption: The company reported 757,000 verified physician registrations and more than 1 million consultations on March 10, 2026. These vendor-reported activity measures do not establish accuracy, patient benefit, or the proportion of U.S. physicians who use the product routinely (OpenEvidence, March 2026) - Evidence: A published evaluation reported favorable clinician ratings for evaluated answers in primary-care cases. Its design, cases, raters, endpoints, and product version should be reviewed before translating the result into clinical-performance claims (PubMed record, 2026) - Independent benchmark: A 2026 Nature Medicine study tested OpenEvidence and UpToDate Expert AI against GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 using 500 MedQA questions, 500 HealthBench items, and 100 de-identified real clinical queries (1,800 model-query annotations with clinician review). Frontier LLMs outperformed both specialized tools across all three stages. UpToDate Expert AI and OpenEvidence performed similarly on the real clinical query stage and similarly to Google AI Overview in that arm (Vishwanath et al., 2026) - Systematic review: A 2026 npj Digital Medicine review, the first systematic review of OpenEvidence, included 11 peer-reviewed studies through January 2026. Where hallucination was measured, fabricated-citation rates were lower than for general-purpose LLMs, and guideline-derived accuracy reached 82% on 130 neurology items and 94.4% (119/126) on osteoarticular infection items (Artsi et al., 2026). On 15 structural-heart questions, fully accurate answers were 26.7% for OpenEvidence versus 66.7% for ChatGPT-4o; the platform often reinforced rather than changed management, 11 heterogeneous small studies are thin relative to claimed daily use, and no patient-outcome evidence was found. - Market context: Trade reporting described a $12 billion valuation after a $250 million financing round in January 2026. Valuation is not clinical evidence and can become stale (STAT News, January 2026) - Cost: Subscription-based; institutional licensing available - Strength: Source-linked retrieval and a workflow designed for clinical questions - Limitation: Requires subscription; independent head-to-head evidence does not show superiority over frontier general-purpose LLMs; evidence synthesis quality depends on underlying literature availability - Verdict: A prominent clinical retrieval product whose cited answers, version-specific performance, data handling, and local workflow effects still require verification
DXplain (Massachusetts General Hospital) - Type: Differential diagnosis generator - Function: Enter findings → generates ranked differential - Evidence: Used since 1980s; comparative studies show solid diagnostic accuracy, scoring 3.45/5 tied with Isabel for top DDx generator performance (Bond et al., 2012); original system description (Barnett et al., 1987) - Cost: Free for medical professionals - Strength: Broad knowledge base - Limitation: Does not narrow differential without clinical judgment
Isabel Healthcare - Type: Differential diagnosis support - Function: Enter patient presentation → suggests diagnoses - Evidence: A comparative evaluation gave Isabel and DXplain the same mean usefulness score of 3.45/5 in the study’s cases (Bond et al., 2012). The company separately reports 96% inclusion of the correct diagnosis among suggestions; that vendor metric requires its original methods, case mix, and denominator before use - Cost: Subscription-based - Limitation: Accuracy variable, requires clinical interpretation
Modern AI-Enhanced CDS
Epic Sepsis Model - Type: EHR-integrated sepsis prediction - Function: Real-time risk score based on vital signs, labs - Evidence: An external validation reported 33% sensitivity and 12% positive predictive value at the evaluated threshold (Wong et al., 2021). The study assessed model performance, not causal patient benefit - Cost: Included with Epic EHR - Strength: Integrated workflow - Limitation: High false positive rates, mixed evidence for clinical benefit - Verdict: Use with caution, understand limitations
WAVE Clinical Platform (ExcelMedical/Hillrom) - Type: Continuous vital sign monitoring + early warning scores - Function: ICU/step-down monitoring, deterioration prediction - FDA record: WAVE Clinical Platform, K171056, 510(k) decision January 4, 2018; the record should be read for the exact monitoring and medical-device-data-system functions - Evidence: UPMC step-down unit study found every medical emergency team event was preceded by an elevated monitoring index, by a mean of 6.3 hours, in a design where staff were blinded to the index and no outcome effect was measured (Hravnak et al., 2008); a later phase of the same program found reduced frequency and duration of serious instability once nurses acted on index alerts (Hravnak et al., 2011) - Cost: Institutional licensing - Use case: Hospital early warning systems
Category 2: Diagnostic AI
Radiology AI (Comprehensive List)
Intracranial Hemorrhage Detection
Aidoc markets multiple radiology products. Company-reported deployment counts describe adoption, not clinical effectiveness. Each hemorrhage, pulmonary embolism, fracture, or abdominal CT claim must remain connected to the intended use and version in its own FDA record; K252970 cannot authorize or validate the full portfolio.
Voter et al. reported 92.3% sensitivity and 97.7% specificity for the evaluated ICH-detection setting (Voter et al., 2021). A separate evaluation across 17 facilities and 101,944 examinations reported lower overall sensitivity, illustrating why performance should not be transferred across settings (2025 multicenter evaluation). Integration, alert routing, contract price, and economic value require local measurement. A hypothetical avoided miss is not a measured return on investment.
Viz.ai combines product-specific detection and care-coordination functions. Observational stroke-workflow studies report time changes in defined implementations (Figurelle et al., 2023). Randomized and observational evidence should be labeled separately, and stroke evidence should not be transferred to pulmonary embolism or other modules.
The platform has expanded to other time-critical conditions. Current indications and versions require exact FDA-record review. Vendor-reported hospital counts and contract terms should be dated and should not be used as evidence of patient benefit.
RapidAI offers neuroradiology products for tasks including ASPECTS, perfusion, and hemorrhage analysis. A CADTH health technology assessment reviewed 13 diagnostic-accuracy studies (NCBI NBK611329). Product-specific evidence in that review is observational or diagnostic unless a trial is explicitly identified; it does not establish universal treatment-time or outcome benefit.
Chest X-Ray AI
Lunit INSIGHT CXR addresses multiple chest-radiograph findings. A head-to-head validation study reported an AUC of 0.93 for lung-nodule detection in its evaluated dataset. The result should remain tied to that task, population, version, operating point, and comparator; it does not establish cross-task or cross-site superiority.
The procurement question is whether the exact product, version, indication, and operating point match the local task. A publication count is not a substitute for evaluating the connected product-study record.
Oxipit ChestLink is designed for autonomous reporting of selected chest radiographs classified as having no abnormality. In the retrospective evaluation reported by Plesner et al., the evaluated algorithm had 99.1% sensitivity for abnormal findings and 99.8% sensitivity for critical findings, while the comparison and operating conditions were specific to that study (Plesner et al., 2023). Those results do not by themselves establish suitability for autonomous use in another jurisdiction, population, acquisition environment, or radiology workflow. Verify the current regulatory record, labeled exclusions, failure pathway, and required oversight where the product would be deployed.
qXR from Qure.ai addresses multiple chest-radiograph findings. A UK service evaluation of 522 radiographs reported 96.88% sensitivity and 75.55% specificity for abnormal classification, with predictive values shaped by the study’s high proportion of normal examinations (BJR|Open, 2024). The RADICAL study evaluated a different lung-cancer-screening setting (BMJ Open, 2024). Transportability depends on prevalence, spectrum, acquisition, reference standard, and workflow. Pricing and value require current local data.
Mammography AI
iCAD ProFound AI is a mammography AI product. A retrospective study of 419 patients screened with mammography plus digital breast tomosynthesis reported an AUC of 0.93 at the evaluated rule-out setting, while human double-reading performed better in that study (Resch et al., 2024). Integration claims should be verified against the local equipment and current product version.
This evidence supports product-specific appraisal, not a universal ranking. Procurement should compare the local population, reading model, recall and detection endpoints, workflow, and cost against relevant alternatives.
Lunit INSIGHT MMG has been evaluated in multiple screening settings. A JAMA Oncology comparison reported study-specific performance across AI systems (Salim et al., 2020). International evaluation can inform transportability, but it does not establish local performance without matching acquisition, screening program, reference standard, workflow, and subgroup uncertainty.
Hologic Genius AI offers native integration with Hologic mammography equipment. Native integration can reduce some workflow friction, but it does not establish comparative accuracy or clinical benefit. Review the exact study, device record, software version, local equipment, and comparator before selecting it over a standalone product.
Other Imaging Modalities
Tempus Pixel Cardio (formerly Arterys) (Cardiac MRI, CT Angiography) - FDA status: The product lineage has multiple device records; verify the exact current record rather than relying on a company-level “first” claim - Function: Automated cardiac chamber quantification, vessel analysis - Note: Arterys was acquired by Tempus Labs; the cardiac MRI AI product is now called Tempus Pixel Cardio (FDA-cleared K203744) - Evidence: Multi-vendor comparison study validated performance (Ruijsink et al., 2022) - Deployment: Academic medical centers, cardiology practices - Appraisal: Verify the current owner, product name, record, version, and intended use before applying older Arterys evidence to Tempus Pixel Cardio
HeartFlow FFR-CT - FDA status: Verify the exact current HeartFlow product and authorization record; do not infer the pathway from the company name - Function: CT-based fractional flow reserve (non-invasive) - Evidence: PRECISE randomized 2,103 participants to a precision strategy or usual testing and reported fewer catheterizations without obstructive coronary disease in the precision-strategy arm; the composite safety endpoint was noninferior, which should not be rewritten as proof of benefit for every downstream outcome (Douglas et al., 2023). PACIFIC reported per-vessel diagnostic performance in a separate prospective comparison (Danad et al., 2019) - Reimbursement: Billing codes do not guarantee payer coverage or payment - Economic evidence: Evaluate the FISH&CHIPS design, local pathway, downstream testing, and costs before transferring its conclusion - Appraisal: Strong product-specific evidence should still be separated into diagnostic, management, safety, outcome, and economic endpoints
Paige Prostate is an adjunct prostate-cancer detection product authorized through De Novo classification under DEN200080. The intended use is adjunct detection, not autonomous Gleason grading. Perincheri et al. evaluated product-specific pathology performance rather than patient outcomes (Perincheri et al., 2021).
The product highlights suspicious regions for pathologist review. Human review is part of the intended workflow, but the phrase “human in the loop” does not itself establish control effectiveness. Deployment is an adoption signal, not proof of reduced false negatives or improved outcomes.
PathAI and Proscia illustrate different pathology uses. FDA Drug Development Tool qualification for PathAI’s AIM-MASH supports a defined drug-development context of use; it is not authorization for clinical diagnosis. Pulaski et al. reported the connected validation evidence (Pulaski et al., 2025). Proscia’s platform study addressed multi-site data in its evaluated setting (Ianni et al., 2020).
Dermatology AI
3Derm has been described by the company as receiving Breakthrough Device Designation for an autonomous skin-cancer function. Breakthrough designation is not marketing authorization. Current authorization and availability should be checked in FDA’s primary records rather than inferred from a dated absence claim. Dermatology models also require prespecified performance analysis across skin tones and acquisition conditions.
SkinVision and other direct-to-consumer dermatology apps have heterogeneous evidence. Freeman et al. identified small, high-bias evaluations with widely varying performance among available apps (Freeman et al., 2020). Availability and European regulatory status change, so both should be checked in current official records. A consumer-app result should not be transferred to clinician-acquired dermoscopy or another product.
Ophthalmology AI
IDx-DR / LumineticsCore was the first autonomous AI diagnostic system authorized by FDA through De Novo classification, under DEN180001. The pivotal study was prospective and single-arm, not randomized, and reported 87.2% sensitivity and 90.7% specificity for its prespecified diabetic-retinopathy endpoint (Abràmoff et al., 2018).
Factors relevant to adoption include a narrow intended use and a specified autonomous workflow, access to retinal acquisition, image-quality requirements, referral capacity, and prospective evaluation in primary care. A CPT code does not establish a fixed payment amount, payer coverage, or local economic value.
Deployment has expanded in primary care offices, endocrinology clinics, and federally qualified health centers. For practices seeing diabetic patients with limited ophthalmology referral access, this system warrants consideration.
EyeArt from Eyenuk is a separate autonomous diabetic-retinopathy product with its own FDA record and evidence. Ipp et al. reported product-specific performance (Ipp et al., 2021); a larger study addressed a different setting and endpoint (Bhaskaranand et al., 2019). RetCAD has European evidence for retinal-image analysis (González-Gonzalo et al., 2020). Current CE and U.S. status should be verified in official databases rather than through product comparison language.
Category 3: Documentation and Ambient Scribe AI
Regulatory Scope Depends on the Exact Function
Ambient documentation systems primarily address language and workflow tasks. Whether a function falls within FDA’s device definition depends on its exact intended use and behavior. A documentation feature should not be used to infer the status of diagnostic, treatment, coding, or ordering functions in the same platform.
Nuance DAX (Dragon Ambient eXperience) is widely deployed in the ambient scribe market. The system listens to patient encounters, transcribes conversations, extracts clinical information, and auto-generates SOAP note drafts. Physicians review, edit, and sign the generated notes.
Vendor-funded studies have reported large documentation-time reductions, while independent studies show smaller and heterogeneous effects. A cohort study at Intermountain Health found documentation time decreased from 5.3 to 4.54 minutes per encounter (Haberle et al., 2024). A Stanford study of 45 physicians reported changes in documentation, after-hours, and total EHR time (Ma et al., 2025). Each result is product-, setting-, and endpoint-specific.
A head-to-head randomized trial comparing two ambient scribe products with control involved 238 outpatient physicians and reported a 9.5% documentation-time reduction for Nabla versus control, while the DAX estimate was smaller and not statistically significant (Lukac et al., 2025). A separate randomized trial at UW Health evaluated Abridge and reported changes in documentation and work-outside-work time (Afshar et al., 2025). Secondary burnout findings should remain labeled secondary.
Ambient scribe adoption is broad, but current price and setup terms vary by product, volume, integration, and contract. Use the written proposal and include implementation, training, support, monitoring, and change-management costs.
Critical caveat: generated notes require effective review. These systems can omit, add, misattribute, or misstate information. Editing can be faster in some workflows, but verification burden must be measured rather than assumed.
Abridge creates patient-shareable visit summaries in addition to clinical documentation. Recordings and written summaries allow patients to review encounters afterward. A quality improvement study at University of Kansas Medical Center found clinicians using Abridge were 7x more likely to find their workflow easy and 5x more likely to complete notes before the next patient visit; 67% reported feeling less at risk of burnout (Albrecht et al., JAMIA Open 2025). A randomized crossover trial showed 46.6% reduction in cognitive load as measured by NASA-TLX (Hudson et al., Mayo Clin Proc Digit Health 2025). Cost is competitive with DAX.
Suki adds features beyond transcription, including order placement, ICD/CPT code lookup, and voice-enabled EHR navigation. A peer-reviewed validation study assessed note quality using the modified PDQI-9 metric across general medicine, pediatrics, OB/GYN, orthopedics, and cardiology, showing high interrater agreement for most specialties (Palm et al., Front Artif Intell 2025). The AAFP Innovation Laboratory found 60% of participating family physicians adopted the solution after a 30-day trial. Deployment is growing, particularly among physicians seeking complete voice-enabled workflows.
Doximity GPT is a physician-facing assistant within the Doximity platform. Plans, access, data handling, and capabilities change. HIPAA status depends on the exact service, account, relationship, configuration, and contract. The tool has been described as supporting documentation, letters, education, clinical questions, and coding, each of which requires its own evidence and review controls.
Doximity acquired Pathway Medical for $63 million to integrate clinical datasets and AI capabilities (FierceHealthcare, 2025). Doximity’s AI suite includes Scribe (ambient documentation, launched as a free product in July 2025 after over a year of beta testing with 10,000+ physicians), GPT (clinical assistant), and Pathway (clinical datasets) (STAT News, July 2025).
Clinical branding and security infrastructure do not establish clinical validity. Corporate revenue and margin are not evidence for documentation accuracy, diagnostic value, or workflow safety and should not guide clinical selection.
Cost: Free for Doximity members (Scribe has tiered pricing for heavy users) Strength: Low-friction access may support a bounded evaluation for eligible users Limitation: Requires Doximity account; Scribe features require separate subscription for high-volume use; independent clinical validation studies not yet published Appraisal: Verify current limits and total workflow cost, then evaluate the exact function before patient-data use
DeepScribe and Freed AI target smaller practices. DeepScribe offers ambient transcription for primary care and specialty clinics. Freed AI has a free tier, making it accessible for solo practitioners or small groups testing ambient documentation.
Epic AI Charting (Art) launched in February 2026 as a native ambient documentation feature embedded directly within Epic’s EHR. The system drafts notes and queues orders from clinician-patient conversations, with full access to medications, problem lists, and prior history. Epic controls 42% of the acute hospital EHR market and 55% of hospital beds (STAT News, February 2026), giving AI Charting immediate distribution to health systems already running Epic. Planned expansions include diagnosis-aware notes and bedside nursing workflows. For health systems evaluating ambient scribe vendors, the native integration may reduce the need for separate contracts, though independent clinical validation studies are not yet available.
Implementation note: Ambient scribe technology can reduce documentation burden in some workflows. Start with a prespecified local evaluation and maintain review of generated content. Menz et al. evaluated a vision-enabled prototype and reported fewer omission errors for tasks requiring visual input, but this does not establish that all commercial products are audio-only or that the prototype is ready for clinical deployment (Menz et al., 2026).
Implementation Considerations:
Benefits: - Potential changes in documentation time and after-hours work, measured locally - Improved patient eye contact - Reduced burnout - After-hours documentation reduced
Considerations: - Requires physician review (AI makes errors) - Patient consent for recording - Privacy/security review for the exact product, workflow, account, configuration, and contract - Cost (ROI depends on time saved, productivity gains) - Learning curve (initial weeks slower as physician adapts)
Category 4: Literature Search and Synthesis
PubMed / MEDLINE (with AI enhancements) - Free, covers most clinical scenarios - New features: AI-powered search refinement (limited) - Verdict: Still the gold standard, but time-consuming
Consensus - Function: AI searches 220+ million peer-reviewed papers, compiles findings - Use case: Quick literature review, evidence synthesis - Evidence: A systematic review found limited peer-reviewed validation; researchers raised concerns about accuracy and transparency (Apata et al., Cureus 2025) - Cost: Free tier, paid for advanced features - Verdict: Useful for rapid evidence gathering; verify key findings independently
Elicit - Function: AI research assistant - finds papers, extracts key info from 125+ million papers - Use case: Literature review, research questions - Evidence: Validation studies show average sensitivity of 39.5% (systematic reviews require ≥90%); useful as complement, not replacement for traditional searching (Lau & Golder, Cochrane Evid Synth Methods 2025) - Cost: Free tier, paid plans - Verdict: Helpful for scoping searches; insufficient alone for systematic reviews
Scite.ai - Function: Citation analysis - shows how papers cite each other (supporting, contrasting); 1.5 billion citation statements indexed - Use case: Evaluating strength of evidence, finding contradictory studies - Evidence: Technical validation showed citation matching F-score of 95.4% (Nicholson et al., Quant Sci Stud 2021); however, independent evaluation found low accuracy for classifying supporting vs. contrasting citations (F-measures 0.0-0.58) (Bakker et al., Hypothesis 2023) - Cost: Subscription (now part of Research Solutions) - Verdict: Valuable for finding citation context; interpret supporting/contrasting classifications with caution
ResearchRabbit - Function: Literature mapping, citation networks (uses PubMed for medical sciences, Semantic Scholar for other fields) - Note: Acquired by Litmaps in November 2025 - Cost: Free - Verdict: Excellent for exploring research landscapes; no peer-reviewed validation studies
Connected Papers - Function: Visual citation networks using co-citation and bibliographic coupling analysis - Use case: Finding related papers - Cost: Free tier (limited to 5 graphs/month) - Verdict: Great visualization tool for exploratory discovery; not suitable for systematic reviews due to coverage limitations
Category 5: Patient Communication
Patient Education
ChatGPT / GPT-4 (with extreme caution) - Capabilities: Generate patient education materials, explain diagnoses - Evidence: Med-PaLM was evaluated on medical question-answering benchmarks and clinician-rated answers (Singhal et al., 2023). Licensing-exam performance does not establish that a consumer chatbot produces accurate, comprehensible, equitable, or safe patient education in a live workflow - Critical limitations: - Hallucinates (makes up plausible-sounding false information) - No access to patient-specific data - Accountability and liability depend on the parties, workflow, representations, and jurisdiction - May generate outdated or incorrect guidance - Appropriate use: - Draft patient education materials (physician reviews/edits) - Simplify complex medical concepts (verify accuracy) - NOT for patient-specific medical advice - Verdict: A drafting aid only when the content is reviewed against current sources and the patient’s actual context; not autonomous patient advice
Google MedGemma - Medical-specific open-weight models whose available sizes, licenses, and checkpoints should be verified in the current documentation - Evolution: Med-PaLM (2022) → Med-PaLM 2 → Med-Gemini → MedGemma (May 2025) - Evidence: Benchmark results are model- and task-specific. Results from Med-PaLM or Med-PaLM 2 should not be transferred to a MedGemma checkpoint without a connected evaluation (Singhal et al., 2023) - Status: Available to developers under model-specific terms. The models are not finished clinical devices, and any medical-device status depends on the downstream intended use and implementation (Google Health AI Developer Foundations) - Verdict: A development resource for controlled evaluation; clinical deployment requires a complete, validated system and applicable regulatory review
Symptom Checkers (Patient-Facing)
Ada Health - Function: Symptom assessment, triage guidance - Evidence: Vignette study of eight symptom apps showed Ada had the highest condition coverage (99.0%), listed the correct diagnosis in its top 3 in 70.5% of vignettes (GP average 82.1%), and gave safe urgency advice in 97.0%, matching the GP average (Gilbert et al., BMJ Open 2020); the study was designed and conducted by a team including Ada Health employees - Use case: Patient triage (ED vs. urgent care vs. PCP) - Verdict: Triage tool, not diagnostic
Buoy Health - Function: Symptom checker + care navigation (developed at Harvard Innovation Laboratory) - Evidence: Published independent validation is limited; the company has reported a 95% safe-triage metric, which should not be used without the original cases, reference standard, denominator, and definition of safety - Partnerships: The company has described health-system navigation relationships; current partners and deployment scope require verification - Verdict: A patient-navigation product requiring independent, use-case-specific validation
K Health - Function: AI symptom assessment + telemedicine - Model: Subscription and visit terms change; verify the current official offering and distinguish the software interaction from licensed clinical care - Evidence: The company has cited a large proprietary data corpus. Corpus size is not evidence of accuracy or benefit, and independent peer-reviewed validation of the exact patient-facing workflow remains necessary - Verdict: An integrated digital-care model whose triage, clinician-access, safety, privacy, and continuity claims require separate appraisal
Caution on Symptom Checkers: - Accuracy limited (patients may not describe symptoms accurately) - Liability unclear if patients rely on recommendations - Best use: Triage, not diagnosis - Physicians should be cautious recommending specific tools
Category 6: Specialty-Specific Tools
Cardiology
HeartFlow FFR-CT (covered above)
Caption Care (GE HealthCare) - FDA Clearance: Yes (510(k) K190887; acquired by GE HealthCare in February 2023) - Function: AI-guided cardiac ultrasound acquisition - Use case: Point-of-care echo by non-experts - Evidence: In a prospective study, nurses without prior sonography experience used the guidance software to acquire limited transthoracic echocardiograms that were evaluated for diagnostic quality (Narang et al., 2021). The result supports the defined acquisition workflow, not universal equivalence to an expert sonographer across examinations or diagnoses - Verdict: A guided-acquisition option whose operator training, exam scope, quality controls, downstream interpretation, and local pathway should be evaluated
Eko Analysis - FDA record: Eko Murmur Analysis Software, K213794, is a Class II 510(k)-cleared decision-support device for evaluation of heart sounds; later Eko records cover different functions and versions - Function: Digital stethoscope + AI murmur detection - Use case: Primary care screening for valvular heart disease - Evidence: A deep-learning algorithm evaluated with digital-stethoscope recordings reported 93.2% sensitivity and 86.0% specificity for moderate-or-greater aortic stenosis and 66.2% sensitivity and 94.6% specificity for moderate-or-greater mitral regurgitation in the study population (Chorba et al., 2021). Connect the study algorithm, software version, indication, and device record before attributing these results to a current commercial configuration - Verdict: A decision-support option for product-specific evaluation, not a replacement for indicated diagnostic assessment
Oncology
Tempus - Function: Genomic sequencing, AI-driven treatment matching, clinical data analysis - Use case: Precision oncology decision support - Evidence: Beaubier et al. reported analytical validation of the xT next-generation-sequencing assay (Beaubier et al., 2019). That evidence concerns assay performance and should not be transferred to every treatment-matching, clinical-trial-matching, or AI function under the Tempus brand - Verdict: A precision-oncology platform whose assay, software function, evidence, and regulatory status should be evaluated separately
Foundation Medicine - Function: Comprehensive genomic profiling (CGP) - Use case: Cancer treatment selection based on tumor mutations - Evidence: FoundationOne CDx received FDA approval (P170019) as companion diagnostic for multiple targeted therapies; analytical validation published (Frampton et al., Nat Biotechnol 2013) - FDA status: FoundationOne CDx, P170019, and FoundationOne Liquid CDx, P190032, are PMA-approved devices with indication-specific supplements. FoundationOne Liquid CDx is not a 510(k)-cleared product - Verdict: Established FDA-authorized genomic-profiling products whose current indications, specimen requirements, limitations, and companion-diagnostic labels must be checked in the applicable PMA record
IBM Watson for Oncology (historical product)
- Status: The product was discontinued. Retrospective studies reported variable concordance with local multidisciplinary recommendations across cancer types and settings, and they generally did not test patient-outcome benefit
- Lesson: Product visibility and concordance studies do not substitute for prospective evidence of clinical utility.
Emergency Medicine
Viz.ai suite (covered above - stroke, PE)
Epic Deterioration Index - Function: Patient deterioration prediction - Evidence: Variable - some validation, implementation challenges - Cost: Included with Epic - Verdict: Use with caution, understand limitations
Gastroenterology
Medtronic GI Genius - FDA Clearance: Yes (De Novo DEN200055, April 2021; first AI device for colonoscopy in US) - Function: Real-time AI-assisted polyp detection during colonoscopy - Evidence: Randomized controlled trial showed approximately 2-fold reduction in colorectal neoplasia miss rate with AI-assisted colonoscopy (Wallace et al., Gastroenterology 2022); meta-analysis of 5 RCTs (4,354 patients) confirmed ADR improvement (36.6% vs. 25.2%; relative risk 1.44) (Hassan et al., Gastrointestinal Endoscopy 2020) - Use case: Improving adenoma detection during screening/surveillance colonoscopy - Verdict: Product-specific randomized evidence supports improved detection endpoints in the evaluated colonoscopy workflows; patient-outcome and local implementation claims require their own evidence
Evaluating AI Tools for a Clinical Practice
Step 1: Identify Clinical Need
Define the intended improvement before reviewing products: - Name the clinical or operational problem - Document the present workflow, baseline performance, and affected population - Specify whether the desired endpoint is diagnostic performance, timeliness, workload, patient experience, cost, safety, or a patient outcome - Identify who would act differently when the system produces an output
Step 2: Evidence Review
Essential questions: - Does the function meet the device definition, and if so, what exact FDA record applies? - Does each publication evaluate the same product, version, intended use, population, comparator, and endpoint? - Is the evidence retrospective, prospective, comparative, randomized, or postdeployment? - Does external validation test transportability across relevant institutions, populations, equipment, and workflows? - What local evidence is needed before the intended claim is defensible?
Step 3: Workflow Assessment
Integration: - EHR-integrated or standalone? - Number of clicks? - Time added or saved? - Who operates it? (Physician, MA, nurse?)
Step 4: Financial Analysis
Costs: - Licensing fees (annual, per-study, per-patient) - Hardware (servers, cameras, specialized equipment) - Personnel (training, IT support, clinical champions) - Maintenance and updates
ROI: - Measured time saved and whether it becomes usable capacity or cash savings - Reimbursement, including current payer rules rather than the existence of a code alone - Quality metrics (value-based care bonuses) - Measured safety or risk changes, without assuming fewer malpractice claims - Patient satisfaction (retention, referrals)
Step 5: Pilot Testing
Before full deployment: - Select a local evaluation design that matches the claim; retrospective testing may be useful but is not sufficient for every workflow or outcome claim - Use a bounded pilot with prespecified users, escalation rules, endpoints, and stop criteria - Collect feedback (physician, patient, staff) - Measure relevant baseline and postimplementation outcomes with uncertainty and an appropriate comparator - Identify failure modes, automation bias, subgroup risks, downtime procedures, and responsibility for action
Step 6: Continuous Monitoring
Post-deployment: - Review at a frequency justified by clinical risk, volume, change rate, and time to detect harm, not an arbitrary universal calendar interval - User feedback collection - False positive/negative tracking - Clinical outcome monitoring - Vendor support responsiveness
Enterprise AI Platforms for Healthcare Organizations
While consumer health AI products like ChatGPT Health target patients directly, a parallel category of enterprise healthcare AI platforms is marketed for deployment within health systems. Vendors may offer enterprise controls, Business Associate Agreements (BAAs), and clinical-workflow integrations. Those features support due diligence but do not make every configuration, connector, data flow, or use case compliant.
OpenAI for Healthcare (January 2026)
OpenAI launched OpenAI for Healthcare in January 2026, a set of enterprise products designed for healthcare organizations, including ChatGPT for Healthcare and API access with BAA support.
Key distinction from ChatGPT Health:
| Product | Target | HIPAA Status | BAA Available |
|---|---|---|---|
| ChatGPT Health (consumer) | Patients directly | HIPAA role depends on parties and function | Verify current terms |
| ChatGPT for Healthcare (enterprise) | Healthcare organizations | Supports covered-entity workflows | BAA pathway described by vendor |
ChatGPT for Healthcare features:
- Evidence retrieval with citations: Responses grounded in peer-reviewed research, clinical guidelines, and public health guidance with transparent citations including titles, journals, and publication dates
- Institutional policy integration: Connects with enterprise tools (Microsoft SharePoint) to incorporate organization-approved policies and care pathways
- Reusable templates: Shared templates for discharge summaries, patient instructions, prior authorization support
- Role-based access controls: Centralized workspace with SAML SSO, SCIM user management
- Data controls: Audit logs, customer-managed encryption keys, data residency options
- Epic chart connector (September 2026): Vendor-announced read-only FHIR/OAuth access to authorized chart context in supported Epic environments; no write-back (OpenAI, September 2026). See HIPAA considerations.
HIPAA and compliance:
- Verify BAA eligibility and scope for the purchased service
- Verify current training, retention, encryption, deletion, access, and subcontractor terms in the contract
- A BAA does not certify security, clinical performance, FDA status, or lawful use
- See HIPAA considerations for workflow-specific analysis
Early hospital partners (as of January 2026):
OpenAI lists eight hospital partners, though independent verification varies:
- Boston Children’s Hospital (verified): John Brownstein (SVP/Chief Innovation Officer) confirmed adoption of ChatGPT Team with custom OpenAI-powered solution and governance foundations
- UCSF (verified): Chancellor’s announcement confirms ChatGPT Enterprise deployment for 9,000 users in early 2026
- Stanford Medicine Children’s Health, AdventHealth, HCA Healthcare, Baylor Scott & White Health, Cedars-Sinai Medical Center, Memorial Sloan Kettering Cancer Center (listed by OpenAI; independent press releases not located as of this writing)
Underlying models: GPT-5.2 and GPT-5.4
ChatGPT for Healthcare runs on GPT-5.2. OpenAI subsequently released ChatGPT for Clinicians, a specialized product configuration running on GPT-5.4 (a distinct frontier model released March 2026), designed specifically for clinical professional use cases. Performance claims cite:
- HealthBench scores (vendor-developed benchmark; see caveats in the Evaluation chapter)
- HealthBench Professional scores for ChatGPT for Clinicians (GPT-5.4): 59.0 overall vs. 43.7 for physicians on 525 clinician tasks (vendor-developed; evaluation conversations were written by physicians testing the product during development)
- GDPval performance (internal OpenAI benchmark claiming 70.9% win/tie rate versus human baselines, not “better across every role” as sometimes characterized)
Clinical evidence: Penda Health study
OpenAI cites a Penda Health preprint from Kenya. The preprint describes one product version and a nonrandomized analysis; its safety findings and effects should not be transferred to a later version or treated as causal (OpenAI et al., 2025, preprint).
- Reported association: 16% relative reduction in diagnostic errors and 13% reduction in treatment errors across 39,849 visits; the design does not establish that the tool caused the difference
- Critical caveats:
- Alert-response findings require the preprint’s exact denominator and product-version context
- Do not convert retrospective reviewer judgments about preventability into a claim that the AI would have prevented deaths
- Clinicians using AI spent more time per patient (16.4 vs. 13.0 minutes median)
- Nonrandomized design
- Preprint, not peer-reviewed
- Pragmatic randomized evidence: A cluster-randomized trial involving 103 clinician clusters and 9,702 encounters found no significant improvement in the primary 14-day treatment-failure endpoint (2.0% versus 2.2%, adjusted odds ratio 0.77, 95% CI 0.55–1.08, P=.13) (Agweyu et al., 2026)
When evaluating enterprise healthcare AI platforms:
- Verify BAA terms: What exactly is covered? What are the liability provisions?
- Understand data handling: Where is PHI stored? Who has access? Is it used for any training?
- Check independent validation: Vendor benchmarks (HealthBench, GDPval) are not substitutes for peer-reviewed clinical validation
- Assess integration requirements: What EHR/enterprise tool integration is needed?
- Calculate total cost of ownership: Licensing, implementation, training, ongoing support
The same due diligence applied to any clinical system applies to enterprise AI platforms.
Claude for Healthcare (Anthropic, January 2026)
Anthropic launched Claude for Healthcare in January 2026, introducing HIPAA-ready enterprise tools for healthcare providers and payers. The announcement followed OpenAI’s healthcare launch by one week and was made at the J.P. Morgan Healthcare Conference.
Key distinction from consumer Claude:
| Product | Target | HIPAA Status | BAA Available |
|---|---|---|---|
| Claude Pro/Max (consumer) | Individual users | HIPAA role depends on parties and function | Verify current terms |
| Claude for Healthcare (enterprise) | Healthcare organizations | Supports covered-entity workflows | BAA pathway described by vendor |
Healthcare connectors:
Claude for Healthcare includes connectors that allow Claude to pull information from industry-standard systems:
- CMS Coverage Database: Local and National Coverage Determinations for prior authorization verification, coverage requirements, and claims appeals
- ICD-10: Diagnosis and procedure code lookup via CMS and CDC data for medical coding and billing accuracy
- National Provider Identifier Registry: Provider verification, credentialing, and claims validation
- PubMed: Access to 35+ million biomedical literature citations for literature reviews and evidence retrieval
Agent skills:
- FHIR development: Skill for building interoperable healthcare data exchanges using the HL7 FHIR standard
- Prior authorization review: Sample skill template (customizable to organization policies) for cross-referencing coverage requirements, clinical guidelines, and patient records
Use cases highlighted by Anthropic:
- Prior authorization: Pull coverage requirements, check clinical criteria against patient records, propose determinations with supporting materials
- Claims appeals: Assemble documentation from patient records, coverage policies, clinical guidelines for appeal preparation
- Care coordination: Triage patient portal messages, identify urgent items, track referrals and handoffs
HIPAA and compliance:
- Verify BAA eligibility and scope for the exact enterprise service
- Verify current connector, memory, retention, training, access, and deletion behavior in the contract and documentation
- A vendor statement about data handling does not establish compliant local configuration or clinical validity
Early adopter (verified):
- Banner Health: The vendor and trade press have described deployment. Reported user counts and survey impressions are adoption signals, not independent measures of accuracy or patient outcomes (Fierce Healthcare, 2026)
Additional partners listed by Anthropic (independent confirmation not located): Stanford Healthcare, Novo Nordisk, Sanofi, AbbVie, Genmab
Performance benchmarks:
Anthropic reports Claude Opus 4.5 with extended thinking (64k tokens) achieves:
- MedCalc-Bench: Anthropic reported 98.1% for a specified model and configuration. The original benchmark evaluated 55 calculator tasks, and company-reported results require independent replication before clinical transfer (Khandekar et al., 2024). See Medical Calculation Limitations
- MedAgentBench: 91.4% on Stanford’s medical agent benchmark. MedAgentBench tests LLM agent capabilities across 300 clinically-derived tasks in a realistic EHR environment (Jiang et al., NEJM AI 2025)
Life sciences features:
Anthropic also expanded Claude for Life Sciences with connectors to Medidata (clinical trial data), ClinicalTrials.gov, bioRxiv/medRxiv, ChEMBL, and Open Targets. These features target pharmaceutical R&D and clinical trial operations rather than clinical care delivery.
API Access with BAA
Organizations building custom healthcare AI applications can access OpenAI’s API with BAA coverage:
- API access to GPT-5.2 models
- Eligible customers can apply for BAA through OpenAI
- Used by ambient documentation companies (Ambience, EliseAI; note that Abridge uses primarily proprietary models)
- Enterprise API customers contact account teams for access
Consumer Health AI: What Your Patients Are Using
Beyond clinical AI tools that physicians deploy, a growing category of consumer health AI products is reaching patients directly. Understanding these tools helps physicians contextualize patient-reported AI recommendations and have informed conversations about their use.
ChatGPT Health (OpenAI, January 2026)
OpenAI launched ChatGPT Health as a dedicated consumer health experience within ChatGPT. Its scale makes the product relevant to clinical conversations, but use volume is not evidence of safety or benefit.
What it does:
ChatGPT Health allows users in supported configurations to connect health information and wellness services. OpenAI has reported high health-query volume, but vendor usage estimates are not evidence of safety, accuracy, or clinical benefit.
Key features:
- Medical record integration: Connection to medical records via b.well’s health data network (U.S. only; supports Epic, Cerner, Meditech EHRs)
- Wellness app integrations: Apple Health (iOS required), Function, MyFitnessPal, Weight Watchers, AllTrails, Instacart, Peloton
- Personalized responses: Conversations grounded in user’s connected health data
- Separate health space: Health conversations isolated from regular ChatGPT chats with separate memories
- Health insurance navigation: The vendor describes support for plan, claim, billing, and coverage questions; current scope and accuracy require verification
Development approach:
OpenAI collaborated with 262 physicians across 60 countries and 26 medical specialties during development. The company created HealthBench, an open-source evaluation framework with 5,000 multi-turn health conversations and 48,562 rubric criteria developed with physician input (OpenAI et al., 2025, preprint). See AI Evaluation and Validation for detailed analysis of HealthBench methodology.
Privacy and security:
- Health conversations not used for model training (by default)
- Purpose-built encryption and isolation for health data
- Conversations stored separately from non-health chats
- Multi-factor authentication recommended
- Important: A consumer product can operate outside HIPAA because it is not acting for a covered entity. That is different from a universal product-level conclusion that every use is “HIPAA compliant” or “not HIPAA compliant.”
Clinical considerations:
| Aspect | Details |
|---|---|
| FDA status | Not FDA-cleared; positioned as wellness/information tool, not medical device |
| HIPAA | Applicability depends on parties and function; consumer use may fall outside HIPAA |
| Availability | U.S. medical records; excluding EEA, Switzerland, UK |
| Target users | Consumers, not healthcare providers |
Patients may present with AI-interpreted lab results, health recommendations, or questions derived from ChatGPT Health conversations. Consider:
- Ask about AI tool usage: “Have you looked up anything about this online or with AI tools?”
- Review source material: If patient shares AI-generated interpretation, compare to actual clinical data
- Provide context: Explain limitations of consumer AI tools for medical decision-making
- Document appropriately: Note when clinical discussion addresses patient-generated AI content
- Stay current: ChatGPT Health and similar tools will evolve; understand what patients are using
The goal is not to dismiss patient engagement with health AI, but to ensure clinical decisions are based on appropriate medical evaluation.
Independent research context:
A Vanderbilt summary describes peer-reviewed research on how evaluated LLMs translated qualitative risk terms into numbers. The finding is model- and prompt-specific and should be traced to the primary paper before quoting exact ranges (Vanderbilt University Medical Center, 2026). Physicians should expect AI-generated risk language to require contextualization and source verification.
Claude Health Integrations (Anthropic, January 2026)
Anthropic introduced consumer health data integrations for Claude Pro and Max subscribers alongside its enterprise healthcare launch (Anthropic, January 2026).
Health data connections (U.S. only, beta):
- Apple Health: iOS integration for fitness and health metrics (rolling out)
- Android Health Connect: Android equivalent for health data access (rolling out)
- HealthEx: Lab results and health records connector (available in beta)
- Function: Health data integration (available in beta)
What users can do:
- Summarize medical history from connected records
- Explain lab results in plain language
- Detect patterns across fitness and health metrics
- Prepare questions for medical appointments
Privacy protections:
- Users must explicitly opt in to enable access
- Users can disconnect or edit permissions at any time
- Health data excluded from Claude’s memory
- Data not used for model training
Clinical considerations:
| Aspect | Details |
|---|---|
| FDA status | Not FDA-cleared; wellness/information positioning |
| HIPAA | Applicability depends on parties and function; consumer use may fall outside HIPAA |
| Target users | Consumers (Pro/Max subscribers), not providers |
As with ChatGPT Health, patients may present with Claude-interpreted health information. The same guidance applies: ask about AI tool usage, review source material, and ensure clinical decisions are based on appropriate medical evaluation.
Other Consumer Health AI Products
Apple Health AI features: Integrated health insights within Apple’s ecosystem, leveraging Apple Watch and iPhone sensor data.
Google Health initiatives: Various consumer health tools including AI-powered features in Google Fit and experimental health applications.
Symptom checker apps: Babylon Health, Ada Health, K Health, and others provide AI-driven symptom assessment with varying levels of validation.
Direct-to-consumer lab interpretation: Services that apply AI to interpret lab results outside clinical context.
| Consumer Health AI | Clinical AI |
|---|---|
| Direct to patients | Deployed by healthcare systems |
| HIPAA may not apply to direct consumer use | Covered-entity workflow may require a BAA and additional controls |
| Consumer information or wellness positioning | Device status depends on exact function and intended use |
| Self-reported symptoms | Clinical data integration |
| Oversight and escalation vary by product | Clinical oversight and action pathways should be specified |
Consumer and clinical AI differ in relationship, data flow, oversight, and intended use, but the table is not a legal classification. Verify the exact product and workflow.
Red Flags in Clinical AI Procurement
A mismatch between the marketed function and the exact FDA record. Some low-risk or nondevice functions do not require device authorization, so absence from an FDA list is not an automatic rejection. A diagnostic or treatment claim that should be regulated, however, requires immediate record-level review.
No evidence matched to the intended claim. A vendor white paper can describe implementation, and not every useful workflow has a randomized trial. The concern is a benefit claim whose product, version, population, comparator, endpoint, or methods cannot be inspected.
No relevant transportability evidence. Single-site development does not prove failure elsewhere, but it increases the need for external and local evaluation across the intended population, equipment, acquisition, and workflow.
Performance data or failure behavior cannot be reviewed. Procurement teams should be able to inspect operating points, denominators, uncertainty, exclusions, subgroup results, reference standards, and known failure modes.
A single accuracy number or replacement claim substitutes for a defined task. Statements such as “99.9% accurate” or “replaces physicians” are uninterpretable without the unit of analysis, prevalence, threshold, comparator, and workflow.
Data relationships are unclear. The contract and architecture should identify data roles, permitted uses, retention, training, access, deletion, incident response, subprocessors, and downstream sharing.
References cannot describe failures as well as successes. Current customers should be able to explain alert burden, downtime, integration effort, override behavior, monitoring, and unmet expectations, not only satisfaction.
Integration burden is omitted from the value case. Complex integration is not necessarily disqualifying, but its cost, delay, maintenance, human factors, and new failure modes belong in the decision.
No prespecified clinical or operational value proposition. A solution without a measurable baseline, endpoint, accountable owner, and stop criterion is difficult to evaluate after deployment.
These are investigation signals, not a universal pass-fail score. Their importance depends on the function, risk, evidence claim, and deployment context.
Questions About Clinical AI Tools
Resources for Staying Current
FDA AI Device Database: https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices
Medical AI Research: - npj Digital Medicine (Nature) - The Lancet Digital Health - JAMA Network Open (AI sections) - Radiology: Artificial Intelligence
Professional Organizations: - Society for Imaging Informatics in Medicine (SIIM) - American Medical Informatics Association (AMIA) - Radiological Society of North America (RSNA) AI sessions
Conferences: - RSNA (annual AI showcase) - HIMSS (health IT focus) - ML4H (Machine Learning for Health - NeurIPS workshop)
AI Safety Evaluation Tools:
- Petri (Anthropic, open-source, 2025): Automated behavioral audit tool for evaluating LLM safety characteristics including sycophancy, deception, power-seeking, and crisis handling. Uses multi-turn conversations with simulated users to probe model behavior across diverse scenarios. Useful for comparing safety profiles across different LLMs. Note: Developed by Anthropic; interpret results with awareness that tool design may favor Claude models (Anthropic Research, 2025)
- Bloom (Anthropic, open-source, 2025): Complementary tool that generates in-depth evaluation suites for specific behaviors, quantifying severity and frequency. Benchmarks available for delusional sycophancy, instructed long-horizon sabotage, self-preservation, and self-preferential bias (Anthropic Research, 2025)
The Clinical Bottom Line
Match evidence to the claim: Regulatory authorization, diagnostic performance, workflow effects, patient outcomes, and economic value are different questions
Start with bounded applications: Autonomous diabetic-retinopathy screening, ambient documentation, and selected radiology tasks have product-specific evidence, not class-wide proof
Evaluate the local setting: Compare the intended population, equipment, data, users, action pathway, and workflow with the evidence
Calculate measured value: Separate time, capacity, quality, reimbursement, cost, and risk rather than combining assumptions into a single return estimate
Pilot before full deployment: Test on your data, collect feedback, identify failures
Avoid red flags: No evidence, no transparency, too-good-to-be-true claims
Continuous monitoring essential: Performance can drift, vigilance required
Patient communication matters: Transparency, consent, addressing concerns
Map responsibility explicitly: Clinical, institutional, vendor, product, and contractual duties are fact- and jurisdiction-specific; FDA authorization does not decide civil liability
Field evolving rapidly: Stay current, re-evaluate tools regularly
Next Chapter: Large Language Models (ChatGPT, GPT-4, MedGemma) for clinical practice: capabilities, limitations, and safe usage guidelines.
Explore by Specialty
The toolkit gives you the framework. These chapters show how it plays out in practice:
- Radiology & Nuclear Medicine: the most mature AI application in medicine
- Cardiology: ECG interpretation, arrhythmia detection, risk prediction
- Emergency Medicine: triage support, sepsis alerts, imaging triage
- Oncology: precision medicine, treatment matching, pathology AI
See also the Clinical Case Study Library for real-world deployments and lessons from AI failures in practice.