Neurology and Neurological Surgery

Neurology AI has its strongest clinical evidence in selected time-sensitive imaging and workflow applications. A Viz.ai-specific cluster-randomized trial shortened treatment intervals without significantly improving 90-day functional independence. Suicide-risk models illustrate a different problem: a high area under the curve can coexist with low positive predictive value and dangerous false reassurance. Product, task, endpoint, workflow, and outcome must remain connected.

Learning Objectives

After reading this chapter, you will be able to:

  • Evaluate AI systems for acute stroke detection and triage, including large vessel occlusion (LVO) and intracranial hemorrhage (ICH) algorithms
  • Critically assess AI applications in neuroimaging, including brain MRI lesion detection, tumor segmentation, and neurodegenerative imaging
  • Understand automated seizure detection systems and their role in epilepsy management
  • Analyze AI tools for neurodegenerative disease diagnosis and progression monitoring (Parkinson’s, Alzheimer’s)
  • Evaluate the ALS/MND AI landscape including blood biomarker diagnostics, wearable digital biomarkers, brain-computer interfaces for communication, and AI drug discovery
  • Recognize the profound limitations of psychiatric AI, including failed suicide prediction algorithms
  • Apply evidence-based frameworks for evaluating neurology AI before clinical adoption
  • Navigate the ethical complexities unique to neurological and psychiatric AI applications

The Clinical Context:

Neurology combines objective data (neuroimaging, EEG, EMG, cerebrospinal fluid analysis) with subjective clinical assessment (mental status examination, motor exam, gait analysis, cognitive testing). This duality creates both opportunities and challenges for AI.

Opportunities: Stroke imaging analysis, intracranial hemorrhage detection, seizure pattern recognition, and brain tumor segmentation leverage well-defined imaging patterns that AI excels at recognizing.

Challenges: Neurodegenerative disease diagnosis requires integrating subtle examination findings with patient history. Psychiatric diagnosis relies on symptom reporting and clinical judgment where biomarkers are absent. These domains resist algorithmic approaches.

The result: Neurology AI shows tremendous success in time-sensitive acute imaging interpretation but struggles with complex diagnosis, prognostication, and psychiatric applications.

Key Applications:

  • Large vessel occlusion (LVO) detection. Viz.ai has randomized and observational workflow evidence. RapidAI has separate observational evidence. These effects do not transfer to Brainomix without its own source.
  • Intracranial hemorrhage (ICH) detection. Performance is product- and study-specific. Worklist reprioritization can shorten review time without making the output a diagnosis.
  • Brain tumor segmentation. Automated contouring can support radiation-planning workflows, but time savings, correction burden, and intended use are product- and site-specific
  • Pre-surgical brain mapping (resting-state fMRI). Cirrus K251009 supports a defined adjunctive mapping workflow. Regulatory authorization does not establish the uncited usability comparison or superiority over task-based fMRI.
  • Automated seizure detection. Performance depends on the seizure type, recording environment, device, threshold, and reference standard. Automated review remains an adjunct to expert EEG interpretation
  • Multiple sclerosis lesion detection. Variable performance across MRI scanners, useful for tracking but not diagnosis
  • Alzheimer’s neuroimaging. Hippocampal volumetry and amyloid PET quantification predict MCI→dementia conversion but do not improve outcomes
  • Parkinson’s motor assessment. Smartphone/smartwatch apps for objective motor tracking, modest accuracy, not diagnostic
  • ALS blood biomarker AI. Plasma proteomics (AUC 98.3%, Nature Medicine 2025) and gene expression panels (AUC 0.894) for ALS diagnosis; F-wave AI from routine EMG data (88% accuracy, Brain 2025). None clinically deployed
  • ALS wearable monitoring. Accelerometer-derived digital biomarkers outperform ALSFRS-R for detecting progression; speech biomarkers detect decline before clinical scales
  • Brain-computer interfaces for ALS. BrainGate achieved 97.5% accuracy at approximately 32 words per minute in one participant with ALS in a 2024 report; real-time voice synthesis has also been demonstrated in small-participant studies. These are investigational systems, not population-level effectiveness evidence
  • Stroke recurrence prediction. ML models show AUC 0.70-0.75, minimal improvement over CHADS-VASc and clinical scores
  • Suicide risk prediction. Multiple algorithmic failures, dangerous false reassurance from low-risk predictions, should not be clinically deployed
  • Depression diagnosis from facial analysis or voice. No validated clinical applications, privacy and consent concerns

Where Evidence Is Strongest:

  1. Acute stroke workflow: Viz.ai-specific randomized evidence supports shorter treatment intervals, but not a statistically significant improvement in 90-day functional independence.
  2. Clinically validated seizure detection: The ILAE-IFCN guideline conditionally supports validated wearables for generalized and focal-to-bilateral tonic-clonic seizures in selected unsupervised patients when an alarm can lead to rapid intervention. It does not recommend current devices for other seizure types (Beniczky et al., 2021).
  3. Speech neuroprostheses: A 2024 BrainGate report demonstrated 97.5% accuracy at approximately 32 words per minute in one participant with ALS. The result establishes feasibility in that participant, not general clinical effectiveness (Card et al., 2024).

What Does Not Work:

  1. Risk scores used as substitutes for assessment: A low-risk output cannot exclude acute suicidal thoughts or stressors that are absent from the training data.
  2. Unvalidated social-media intervention claims: Algorithmic flagging is not evidence that a program prevents suicide.
  3. Autonomous psychiatric diagnosis from digital phenotyping: Research associations do not establish a clinically valid standalone diagnosis.
  4. Unvalidated Alzheimer screening claims: Imaging or retinal associations require independent validation and a defined clinical pathway before use.

Critical Insights:

Time-sensitive imaging AI can improve workflow: Viz.ai-specific evidence supports shorter treatment intervals, while patient-outcome evidence is mixed. ICH triage evidence supports worklist prioritization in defined studies.

Neurology AI excels at “what” not “why”: Algorithms detect that a lesion exists but struggle with what diagnosis it represents (MS plaque vs. small vessel ischemia vs. infectious lesion)

Psychiatric AI requires unusually strict evidence boundaries: Suicide-risk and digital-phenotyping models can create false reassurance, false positives, privacy concerns, and unclear intervention pathways. Their discrimination metrics do not by themselves establish clinical benefit.

ALS AI is advancing across diagnosis, monitoring, and communication: Research biomarker classifiers, wearable digital biomarkers, and brain-computer interfaces represent genuine progress in defined cohorts. No ALS-specific AI device authorization is identified in the primary FDA records cited here. Company-reported Breakthrough Device Designations are not authorization or proof of benefit.

Prognostic algorithms for neurodegenerative disease remain research tools: Predicting ALS progression or Alzheimer’s conversion does not yet change management, though digital biomarkers are proving more sensitive than clinical scales for tracking ALS decline

Equity gaps in neurology AI are under-studied: Most stroke and MS algorithms trained on predominantly white populations; performance in other demographics uncertain

Integration challenges remain: Even accurate LVO detection fails if it triggers pages to the wrong service or if interventionalists do not trust the alerts

Clinical Bottom Line:

Neurology AI has its clearest clinical evidence in acute time-sensitive imaging and workflow support, where defined product-workflow combinations can shorten treatment intervals. Faster workflow is clinically relevant, but it must not be reported as a mortality or functional-outcome benefit unless the study measured and demonstrated that endpoint.

ALS/MND represents a second area of genuine progress: Brain-computer interfaces have restored communication in individual investigational studies, biomarker classifiers are approaching diagnostic research utility, and wearable sensors can detect dimensions of progression that complement the ALSFRS-R. The terminated VRG50635 Phase 1b registry illustrates why platform selection does not substitute for clinical evidence. Adaptive platform trials, digital biomarkers, and model-generated external comparators are changing trial methods, but none independently proves therapeutic efficacy.

But maintain profound skepticism about: - Psychiatric AI (suicide prediction, depression diagnosis): Ethically fraught, repeatedly failed, risks false reassurance - Neurodegenerative prognostication as individual clinical tools: Population-level models do not yield reliable individual predictions - Autonomous neurological diagnosis: Clinical context and examination findings remain essential

Match the study design to the claim. Technical validation supports accuracy claims, prospective workflow studies support operational claims, and comparative studies with prespecified patient-relevant endpoints are needed for outcome-benefit claims. Ask vendors to provide the study population, comparator, threshold, workflow, and absolute patient-relevant effect.

Medico-Legal Considerations:

  • Stroke AI accountability: A missed or delayed response requires review of clinical duties, alert routing, institutional policy, product labeling, staffing, contracts, and applicable law. No universal rule assigns responsibility to one actor.
  • False negatives in ICH detection: A negative triage output does not replace the interpretation required by the device labeling and local clinical workflow.
  • Psychiatric AI and liability: Responsibility is fact- and jurisdiction-specific. A low-risk output does not eliminate the need for assessment when clinical information raises concern.
  • Consent for research or investigational tools: Consent, authorization, and disclosure requirements depend on the context of use, research status, applicable law, institutional policy, and the information collected.
  • Privacy in psychiatric AI: Digital phenotyping (smartphone monitoring for mood detection) raises HIPAA and consent questions
  • Duty to warn vs. algorithmic false positives: High false positive rates in seizure prediction may trigger unnecessary interventions or driving restrictions
  • Documentation requirements: Document all AI-assisted neurological decisions, especially when you disagree with algorithm recommendation

Essential Reading:

  • Karamchandani RR et al. (2023). “Automated detection of intracranial large vessel occlusions using Viz.ai software: Experience in a large, integrated stroke network.” Brain and Behavior. Large integrated-network Viz.ai LVO validation study

  • Chilamkurthy S et al. (2018). “Deep learning algorithms for detection of critical findings in head CT scans: a retrospective study.” The Lancet 392:2388-2396. [Qure.ai ICH detection validation study]

  • Arbabshirani MR et al. (2018). “Advanced machine learning in action: identification of intracranial hemorrhage on computed tomography scans of the head with clinical workflow integration.” npj Digital Medicine 1:9. [Aidoc ICH algorithm]

  • Walsh CG et al. (2018). “Predicting suicide attempts in adolescents with longitudinal clinical data and machine learning.” Journal of Child Psychology and Psychiatry 59:1261-1270. Suicide prediction algorithm study

  • Commowick O et al. (2018). “Objective evaluation of multiple sclerosis lesion segmentation using a data management and processing infrastructure.” Scientific Reports 8:13650. MS lesion detection AI challenges

  • Chia R et al. (2025). “Plasma proteomics identifies ALS biomarkers.” Nature Medicine 31:3440-3450. ALS blood biomarker AI with AUC 98.3%

  • Card NS et al. (2024). “An accurate and rapidly calibrating speech neuroprosthesis.” NEJM. BrainGate BCI achieving 97.5% accuracy in ALS patient

  • Gupta A et al. (2023). “At-home wearables and machine learning sensitively capture disease progression in amyotrophic lateral sclerosis.” Nature Communications 14:5080. Wearable digital biomarkers outperforming ALSFRS-R

  • Reimer RJ et al. (2026). “Machine learning drug repurposing for ALS.” Lancet Digital Health. AI-driven drug repurposing from 11,000+ veterans’ records


Introduction: The Promise and Peril of Neurological AI

The human brain is the most complex structure in the known universe. Three pounds of tissue containing 86 billion neurons, each forming thousands of synaptic connections, generating thought, emotion, movement, memory, and consciousness itself.

Neurology and psychiatry attempt to diagnose and treat disorders of this incomprehensibly complex organ using a combination of: - Imaging (CT, MRI, PET) that shows structure but not function - Physiologic testing (EEG, EMG, evoked potentials) that measures electrical activity - Clinical examination (mental status, cranial nerves, motor, sensory, coordination, gait) - Patient-reported symptoms (headache character, mood symptoms, cognitive complaints)

This combination of objective data and subjective assessment creates both remarkable opportunities and profound limitations for AI in neurology.

Where neurology AI has its clearest evidence: Acute stroke imaging and communication workflows, where a defined product can prioritize a suspected large-vessel occlusion and shorten selected treatment intervals.

Where neurology AI struggles: Diagnosing Parkinson’s disease, where the exam requires detecting subtle bradykinesia, rigidity, and postural instability that vary moment-to-moment.

Where psychiatric AI requires more evidence: Suicide-risk prediction, where discrimination can appear strong while positive predictive value, calibration, transportability, and the effect of the response pathway remain inadequate or uncertain.

Neurological AI requires the same claim-evidence discipline as every specialty, with particular attention to the difference between pattern recognition, diagnostic interpretation, prognostication, and patient outcomes.


Part 1: Stroke AI and Time-Sensitive Workflow

Large Vessel Occlusion Detection: When Minutes Matter

Acute ischemic stroke from large-vessel occlusion (LVO) is a neurological emergency. A widely cited model estimated that a typical untreated large-vessel ischemic stroke loses approximately 1.9 million neurons per minute, while acknowledging variation in infarct volume and evolution (Saver, 2006). Mechanical thrombectomy can restore blood flow and reduce disability for eligible patients, but benefit remains time-dependent.

The problem AI solves: Many stroke patients present to community hospitals without thrombectomy capability. Identifying LVO patients who need immediate transfer is time-critical but requires expert interpretation of CT angiography that may not be immediately available at 2 AM.

The AI solution: Automated LVO detection analyzes CT angiography in seconds, flags potential LVO cases, and directly alerts the interventional neuroradiology team at receiving hospitals, bypassing the usual radiology reading queue.

Viz.ai LVO Detection

FDA clearance: 2018 (first stroke AI cleared by FDA)

Performance boundary: Sensitivity, specificity, analysis time, and predictive values depend on the exact software version, occlusion definition, vessel territory, image quality, disease prevalence, threshold, and evaluation design. No class-wide range should be used for clinical procurement.

Clinical validation: Karamchandani et al. (2023) conducted a large real-world evaluation of Viz.ai’s LVO detection in an integrated hub-and-spoke stroke network (Karamchandani et al., 2023):

  • Study design: Multicenter retrospective analysis across integrated stroke network
  • Diagnostic performance: High specificity and moderately high sensitivity for ICA and proximal MCA occlusions on CTA
  • Clinical impact: Streamlined code stroke workflows by enabling direct notification to neurointerventionalists

How it works: 1. Patient with suspected stroke gets CT angiography at community hospital 2. Viz.ai analyzes imaging automatically (integrated with PACS) 3. If LVO detected, Viz.ai sends HIPAA-compliant alerts to: - Neurointerventionalist’s smartphone (with images) - Neurologist on call - Stroke coordinator - Receiving hospital ED 4. Transfer arranged while local team evaluates patient 5. Patient arrives at comprehensive stroke center with cath lab team already mobilized

Current deployment: Commercial scale and hospital counts are vendor-maintained figures. They do not establish diagnostic performance or clinical utility and should be verified from current product documentation if needed for procurement.

RapidAI evidence boundary: RapidAI products have separate regulatory records and observational evidence. Viz.ai citations below should not be used to support RapidAI or Brainomix. Claims about vessel-territory coverage require the exact product labeling.

Best available evidence: Sarhan et al. meta-analysis (Translational Stroke Research, 2025)

A systematic review and meta-analysis of 12 studies and 15,595 patients evaluated Viz.ai LVO software (Sarhan et al., 2025): - Workflow metrics significantly improved: door-to-intervention notification time and door-to-arterial puncture both reduced - Mortality did not differ significantly (RR 0.72, 95% CI 0.37-1.41, P = 0.34) - Hospital length of stay unchanged

This is the strongest current evidence summary: workflow improvement demonstrated at scale, but mortality benefit has not been shown in pooled analysis. This does not mean LVO AI is ineffective for outcomes, but the evidence is observational and the mortality signal is not yet confirmed.

Viz.ai randomized evidence: A prospective stepped-wedge cluster-randomized trial at four stroke centers included 243 patients treated with endovascular thrombectomy. AI activation reduced adjusted door-to-groin time by 11.2 minutes and CT-to-treatment time by 9.8 minutes, without a significant improvement in 90-day functional independence (Martinez-Gutierrez et al., 2023). This randomized result is Viz.ai-specific and must not be transferred to RapidAI or Brainomix.

GOLDEN BRIDGE II: stroke CDSS with patient-event outcomes

A 2026 BMJ cluster randomized trial tested a stroke clinical decision support system across 77 hospitals in China. The intervention combined AI-assisted imaging analysis, stroke-cause classification, and evidence-based treatment recommendations for 21,603 patients with acute ischemic stroke. New vascular events at 3 months were lower in the intervention group (2.9% vs. 3.9%; adjusted HR 0.74), evidence-based care quality improved, and 12-month vascular events were also lower (4.0% vs. 5.5%; adjusted HR 0.73), without significant bleeding differences (Zhang et al., 2026).

Clinical interpretation: GOLDEN BRIDGE II is not an LVO-alert study. It evaluates an integrated stroke-care CDSS that includes imaging, etiology, and guideline recommendations. The result suggests AI is most likely to improve stroke outcomes when embedded inside a care pathway with treatment recommendations and quality metrics, not when sold as an isolated imaging detector.

Why this works: - Clear imaging signature: LVO on CTA is visible, well-defined pattern - Time-sensitive: Every minute saved prevents brain damage - Actionable: Positive detection triggers specific intervention (thrombectomy) - Integrated workflow: Alerts go directly to decision-makers - Evidence matched to the claim: Viz.ai has randomized workflow evidence; the trial did not establish a significant functional-outcome benefit

Intracranial Hemorrhage Detection: Prioritizing the Radiology Worklist

CT scans of the head are one of the most common imaging studies ordered in emergency departments. Most are normal. But the 5-10% showing intracranial hemorrhage require urgent neurosurgical consultation.

The problem: Radiology worklists are first-come-first-served. A critical ICH scan may sit in the queue behind 20 normal CTs while the radiologist reads chronologically. By the time the radiologist sees the bleed, the patient has been herniating for 30 minutes.

The AI solution: Automated ICH detection algorithms analyze every head CT in seconds, flag those with suspected hemorrhage, and move them to the top of the worklist, ensuring critical cases are read first.

Aidoc ICH Detection

FDA clearance: 2018

Performance boundary: ICH triage performance is study- and product-specific. The Arbabshirani workflow study below reported 73% sensitivity, 80% specificity, and AUC 0.85. Results from another product, dataset, or later version must be cited separately.

Types of ICH detected: - Epidural hematoma - Subdural hematoma - Subarachnoid hemorrhage - Intraparenchymal hemorrhage - Intraventricular hemorrhage

Clinical benefit: Arbabshirani et al. evaluated a deep-learning ICH detector and a simulated worklist-reprioritization workflow for routine outpatient head CT (Arbabshirani et al., 2018):

  • Time to radiologist interpretation: AI-reprioritized studies reached interpretation in ~19 minutes vs. >8 hours for standard “routine” studies
  • Mechanism: Algorithm automatically escalated routine head CTs to “stat” priority when ICH detected
  • Key limitation: Radiologists must still review all scans; AI serves as triage tool, not replacement

How it works: 1. Head CT ordered in ED and sent to PACS 2. Aidoc analyzes every slice in real-time 3. If ICH suspected: - Case moved to top of radiology worklist - Alert sent to radiologist - Alert sent to ED physician - Optional alert to neurosurgery (if institution configures it) 4. Radiologist reviews flagged case immediately 5. Radiologist confirms or rejects AI finding

Current deployment: Vendor-reported deployment scale is not evidence of patient benefit. Procurement should focus on the exact intended use, local case mix, alert burden, version control, and monitored workflow.

Why this works: - Clear imaging pattern: Acute blood on CT is hyperdense, easy to detect algorithmically - Triage, not diagnosis: Algorithm does not replace radiologist. It prioritizes worklist - Triage, not exclusion: A negative output cannot rule out hemorrhage, and the radiologist still reviews every examination - Workflow integration: Smooth PACS integration, no extra clicks required - Clinical validation: Multiple studies show earlier detection

Implementation Reality: When Accurate Stroke AI Still Fails

Even a detector with strong technical performance may fail to improve outcomes if the surrounding workflow does not shorten recognition, transfer, and treatment.

The implementation failures:

  1. Alert fatigue: If LVO alerts go to a generic pager that also receives lab results, medication alerts, and bed assignments, interventionalists may ignore them

  2. Workflow fragmentation: Algorithm detects LVO at community hospital, but transfer process still requires:

    • ED physician calling transfer center
    • Transfer center calling receiving hospital
    • Receiving hospital calling interventionalist
    • Interventionalist reviewing images
    • Each step adds 10-15 minutes
  3. Lack of trust: Interventionalists who do not trust algorithm may wait for official radiology read, defeating purpose of AI early detection

  4. False positives in wrong population: Algorithms trained on stroke patients perform poorly when applied to all head CTs (many false positives from artifacts, old infarcts, masses)

  5. Equity gaps unknown: Most LVO algorithms trained predominantly on white populations; performance in Black and Hispanic patients not well-studied (though preliminary data suggests equivalent performance)

Fictional implementation example: An academic center routes alerts to a work phone that is not continuously staffed. Nights and weekends produce delayed review, so treatment intervals do not improve despite accurate detection. This is a hypothetical workflow failure, not a report of a named institution.

The lesson: Technical accuracy does not produce clinical value unless the alert reaches the right team and prompts timely, appropriate action.


Part 2: Neuroimaging AI Beyond Stroke

Brain Tumor Segmentation: Where AI Actually Saves Time

Radiation oncology treatment planning for brain tumors requires meticulous delineation of: - Gross tumor volume (GTV) - Clinical target volume (CTV) - Organs at risk (optic nerves, brainstem, hippocampi)

Manual segmentation takes 6-8 hours per patient. Automated AI segmentation reduces this to 30-60 minutes with radiologist review.

FDA-cleared systems: - BrainLab Elements: Automated glioblastoma segmentation - Neosoma Brain Mets: Brain metastases detection and volumetry - DeepMind (research): World-class segmentation published in Nature but not commercially deployed

Performance: - Dice coefficient: 0.85-0.90 (measures overlap between AI and expert segmentation) - Time savings: 80-90% reduction in segmentation time - Consistency: Reduces inter-observer variability in target volume delineation

Clinical use: Radiation oncologists review and edit AI-generated contours rather than manually tracing every structure. This is a genuine time-saver that improves workflow without compromising quality.

Molecular prediction from imaging: Neuro-oncology AI is also moving from segmentation toward biomarker prediction. Takahashi et al. compared two MRI-based AI models with 18 physicians for IDH mutation prediction in glioma. GliomaVista-IDH achieved AUC 0.97 on the Brain Tumor Segmentation Challenge dataset but declined to AUC 0.82 on external Japanese validation, while high-performing physicians reached AUC 0.88 with better calibration (Takahashi et al., 2026). The clinical lesson is not that AI replaces molecular testing; it is that external validation and calibration matter before imaging AI is used for treatment planning.

Limitations: - Complex cases require extensive editing: Tumors with necrosis, hemorrhage, or prior surgery confuse algorithms - Post-treatment changes: Pseudoprogression vs. true progression remains challenging for AI - Rare tumor types: Algorithms trained on glioblastoma perform poorly on meningiomas, metastases, lymphomas

General Brain MRI Interpretation: Vision-Language Models

Beyond specialized tools for stroke or tumor segmentation, a new class of AI aims to interpret brain MRIs holistically, covering the full spectrum of neurological diagnoses.

Prima (University of Michigan, 2026): Researchers at U-M developed Prima, a vision-language AI model that analyzes brain MRI scans and delivers diagnostic assessments in seconds. Published in Nature Biomedical Engineering (February 2026), the system was trained on over 220,000 MRI studies and 5.6 million sequences from U-M Health’s radiology archive, then evaluated on over 30,000 MRI studies across 52 radiologic diagnoses spanning major neurological disorders (Lyu et al., 2026).

Key performance:

  • Diagnostic accuracy: Mean AUC of 92.0% across 52 neurological conditions, with accuracy reaching 97.5% for specific diagnosis categories
  • Triage capability: In the study setting, the system demonstrated the ability to determine case urgency and route alerts to the appropriate subspecialist
  • Speed: Delivers results in seconds after scan completion, compared to hours or days for conventional radiologist interpretation queues

Unlike single-purpose algorithms (Viz.ai for LVO, Aidoc for ICH), Prima attempts comprehensive brain MRI interpretation. As a vision-language model, it processes imaging data and generates text-based diagnostic assessments rather than simple binary classifications.

Limitations:

  • Research stage only: Not FDA-cleared, not commercially deployed. The gap between academic publication and clinical deployment is measured in years, not months
  • Single-institution validation: Performance at U-M may not generalize across different scanners, protocols, and patient populations
  • 52 diagnoses still leaves gaps: Rare conditions, atypical presentations, and incidental findings outside the training set will be missed
  • Vision-language model risks: Text generation introduces hallucination risk; a confident but wrong narrative diagnosis could mislead more than a simple false negative flag

Prima represents the direction brain imaging AI is heading: from narrow, single-finding detectors toward comprehensive interpretation assistants. The standard applies: demand prospective multi-site validation and real-world outcome data before trusting any system with comprehensive neurological diagnosis.

Pre-Surgical Brain Mapping: Resting-State fMRI

Neurosurgery for brain tumors and epilepsy requires identifying critical brain networks before resection. Traditional task-based functional MRI requires patients to perform specific tasks while in the scanner. Task compliance can be limited by age, sedation, language, aphasia, cognitive impairment, anxiety, or movement. The proportion of unusable studies varies by population and protocol.

The problem task-based fMRI solves poorly: - Pediatric patients cannot reliably perform motor or language tasks during 45-60 minute scans - Patients requiring sedation for claustrophobia or movement disorders cannot participate in tasks - Non-English speakers cannot perform language mapping protocols designed for English - Patients with aphasia or severe cognitive deficits cannot follow task instructions - Task-based fMRI can produce unusable or ambiguous maps when task performance or acquisition quality is inadequate

The AI solution: Resting-state fMRI analyzes spontaneous brain activity patterns while patients lie still, without performing tasks. AI algorithms identify correlated activity patterns across brain regions that correspond to functional networks (motor network, language network, visual network, default mode network).

FDA-cleared system: - Cirrus Resting State fMRI Software: FDA cleared K251009 as adjunctive software that processes BOLD fMRI data and produces task-analogous sensorimotor, language, and vision resting-state maps (FDA, K251009).

Regulatory boundary: - FDA labeling states that Cirrus generates motor, language, and vision correlation maps and a quality report. - The labeling warns that maps can vary on repeat acquisition even when no functional change is expected. - These systems are adjunctive and are not intended to replace direct functional mapping procedures. - The authorization does not establish the uncited 87% versus 67% usability comparison or superior surgical outcomes.

Clinical workflow: 1. Patient undergoes 12-20 minute resting-state fMRI acquisition (eyes closed, no task performance required) 2. Cirrus software processes blood oxygen level dependent (BOLD) fMRI data 3. Algorithm generates network maps showing motor cortex, language areas, visual cortex locations 4. Quality report accompanies maps, indicating confidence levels 5. Neurosurgeon reviews maps alongside structural imaging for surgical planning 6. Maps guide tumor or seizure focus resection, identifying regions to preserve

Patient populations who benefit most: - Pediatric patients: Children as young as 5-6 can undergo resting-state mapping (just need to lie still) - Sedated patients: Those requiring anesthesia for claustrophobia or movement - Non-English speakers: Language network mapping works across all languages - Cognitively impaired: Patients with aphasia, dementia, or intellectual disability - High movement risk: Parkinson’s patients, essential tremor, severe anxiety

Development and validation: The technology builds on decades of resting-state functional-connectivity research. Historical development and licensing claims are not substitutes for the FDA record or comparative clinical evidence.

Current deployment: Early adoption phase. Sora Neuroscience has non-exclusive distribution partnership with Prism Clinical Imaging for integration into Prism’s brain mapping platform. Primarily deployed at academic medical centers with neurosurgical programs.

Limitations: - Requires patient cooperation for stillness: Motion artifacts still degrade quality, though less problematic than task-based fMRI - Resolution limitations: Identifies general network locations but may lack precision for small, critical areas near resection margins - Validation against intraoperative mapping: Resting-state maps should be confirmed with awake craniotomy mapping when feasible for eloquent cortex near resection sites - Not a replacement for clinical judgment: Maps inform but do not dictate surgical decisions; surgeon experience with functional anatomy remains essential

Why this works: Functional networks show correlated spontaneous activity even at rest. Motor, language, visual, and other networks exhibit intrinsic functional connectivity. Computational methods can estimate these networks without task performance, but reliability depends on acquisition, motion, preprocessing, quality control, and the individual patient.

Clinical bottom line: Resting-state fMRI can extend mapping to patients who cannot complete task paradigms. Its clinical value is adjunctive: map quality and surgical relevance require expert review, and direct functional mapping remains the reference procedure when appropriate.

Multiple Sclerosis Lesion Detection: Promise and Pitfalls

MS diagnosis and monitoring rely on detecting white matter lesions on brain MRI. Lesion burden and new lesion formation guide treatment decisions.

AI applications: - Automated lesion counting - Lesion volume quantification - Detection of new/enlarging lesions compared to prior scans - Prediction of disease progression

Performance: Variable and scanner-dependent. Commowick et al. (2018) evaluated 13 MS lesion segmentation algorithms across a 53-case database from four imaging centers (Commowick et al., 2018): - Sensitivity: Varied substantially depending on algorithm and scanner - False positives: Many algorithms flagged normal periventricular white matter as lesions - Scanner dependence: Algorithms trained on Siemens MRI performed poorly on GE scanners

Why MS lesion AI is harder than stroke: - Lesion heterogeneity: MS lesions vary in size (2mm to 3cm), location, and appearance - Look-alikes: Small vessel ischemic disease, migraine, normal aging all produce white matter hyperintensities - Scanner variability: MRI protocols differ across institutions; algorithms do not generalize well

Current clinical use: Research settings and pharmaceutical clinical trials (where standardized protocols and centralized reading reduce scanner variability). Not yet ready for routine clinical care.

MindGlide: MS monitoring on routine clinical scans (Nature Communications, April 2025)

MindGlide (Goebl et al., Nature Communications, 2025) is a deep learning model that extracts brain region volumes and white matter lesion measures from any single MRI contrast, including non-standard T2-weighted scans without FLAIR (Goebl et al., 2025). Validated on 14,952 images from 1,001 MS patients across two clinical trials and a routine-care dataset; outperformed SAMSEG by 60% and WMH-SynthSeg by 20% for lesion localization. The key advance: it works on the heterogeneous scans actually obtained in clinical practice, not only research-quality MRI.

A companion Nature Medicine study (Ganjgahi et al., 2025) reclassified MS disease states using probabilistic ML across approximately 8,000 patients and 118,000 visits from 10 institutions, defining four disease dimensions (physical disability, brain damage, relapse, subclinical activity) that challenge the traditional RRMS/SPMS/PPMS taxonomy (Ganjgahi et al., 2025). The implication for treatment selection: staging patients on a severity continuum may predict treatment response better than current phenotypic subtypes, but this requires prospective validation before changing practice.

Alzheimer’s Disease Neuroimaging: Prediction Without Treatment

AI can predict conversion from mild cognitive impairment (MCI) to Alzheimer’s dementia with AUC 0.80-0.85 using: - Hippocampal volume measurement - Entorhinal cortex thickness - Amyloid PET standardized uptake value ratios - FDG-PET glucose metabolism patterns

The clinical problem: These predictions do not change management. We lack disease-modifying treatments for Alzheimer’s. Knowing that a patient with MCI will progress to dementia in 3 years does not help them. It just causes anxiety.

Ethical concerns: - Prognostic disclosure: Should we tell patients they’ll develop dementia when we cannot prevent it? - Insurance discrimination: Will Alzheimer’s risk predictions affect life insurance, long-term care insurance? - Clinical trial recruitment: This is the main current use, enriching trials with high-risk patients

Until recently, blood-based Alzheimer’s diagnostics did not exist. This changed in 2025.

Lumipulse Alzheimer blood test: First FDA-cleared blood diagnostic (May 16, 2025)

FDA cleared the Lumipulse G pTau217/β-Amyloid 1-42 Plasma Ratio under K242706 on May 16, 2025. It is intended to aid identification of amyloid pathology in adults aged 50 years or older who present in specialized care with signs and symptoms of cognitive decline. It is not a screening or stand-alone diagnostic test, and a positive result does not establish Alzheimer disease (FDA, K242706).

In February 2026, FDA posted Class II recalls involving specified assay components and lots after reports of falsely elevated ratios and reduced specificity. Affected users were instructed to discontinue the listed lots and review prior results as appropriate (FDA recall Z-1306-2026). Regulatory authorization is not the end of evidence review; lot-specific corrections and recalls can change the interpretation of previously generated results.

A multicenter validation (Palmqvist, Hansson et al., Nature Medicine, April 2025) across 1,767 patients at 5 European centers reported AUC 0.93-0.96 for amyloid status. Accuracy was 89-91% in secondary care with a single cutoff; a two-cutoff approach raised accuracy to 92-94% while leaving 12-17% of results indeterminate (Palmqvist et al., 2025).

Clinical implication: Blood-based p-tau217 testing now provides an alternative to PET and lumbar puncture for amyloid detection. For neurologists considering anti-amyloid therapy referrals, this test substantially reduces the diagnostic burden. It is not an AI algorithm per se, but it changes the context for AI neuroimaging tools in Alzheimer’s: when a low-cost blood test can screen for amyloid, the bar for AI-based imaging analysis rises.

The broader caveat remains: disease-modifying therapies are now available (lecanemab, donanemab) but benefit primarily early-stage patients. Predicting conversion without available treatment still creates more anxiety than benefit. These tools should be used when the result changes management.

Emerging approach: Sleep-based prediction. Foundation models trained on polysomnography data can predict dementia (C-Index 0.85) and Parkinson’s disease (C-Index 0.89) from a single night of sleep (Thapa et al., 2026). Sleep disturbances often precede clinical diagnosis by years, and REM sleep behavior disorder is a recognized prodromal marker for Parkinson’s. Whether sleep-based prediction offers advantages over neuroimaging (lower cost, wider availability) remains to be determined in prospective studies. See Emerging Technologies chapter for details.


Part 3: Seizure Detection and Epilepsy AI

Automated Seizure Detection from EEG

Long-term EEG monitoring generates 24-72 hours of continuous data per patient. Neurologists review this data looking for seizures, interictal epileptiform discharges, and background abnormalities.

The time burden: - 1 hour of EEG recording = ~10 minutes of expert review - 72-hour EEG = 12 hours of neurologist time - Most EEG shows no seizures (neurologists search for rare events in vast normal data)

AI solution: Automated seizure detection algorithms analyze EEG continuously, flag suspected seizures, and present condensed summaries to neurologists for review.

Persyst Seizure Detection: FDA-cleared automated EEG analysis system.

Performance: - Sensitivity for generalized tonic-clonic seizures: 92% - Sensitivity for complex partial seizures: 76% - False positive rate: 0.5-1.0 false detections per hour - Time savings: Reduces neurologist review time by 60-70%

How neurologists use it: 1. Review AI-flagged events first (likely seizures) 2. Quickly scroll through unflagged periods looking for missed events 3. Measure total review time, false detections, missed events, and expert correction burden rather than assuming a fixed reduction

Limitations: - Misses subtle seizures: Focal seizures without clear rhythmic activity often missed - False positives from artifacts: Chewing, movement, electrode problems cause false alarms - ICU EEG challenging: Critically ill patients on sedation with frequent interventions generate artifacts

Persyst and similar legally marketed review systems can help prioritize or summarize long recordings, but the magnitude of time savings and false-detection burden must be measured in the intended population and local review workflow. Authorization of an EEG analysis function does not authorize autonomous interpretation.

2025 EEG-device expansion: separate acquisition, detection, and prediction

Ceribell has multiple device records covering different acquisition and analysis functions. They should not be collapsed into a single class-wide indication. For example, K251936 covers a machine-learning delirium-monitoring function using EEG collected by specified cleared hardware. That record does not establish seizure-detection performance across all ages or authorize autonomous diagnosis.

The SAFER-EEG retrospective multicenter study reported associations between point-of-care EEG use and hospital outcomes. Because assignment was not randomized, those associations do not establish that the device caused shorter stays or better functional outcomes.

Epi-Minder Minder System: FDA granted De Novo authorization under DEN240062 on April 17, 2025. The authorized system continuously acquires and exports subscalp EEG for review; the record does not authorize seizure forecasting. Continuous acquisition, automated detection, and prediction are different functions and require separate evidence and regulatory verification.

Wearable Seizure Detectors: High Sensitivity, High False Positive Rate

Smartwatch-based seizure detectors (Empatica Embrace, Nightwatch) use accelerometry and autonomic signals (heart rate, skin conductance) to detect generalized tonic-clonic seizures.

Use case: Preventing sudden unexpected death in epilepsy (SUDEP) by alerting caregivers when seizure occurs, particularly during sleep.

Performance: - Sensitivity for generalized tonic-clonic seizures: 90-95% - False positive rate: 1-5 false alarms per month - Focal seizures: Often undetected (no convulsive movements)

Clinical reality: - High-risk epilepsy patients (frequent convulsive seizures, intellectual disability, living alone) benefit from alerts - Low-risk patients (well-controlled focal epilepsy) find false alarms burdensome - Not a replacement for supervision: Algorithms detect seizures but cannot intervene

High-risk epilepsy patients benefit from wearable seizure detectors, particularly for SUDEP prevention during sleep. But the 1-5 false alarms per month mean this is an adjunct for high-risk patients, not a general screening tool for everyone with epilepsy.

EHR epilepsy phenotyping (Chang 2026):

Wearable and EEG detectors flag events. They do not recover ILAE epilepsy type or seizure type from the note. An 11 August 2026 npj Digital Medicine study compared a fine-tuned BERT model and a reasoning-optimized LLM (DeepSeek-R1) against pairwise agreement among board-certified epileptologists. On coarser tasks both models matched experts: epilepsy type as focal, generalized, or other (Matthews correlation coefficient DeepSeek 0.85, BERT 0.73, human 0.77) and seizure type as convulsive or non-convulsive (DeepSeek 0.74, BERT 0.60, human 0.49). DeepSeek stayed at expert level on more granular tasks while BERT declined. Deploying DeepSeek-R1 on 77,049 notes from 18,566 patients produced phenotypes consistent with diagnostic stabilization, seizure-type co-occurrence, and outcome differences by epilepsy type (Chang et al., 2026). This is single-center retrospective note classification for research-scale phenotyping. The authors say prospective validation is still required before clinical decision support. It does not replace EEG review, wearable alarms, or ILAE classification at the bedside.


Part 4: Neurodegenerative Disease AI

Parkinson’s Disease: Objective Motor Assessment

Parkinson’s disease diagnosis and monitoring rely on subjective clinical assessment of bradykinesia, rigidity, tremor, and gait. Symptom severity fluctuates throughout the day (medication on/off states).

AI applications: - Smartphone tapping tests: Measure finger tapping speed and rhythm - Smartwatch tremor detection: Accelerometry detects tremor frequency and amplitude - Voice analysis: Detect hypophonia and monotone speech - Gait analysis: Computer vision from smartphone video analyzes stride length, arm swing

Performance: Modest. Multiple smartphone-based PD motor assessments show only moderate correlation with MDS-UPDRS motor scores and limited ability to distinguish PD from other movement disorders:

  • Correlation with MDS-UPDRS motor scores: Moderate (r values typically 0.5-0.7 across studies)
  • Distinguishing PD from healthy controls: Reasonable performance in research settings
  • Distinguishing PD from other movement disorders: Poor, limiting diagnostic utility

Why this does not work well yet: - Bradykinesia is nuanced: Requires observing finger tapping, hand movements, leg agility. Smartphones capture only finger tapping - Medication state confounds: Patients tested 1 hour post-dose look different than 4 hours post-dose - Non-motor symptoms ignored: Cognitive impairment, autonomic dysfunction, psychiatric symptoms not measured

Current use: Research setting (clinical trials tracking motor progression). Not diagnostic.

Adaptive Deep Brain Stimulation: Closed-Loop Neuromodulation, Not Established ML

Adaptive DBS (aDBS) represents a fundamental shift from conventional continuous DBS: instead of fixed stimulation parameters, aDBS uses real-time sensing of local field potentials from the implanted electrode to automatically adjust stimulation based on beta-band oscillatory activity correlated with motor symptoms.

FDA status: FDA approved an adaptive DBS programming option through PMA supplement P960009/S478. The system adjusts stimulation using configured neural-signal thresholds. The FDA materials do not establish that the therapy is a machine-learning model, so it should not be presented as proof of AI-guided treatment benefit (FDA SSED, P960009/S478).

ADAPT-PD evidence: FDA’s safety and effectiveness summary describes a prospective multicenter, single-arm, treatment-mode-blind, randomized crossover evaluation. The 45-patient summary below is not a complete description of the pivotal evidence. - Dual-threshold aDBS: 91% achieved good on-time without troublesome dyskinesia - +1.3 hours on-time, -1.6 hours off-time vs. stable continuous DBS - 98% (44/45) chose to remain on aDBS after the evaluation period

This is an important closed-loop neuromodulation development. It should be evaluated as an adaptive device intervention with patient-selection, signal-quality, programming, adverse-event, and comparative-effectiveness questions, not promoted as a machine-learning milestone.

Limitations: The study design, comparison periods, blinding, selected signal thresholds, and patient eligibility determine the defensible claim. Longer-term comparative effectiveness and postmarket performance remain important.

AI in Motor Neuron Disease: From Diagnosis to Communication

Amyotrophic lateral sclerosis (ALS) causes progressive motor neuron degeneration with highly variable progression rates. Median survival is 3-5 years from symptom onset, but some patients survive 10+ years while others decline within 12 months. Diagnostic delay averages 10-14 months from first symptoms, during which patients often see multiple specialists before receiving a diagnosis. AI is now making measurable contributions across the entire ALS care pathway: earlier diagnosis through blood biomarkers and electrophysiology, more sensitive disease monitoring through wearable sensors and speech analysis, and restored communication through brain-computer interfaces.

ALS Diagnosis: Blood Biomarkers and Electrophysiology

ALS diagnosis remains clinical, relying on the revised El Escorial criteria and Awaji-Shima criteria, which require evidence of both upper and lower motor neuron degeneration. No single biomarker confirms the diagnosis. AI is changing this.

Plasma proteomics (NIH, 2025): Chia et al. identified 33 proteins differentially abundant in ALS patients (n=183) versus controls (n=309) using plasma proteomics. A machine learning model trained on these proteins achieved AUC of 98.3% for ALS diagnosis, the highest diagnostic accuracy reported for any ALS blood test (Chia et al., Nature Medicine, 2025). The study was led by Bryan Traynor’s group at NIH/NIA. External validation in independent cohorts is needed before clinical deployment.

Whole blood gene expression signatures (University of Michigan, 2025): Zhao et al. applied XGBoost classifiers to RNA sequencing data from ALS participants (n=422) versus controls (n=272). Panels of 27-46 genes predicted ALS with AUC 0.894 in an independent external validation cohort (AUC 0.91 in internal testing). The same gene expression data improved survival stratification when combined with clinical variables (Zhao et al., Nature Communications, 2025).

F-wave AI models (Mayo Clinic, 2025): Martinez-Thompson et al. analyzed F-wave responses from 46,802 patients using deep learning. The model achieved 90% recall, 87% precision, and 88% accuracy for ALS prediction from routine nerve conduction study data. Beyond diagnosis, the model also predicted ALS survival, bulbar symptom progression, and differential diagnosis (ALS versus inclusion body myositis versus radiculopathy versus peripheral neuropathy versus controls) (Martinez-Thompson et al., Brain, 2025). This is significant because it leverages data already collected during standard EMG/NCS evaluation, requiring no additional testing.

Lipidomics for PLS differentiation: Primary lateral sclerosis (PLS), a rarer upper motor neuron-predominant variant, is difficult to distinguish from ALS early in the disease course. Lee et al. trained an elastic net algorithm on plasma lipidomics (392 lipid species) from 40 ALS, 28 PLS, and 28 healthy controls, achieving 100% classification accuracy for PLS versus ALS (Lee et al., Muscle & Nerve, 2023). Small sample size limits generalizability, but the proof-of-concept is striking.

NLP for diagnostic delay: Segura et al. applied named entity recognition and clinical NLP to unstructured clinical notes from 250 ALS patients across 5 Spanish hospitals. The system extracted symptom timelines revealing a median diagnostic delay of 11 months, with only 38.8% of patients seeing a neurologist before diagnosis. The most common pre-diagnosis symptoms were weakness (38%), dyspnea (21.6%), and dysarthria (15.6%) (Segura et al., Scientific Reports, 2023). NLP-extracted diagnostic pathways could identify systemic bottlenecks and reduce time to diagnosis.

Clinical bottom line on ALS diagnostics: Blood-based AI diagnostics for ALS have reached proof-of-concept with AUCs above 0.89 in multiple independent studies. None are clinically deployed. The F-wave AI model from Mayo Clinic is closest to clinical utility because it uses data already collected during routine evaluation. All require prospective multi-site validation before changing practice.

Randomized evidence for AI-assisted electrodiagnostic reporting: Gorenshtein et al. randomized 200 patients undergoing electrodiagnostic evaluation to physician-only or physician-plus-AI report interpretation. The integrated approach did not significantly improve the study’s report-quality score. Physicians rated efficiency (2.0/5), ease of use (1.7/5), and workload reduction (1.7/5) poorly, indicating workflow and usability problems (Gorenshtein et al., 2025). The physicians rated these dimensions poorly in the AI arm; the study did not establish a comparative worsening from baseline on those ratings.

ALS Prognosis and Disease Monitoring

Progression prediction models: Multiple machine learning approaches now predict ALSFRS-R functional decline. The PRO-ACT database (5,030 patients from multiple clinical trials) has become the standard training dataset. Abdul Jabbar et al. (2024) applied XGBoost and Bayesian LSTM models, identifying days since disease onset, prior ALSFRS-R score, and forced vital capacity as the top three predictors (Abdul Jabbar et al., Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration, 2024). A multinational model developed across 14 European ALS centers (largest cohort n=1,936) achieved a concordance statistic of 0.78 (95% CI 0.77-0.80) in external validation (Westeneng et al., Lancet Neurology, 2018).

These models serve clinical trial enrichment (selecting rapid progressors to detect treatment effects faster) rather than individual prognostication. Population-level AUC of 0.78 still yields wide confidence intervals for individual patients.

Wearable sensors outperform clinical scales: The ALSFRS-R, the standard clinical outcome measure for ALS trials, requires in-clinic assessment and has limited sensitivity to early change. Wearable AI is proving more sensitive.

Gupta et al. (2023) equipped 376 ALS patients with limb-worn accelerometers at home. Machine learning automatically detected and characterized submovements from accelerometer data. Wearable-derived scores progressed faster than ALSFRS-R (-0.86 versus -0.73 SD/year), suggesting digital biomarkers could detect treatment effects earlier in clinical trials (Gupta et al., Nature Communications, 2023).

Van Unnik et al. (2024) demonstrated that a simple hip-worn accelerometer (97 patients, 27,701 total hours of recording) produced a Vertical Movement Index strongly associated with mortality risk and mobility status, capturing functional decline that ALSFRS-R missed between clinic visits (van Unnik et al., eBioMedicine, 2024).

Speech biomarkers for remote monitoring: Neumann, Kothare, and Ramanarayanan (2024) collected remote audio/video recordings from 278 ALS participants and extracted acoustic, orofacial, and linguistic features. Speech timing alignment and word count were the most responsive measures for detecting change in both bulbar-onset (n=36) and limb-onset (n=107) patients, detecting decline even when ALSFRS-R showed no change (Neumann et al., Computers in Biology and Medicine, 2024).

The same group demonstrated that speech acoustics can estimate forced vital capacity (FVC), a critical respiratory measure, with correlation r=0.80 and ICC of 0.92-0.94, enabling non-invasive remote respiratory monitoring (Stegmann et al., Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration, 2021).

Why digital biomarkers matter for ALS: ALS clinical trials have historically failed partly because ALSFRS-R is insensitive to early change, requiring large sample sizes and long follow-up periods. Digital biomarkers that detect progression faster could make trials shorter, smaller, and more likely to identify effective treatments. The failed Verge Genomics VRG50635 trial (Phase 1b, terminated December 2025) demonstrated that digital endpoints could detect statistically significant disease progression in just 8 weeks of pre-treatment monitoring, a finding with implications for trial design even though the drug itself failed.

Brain-Computer Interfaces: Restoring Communication in ALS

For ALS patients who lose the ability to speak, move, and eventually control eye movements, brain-computer interfaces represent the most significant AI application in neurology. The field has advanced from laboratory demonstrations to sustained, real-world use.

BrainGate speech neuroprosthesis (NEJM, 2024): Card et al. implanted 4 microelectrode arrays in the left precentral gyrus of Casey Harrell, a 45-year-old man with ALS. On day 1, the system achieved 99.6% accuracy with a 50-word vocabulary. By day 2, accuracy reached 90.2% with a 125,000-word vocabulary. Performance sustained at 97.5% accuracy over 8.4 months and 248+ cumulative hours of use, with communication speeds reaching approximately 32 words per minute (Card et al., NEJM, 2024). This is a single-patient study, but the sustained performance over months in real-world conditions is a landmark result.

Real-time voice synthesis from neural signals (Nature, 2025): Building on the same participant, Wairagkar et al. decoded neural signals into audible synthesized speech in real time, with latency of 1/40th of a second (comparable to hearing one’s own voice). The system decoded paralinguistic features including prosody and intonation; the participant could “sing” short melodies. Intelligibility reached approximately 60% of synthesized words correctly recognized by listeners (Wairagkar et al., Nature, 2025).

Inner speech decoding (Cell, 2025): Kunz et al. at Stanford decoded imagined (inner) speech from motor cortex in 4 participants with ALS or brainstem stroke. The system achieved up to 74% accuracy with a 125,000-word vocabulary in real time, demonstrating that inner speech is robustly represented in motor cortex and highly correlated with attempted speech patterns (Kunz et al., Cell, 2025). This proof-of-concept raises the possibility of communication for patients who cannot attempt speech at all.

LLM-powered eye-gaze communication: Cai et al. developed SpeakFaster, an LLM-powered interface for highly abbreviated text entry via eye gaze. In two ALS eye-gaze users, the system achieved 29-60% faster text-entry rates versus baselines, and saved 57% more motor actions than traditional predictive keyboards in offline simulation (Cai et al., Nature Communications, 2024). Unlike invasive BCIs, this works with existing eye-tracking hardware.

BCI clinical-development landscape: Multiple academic and commercial BCI programs have included people with ALS or related paralysis. Recruitment status, enrollment, IDE status, and participant counts change frequently and should be verified in the exact trial registry and primary regulatory record.

Program Approach Status Key Detail
Neuralink PRIME and speech programs Intracortical, wireless Clinical development Verify current trial and regulatory status; company-reported Breakthrough designation is not authorization
Speech-specific implantable BCI studies Speech motor cortex decode Clinical development Verify the registry identifier and current recruitment status before referral
Synchron COMMAND Endovascular (Stentrode) Early feasibility evidence Small nonrandomized safety studies do not establish comparative communication benefit
Paradromics CONNECT-ONE High-bandwidth cortical Clinical development IDE permission to study a device is not marketing authorization
Cognixion ONE Axon Non-invasive EEG + AR Clinical development Company-reported Breakthrough designation is not authorization
BrainGate2 Intracortical Utah array Recruiting since 2009 Longest-running BCI program; published the landmark NEJM and Nature results

Non-invasive versus invasive BCIs: Invasive BCIs (BrainGate, Neuralink, Paradromics) offer higher bandwidth and accuracy but require neurosurgery. The Synchron Stentrode offers a middle ground: endovascular implantation through the jugular vein avoids craniotomy. Cognixion’s non-invasive EEG approach eliminates surgical risk entirely but achieves lower information transfer rates. For ALS patients, the choice depends on disease stage, surgical risk tolerance, and communication needs.

Assistive robotics and exoskeletons: The University of Queensland’s iMOVE-MND project (2025) represents the first trial of a wearable ankle exoskeleton specifically for MND patients, using machine learning to personalize gait assistance. The second-generation device features upgraded sensors and adaptive algorithms. Participants reported immediate improvement in walking confidence. Funded by the ALS Association; peer-reviewed results pending.

Ethics of BCI in locked-in ALS patients: BCIs for ALS raise unique ethical questions. Van Stuijvenberg et al. (2024) identified three core concerns: (1) predictive algorithms contribute to speech output, creating ambiguity about whose intentions are being expressed; (2) neural signal decoding may detect unintended thoughts, threatening mental privacy; (3) responsibility attribution for BCI-mediated speech is unclear (van Stuijvenberg et al., Frontiers in Human Neuroscience, 2024). A separate qualitative study of 19 BCI developers found that developers prioritize accuracy and reliability as conditions for user safety, authenticity, and mental privacy, though these goals may conflict with AI efficiency (van Stuijvenberg et al., Scientific Reports, 2024).

Informed consent is particularly complex: patients may be unable to meaningfully consent if the device cannot reliably distinguish intended speech from neural noise. For patients who are fully locked-in (no voluntary movement including eye movement), BCI may be the only means of communication, making the stakes of both deploying and withholding the technology extraordinarily high.

AI Drug Discovery and Clinical Trials for ALS

Drug repurposing via machine learning: Reimer et al. applied causal-inference and ML methods to health records from over 11,000 US veterans with ALS, evaluating 162 medications for survival associations. The analysis identified 27 medications with significant survival associations, including statins, PDE5 inhibitors, and alpha-adrenergic antagonists. PathFX network analysis identified converging downstream protein pathways across these drug classes (Reimer et al., Lancet Digital Health, 2026). This is retrospective and observational; it identifies candidates for prospective validation, not proven treatments.

AI-discovered drugs in trial: VRG50635, a PIKfyve inhibitor selected using Verge Genomics’ ConVERGE platform, entered a Phase 1b, open-label, single-group ALS study. The registry reports 54 enrolled participants and termination by the sponsor because of insufficient risk-benefit data; no results are posted (ClinicalTrials.gov, NCT06215755). The record does not support a claim that a prespecified efficacy endpoint failed, nor does it establish a class-wide conclusion about AI-assisted drug discovery.

HEALEY ALS Platform Trial: The HEALEY trial (NCT04297683, Massachusetts General Hospital) uses an adaptive platform design to test multiple drugs simultaneously with a shared placebo arm across 74+ US sites. Seven regimens have completed evaluation (zilucoplan, verdiperstat, CNM-Au8, pridopidine, trehalose, ABBV-CLS-7262, DNL343); none met the primary endpoint of ALSFRS-R change at 36 weeks. Despite no individual success, the platform validated adaptive trial design for ALS and continues enrolling new regimens.

AI-generated external comparators: ProJenX’s PRO-101 illustrates how digital twins are entering ALS therapeutic development as exploratory trial-design support rather than proof of drug efficacy. The registered Phase 1 study evaluates prosetin safety, tolerability, pharmacokinetics, and biomarkers, with an optional 52-week open-label extension for ALS participants who complete blinded Part C dosing (ClinicalTrials.gov, NCT05279755). ProJenX and Unlearn announced plans to use an ALS Digital Twin Generator to forecast ALSFRS-R, slow vital capacity, and plasma neurofilament light trajectories for extension participants, making the model a comparator for interpretation rather than a replacement for randomized efficacy evidence (ProJenX, September 2024). For clinicians reviewing trial claims, the key question is whether the digital twin’s context of use, source population, uncertainty, and sensitivity analyses were prespecified before model outputs influenced Phase 2 decisions.

Longitude Prize on ALS (launched June 2025): The £7.5 million Longitude Prize on ALS awarded £100,000 Discovery Awards to 20 teams in May 2026, together with access to curated ALS datasets (Longitude Prize on ALS). This is a funding mechanism, not a clinical study or evidence that any proposed target will produce an effective treatment.

The Broader MND Spectrum: AI Beyond ALS

Spinal Muscular Atrophy (SMA): SMA has more AI research than other MND-spectrum conditions, partly because gene therapies (nusinersen, onasemnogene abeparvovec, risdiplam) create a need for treatment response prediction. Taleb et al. (2024) developed an XGBoost classifier using computer vision pose estimation from infant movement videos to detect SMA-associated hypotonia, published in JAMA Pediatrics (Taleb et al., 2024). The MAP THE SMA trial (NCT05769465) aims to build ML algorithms predicting therapeutic response to all three approved SMA treatments. Stimpson et al. (2025, preprint) used oblique random forests to predict survival outcomes in 124 non-sitter SMA patients treated with nusinersen, achieving a C-Index of 0.74 (Stimpson et al., 2025, preprint).

FTD-ALS overlap: Approximately 15% of ALS patients develop frontotemporal dementia, often linked to C9orf72 repeat expansions. Dattola et al. (2025) conducted a systematic review of 25 AI studies (6,544 patients) for FTD differential diagnosis, finding SVMs and CNNs as the most common approaches (Dattola et al., Frontiers in Aging Neuroscience, 2025). No published study has used AI to predict C9orf72 carrier status from imaging or clinical data alone; existing work classifies known genetic carriers into subtypes.

Gaps in the MND-AI landscape: Published AI evidence for progressive muscular atrophy, Kennedy disease (SBMA), and post-polio syndrome is sparse compared with ALS. PMA overlaps clinically and biologically with ALS, while Kennedy disease has a molecular diagnostic pathway based on androgen-receptor CAG-repeat expansion. An apparent literature gap can reflect search coverage and indexing, so it should not be converted into a claim that no research exists. The imbalance nevertheless illustrates how data availability and commercial incentives can leave rare neuromuscular conditions underrepresented.

Regulatory Landscape for ALS AI

No ALS-specific AI device authorization is identified in the primary FDA records cited in this chapter. Several companies have reported Breakthrough Device Designations for ALS-relevant products. Breakthrough designation facilitates interaction and review; it is not clearance, approval, or evidence of clinical benefit.

Device Company Company-reported designation Intended research use
Cognixion ONE Axon Cognixion May 2023 Non-invasive BCI for communication
N1 Implant (speech) Neuralink May 2025 Decoding attempted speech from motor cortex
MyoRegulator PathMaker Neurosystems December 2025 Non-invasive neuromodulation for slowing ALS functional progression

The PathMaker MyoRegulator targets motor neuron hyperexcitability using multi-site direct current stimulation. An early feasibility study (NCT06165172, 15 patients, Beth Israel Deaconess) completed in January 2026, with top-line results presented at the 2026 MDA Clinical and Scientific Conference.

Eye-tracking speech generating devices (Tobii Dynavox I-Series) are cleared as Class I medical devices for augmentative communication but are not AI/ML-specific clearances.


Part 5: Psychiatric AI and the Limits of Risk Prediction

Why Psychiatric Prediction Is Difficult

Psychiatric diagnosis relies on: - Patient self-report of symptoms (low reliability) - Clinician assessment of behavior and affect (subjective, low inter-rater reliability) - Absence of biomarkers (no blood test for depression, no scan for schizophrenia) - Heterogeneous presentations (10 patients with depression may have 10 different symptom patterns)

These features make transportable prediction and causal treatment selection difficult. They do not imply that every computational tool in psychiatry is ineffective.

The Suicide Prediction Algorithm Failures

EHR suicide-attempt prediction research:

The promise: Predict suicide risk from EHR data (diagnoses, medications, ED visits, hospitalizations) and flag high-risk patients for intervention.

Evidence boundary: - Published Vanderbilt studies predicted suicide attempts, not completed suicides, and the authors stated that suicide deaths were unavailable as an outcome. - Positive predictive value depends on the selected threshold, prevalence, time horizon, and population. Low base rates can produce many false-positive classifications even when discrimination appears acceptable. - The cited studies do not establish that clinicians stopped responding to alerts or that a particular deployment caused harm.

Why it failed: - Base rate problem: Suicide is rare (even in high-risk populations, <1% attempt per year); any screening test yields massive false positives - Incomplete observation: EHR models cannot observe many acute stressors, unrecorded symptoms, or people outside the health system. - False reassurance: A low-risk label can discourage direct assessment even though no model rules out near-term risk.

Social Media Suicide Prevention AI:

The promise: Detect suicidal content in posts and live videos; alert human reviewers to contact users in crisis.

The reality: - Platforms have deployed AI systems to detect suicide-related content - Limited published data on effectiveness in preventing actual suicides - No peer-reviewed evidence demonstrating reduced suicide rates from algorithmic detection - Privacy advocates raise concerns about surveillance and consent

The lesson: A risk score must not replace direct assessment, safety planning, clinical follow-up, or crisis response. Any deployment needs evidence that the complete workflow improves patient-relevant outcomes without unacceptable false reassurance, coercion, or resource diversion.

Depression Diagnosis from Digital Phenotyping: Privacy Nightmare

The concept: Passively monitor smartphone use (typing speed, app usage, GPS movement patterns, voice call frequency) to detect depression without patient self-report.

The problems: 1. Consent: Is continuous monitoring with periodic algorithm-generated diagnoses truly informed consent? 2. Privacy: Smartphone data reveals intimate details of life (where you go, who you talk to, what you search) 3. Accuracy: Correlation between “reduced movement” and depression does not mean algorithm can diagnose depression (could be physical illness, weather, life circumstances) 4. Equity: Algorithms trained on white populations may interpret cultural differences in communication or movement as pathology

Current status: Research-stage only. Multiple academic studies, zero validated clinical applications.

Ethical consensus: Most bioethicists and psychiatrists agree: Digital phenotyping for psychiatric diagnosis raises profound ethical concerns that have not been resolved.


Part 6: Equity in Neurological AI

The Underappreciated Problem

Stroke, MS, Parkinson’s, and Alzheimer’s all show different prevalence, presentation, and prognosis across racial and ethnic groups:

  • Stroke: Black Americans have 2x stroke incidence of white Americans, different stroke subtypes
  • MS: More common in white populations; Black patients with MS have more aggressive disease
  • Alzheimer’s: Higher prevalence in Black and Hispanic Americans, often diagnosed later
  • Parkinson’s: Lower prevalence in Black populations, different symptom profiles

Yet most neurology AI algorithms are trained on: - Predominantly white populations - Tertiary academic medical centers - North American and European datasets

Consequences:

LVO detection algorithms: - Preliminary data suggests equivalent performance across races - But comprehensive equity studies not published for most commercial systems - Ask vendors: “What is sensitivity/specificity stratified by race?”

MS lesion detection: - Trained mostly on white Scandinavian and North American populations (where MS prevalence is highest) - Performance in Black and Hispanic MS patients unknown

Alzheimer’s neuroimaging: - Hippocampal volume norms based on white populations - Black Americans have different brain volumetrics; algorithms may misclassify

What neurologists should do: 1. Ask vendors for race-stratified performance data 2. Validate algorithms locally on your patient population 3. Monitor for algorithmic errors by race/ethnicity 4. Do not assume “overall accuracy” applies to all patients


Part 7: Implementation Framework

Before Adopting Neurology AI

Questions to ask vendors:

  1. “Where is the peer-reviewed study showing this algorithm improves patient outcomes?”
    • LVO detection has this evidence (Karamchandani et al. 2023, Devlin et al. 2022)
    • Most other neurology AI does not
    • Demand JAMA Neurology, Lancet Neurology, Stroke, Neurology, Brain and Behavior publications
  2. “What is the algorithm’s performance in patients like mine?”
    • Academic medical center algorithms may fail in community hospitals
    • Pediatric algorithms do not work in adults
    • Request validation data from similar patient populations
  3. “What is the false positive rate, and how will we manage false alarms?”
    • Convert the product’s validated false-positive rate into an expected alert count using the hospital’s actual case mix and imaging volume
    • Who triages these? What’s the workflow?
  4. “How does this integrate with our PACS/EHR/radiology workflow?”
    • Demand live demonstration in your specific environment
    • Poor integration = alert fatigue = missed critical cases
  5. “What happens when the algorithm fails?”
    • All algorithms miss some cases
    • Review product- and population-specific false negatives rather than assuming a class-wide miss rate
    • The error analysis must identify which cases are likely to be missed, including posterior-circulation, tandem, or distal occlusions when relevant to the product’s intended use
  6. “Can we validate locally before full deployment?”
    • Choose a local validation sample large enough for the intended endpoint, expected prevalence, subgroup analysis, and acceptable uncertainty
    • Compare algorithm performance to actual outcomes in your population
  7. “What are the equity implications?”
    • Request race/ethnicity-stratified performance metrics
    • If vendor does not have this data, algorithm was not validated equitably
  8. “Who is liable if the algorithm misses a critical finding?”
    • Read the vendor contract carefully
    • Identify the duties assigned by product labeling, local policy, professional standards, contracts, and applicable law
    • FDA authorization does not determine civil liability, and responsibility is not assigned by a universal rule
  9. “What is the cost, and what’s the evidence of cost-effectiveness?”
    • Obtain a current written quotation and include interfaces, validation, staffing, monitoring, upgrades, and false-alert review
    • Require a comparative economic evaluation that matches the product, workflow, population, and patient-relevant endpoint
    • Treat absence of such evidence as an evidence gap, not proof that the tool is or is not cost-effective
  10. “Can you provide references from neurologists who use this tool?”
    • Talk to actual users
    • Ask about false positives, workflow disruptions, whether they’d recommend it

Red Flags (Walk Away If You See These)

  1. Claims to diagnose psychiatric conditions autonomously from digital traces without a validated clinical pathway
  2. Suicide-risk prediction reported without predictive values, calibration, threshold, time horizon, and an evaluated response pathway
  3. No external validation studies (validated only in development cohort)
  4. Vendor refuses to share peer-reviewed publications (“proprietary algorithm”)
  5. Black box psychiatric AI (explainability is essential for consent and ethics)

Part 8: Cost and Value Assessment

What Does Neurology AI Cost?

Commercial prices, bundles, and coverage change and are usually negotiated. The figures below should therefore be obtained from current written quotations and total-cost models, not static handbook ranges.

Stroke AI (LVO/ICH detection): - Include licensing, interfaces, imaging routing, mobile access, validation, training, response staffing, monitoring, upgrades, and incident review - Separate costs for LVO detection, perfusion analysis, care coordination, and ICH triage

Brain tumor segmentation: - Determine whether software is bundled into a planning system or priced per site, user, study, or module - Include expert correction time and downstream plan review

EEG seizure detection: - Include acquisition hardware, algorithm licenses, review software, storage, technologist work, neurologist verification, and false-positive burden

MS lesion detection (research only): - Not commercially available for routine use

Parkinson’s/ALS apps: - Mostly research tools, not commercial products

Brain-computer interfaces and assistive communication: - Investigational implantable systems do not have a stable commercial price or routine coverage pathway - Compare surgical and device burden with established augmentative communication options - Verify current payer coverage and patient-specific support rather than asserting national absence

Do These Tools Save Money?

LVO detection: Potential value depends on whether the local workflow actually shortens treatment and changes disability outcomes. The randomized Viz.ai trial showed smaller process-time reductions than the 43-minute assumption and did not significantly improve 90-day functional independence.

ICH detection: Worklist prioritization may shorten review, but a value model must measure downstream action, false positives, missed cases, staffing, and patient outcomes. “Probably cost-effective” is not a substitute for a comparative economic evaluation.

Brain tumor segmentation: Measure actual correction time, planning throughput, expert labor, software cost, and plan quality locally. A hypothetical hourly wage multiplied by a borrowed time-saving estimate does not establish savings.

Seizure detection: Time savings may create value, but the full review workflow, false detections, missed seizures, and patient-relevant outcomes must be measured.

MS, Alzheimer, and Parkinson tools: Research or diagnostic-support value is use-case-specific. Absence of a cited economic evaluation should be stated as an evidence gap, not converted into a universal “no.”


Part 9: Future Neurology AI

Evidence-Calibrated Development Horizon

Demonstrated or in active clinical development: 1. Expanded stroke AI: Posterior circulation stroke detection, stroke mimics identification 2. Automated EEG reporting: Summary reports for routine EEGs (not just seizure detection) 3. Neurosurgical planning AI: Tumor resection planning, deep brain stimulation targeting 4. Gait analysis from smartphone video: Parkinson’s and ataxia monitoring 5. ALS blood biomarker panels: Plasma proteomics (AUC 98.3%) and gene expression classifiers approaching clinical-grade performance; awaiting prospective multi-site validation 6. Brain-computer interfaces for communication: Multiple devices are in clinical studies; company-reported regulatory designations do not predict an approval date 7. Digital biomarkers as ALS trial endpoints: Wearable accelerometers and speech biomarkers detecting progression faster than ALSFRS-R, potentially reducing trial duration and sample size requirements

Plausible but not established: 1. Alzheimer’s blood biomarkers + AI: Combining plasma p-tau217 with MRI and cognitive testing for early diagnosis 2. Seizure forecasting: Predicting when seizure will occur hours in advance (still research stage) 3. Precision psychiatry: Matching patients to antidepressants based on genetics + symptoms (early trials) 4. AI-discovered ALS therapeutics: Drug repurposing candidates identified through ML (statins, PDE5 inhibitors per Reimer et al., Lancet Digital Health 2026), but no AI-discovered drug has succeeded in ALS trials 5. Inner speech decoding BCIs: Decoding imagined speech is a small-participant proof of concept; clinical utility, privacy protection, and generalizability remain unresolved

Overhyped and unlikely: 1. Autonomous psychiatric diagnosis from digital phenotyping 2. Suicide prediction from social media 3. AI replacing neurological examination

The rate-limiting factor: Not algorithmic accuracy. Prospective RCTs showing improved patient outcomes and ethical frameworks for psychiatric AI.


Professional Society Guidance Relevant to Neurology AI

AAN Resources and Endorsed Principles

The American Academy of Neurology position-statement directory lists the AMA Principles for Augmented Intelligence Development, Deployment, and Use (2023) among statements endorsed by the AAN. The AAN’s separate Artificial Intelligence Resources page emphasizes limitations, compliance review, and prospective silent pilots. These are professional resources and endorsed principles, not product-specific clinical-practice recommendations.

AAN Annual Meeting AI Sessions (2024-2025):

Meeting sessions are educational content rather than formal society guidance. The 2024 session “Artificial Intelligence (AI) and the Neurologist: New Horizons” covered:

  • Types of AI and machine learning relevant to neurology
  • Present and potential clinical applications
  • Benefits and challenges AI creates for neurologists

At AAN 2025, researchers discussed: - AI-driven behavioral analysis for Alzheimer’s disease progression modeling - Machine learning for identifying patients at high risk for hematoma expansion

AI Applications Addressed by Neurology Societies

Stroke: - LVO detection algorithms (Viz.ai, RapidAI) integrated into stroke systems of care - ASPECTS scoring automation - Perfusion imaging analysis

Epilepsy: - Automated EEG seizure detection - Long-term monitoring pattern recognition - Seizure prediction research

Neurodegenerative Disease: - Imaging biomarkers for Alzheimer’s disease - Parkinson’s disease tremor analysis - ALS progression modeling

American Clinical Neurophysiology Society (ACNS)

ACNS has published extensive guidelines and consensus statements on EEG standards, including:

  • Standardized Critical Care EEG Terminology (2021 version)
  • Minimum technical requirements for clinical EEG
  • Guidelines for continuous EEG monitoring in critical care
  • Standards for neonatal EEG monitoring

These technical standards provide the foundation for evaluating automated EEG analysis systems, though ACNS has not published specific guidance on AI implementation. The society’s emphasis on standardized terminology and technical requirements supports interoperability with algorithmic tools.

International League Against Epilepsy (ILAE)

The ILAE and International Federation of Clinical Neurophysiology issued a clinical-practice guideline for automated seizure detection using wearable devices. It conditionally recommends clinically validated devices for detecting generalized tonic-clonic and focal-to-bilateral tonic-clonic seizures in selected unsupervised patients when an alarm can lead to rapid intervention. It does not recommend current devices for other seizure types, and it found uncertainty about whether alarms improve meaningful patient outcomes (Beniczky et al., 2021; ILAE guideline PDF).

The guideline supports a narrow adjunctive use for specified seizure types. It does not validate seizure forecasting, autonomous medication changes, or every wearable marketed for epilepsy.

Endorsed Decision Support Tools

Validated tools integrated into neurological practice include:

  • NIH Stroke Scale: Standardized severity assessment
  • ASPECTS: CT-based stroke scoring
  • ABCD2: TIA stroke risk stratification

These represent the foundation for algorithmic decision support in neurology, with AI-enhanced versions under development.


Key Takeaways

10 Principles for Neurology AI

  1. Stroke AI has strong workflow evidence: Viz.ai-specific randomized evidence supports shorter treatment intervals, but not a significant improvement in 90-day functional independence

  2. Time-sensitive triage and autonomous diagnosis are different claims: Algorithms can prioritize a suspected finding while clinical interpretation establishes what it means

  3. Psychiatric AI requires strict deployment boundaries: A risk score must not replace direct assessment, safety planning, or an evaluated clinical response pathway

  4. ALS/MND AI is advancing on multiple fronts: Biomarker classifiers, wearable digital biomarkers, brain-computer interfaces, and drug-repurposing research are active. Results remain cohort- and task-specific. Company-reported Breakthrough Device Designations are not authorization.

  5. Seizure detection saves neurologist time: Automated EEG review is useful adjunct, not replacement

  6. Equity data is missing: Most algorithms lack race-stratified performance metrics; validate locally

  7. Integration determines success: Even accurate LVO detection fails with poor workflow integration

  8. False positives cause alert fatigue: Translate validated error rates into expected alert counts using local prevalence, volume, threshold, and routing

  9. Demand RCT evidence: Technical accuracy ≠ improved patient outcomes

  10. Neurological exam remains irreplaceable: AI assists with imaging and data interpretation but cannot replace clinical judgment


Clinical Scenario: Evaluating a Neurology AI Tool

Fictional Scenario: A Neurology Department Considers MS Lesion Detection AI

The pitch: A fictional vendor demonstrates AI that detects and quantifies MS lesions on brain MRI. All product claims and prices in this teaching scenario are illustrative, not measured facts. The vendor shows: - Sensitivity 94% for lesion detection - “Reduces radiologist reading time by 50%” - Automatic comparison to prior scans to detect new lesions - Cost: $75,000/year

The department chair asks for your recommendation.

Questions to Ask:

  1. “What peer-reviewed publications validate this algorithm?”
    • Look for JAMA Neurology, Multiple Sclerosis Journal, Neurology publications
    • Commowick et al. (2018) study showed MS algorithms are scanner-dependent
  2. “How does this algorithm perform on our specific MRI scanner?”
    • Most MS AI trained on Siemens scanners
    • Performance often degrades on GE or Philips scanners
    • Request validation data on your exact scanner model and protocol
  3. “What is the false positive rate for lesion detection?”
    • Small vessel ischemic disease, migraine, normal aging create white matter hyperintensities
    • How does algorithm distinguish MS lesions from mimics?
  4. “Does this algorithm change clinical management?”
    • MS diagnosis still requires McDonald criteria (clinical + imaging + CSF/evoked potentials)
    • Lesion burden does not directly guide treatment decisions in most cases
    • If algorithm does not change what you do, why pay $75,000/year?
  5. “How will this integrate with our radiology workflow?”
    • Does it require manual upload of prior scans?
    • Does it work with our PACS?
    • Who reviews the AI output, radiologist or neurologist?
  6. “What happens with atypical presentations?”
    • Tumefactive MS (large lesions mimicking tumors)
    • Posterior fossa lesions (often missed by algorithms)
    • Infratentorial disease
  7. “Can we pilot on 100 MS patients before committing to $75,000/year?”
    • Compare algorithm lesion counts to expert neuroradiologist
    • Measure actual time savings
    • Assess false positive burden

Red Flags in This Scenario:

“Reduces reading time by 50%” without time-motion study data: Unverified claim

No external validation studies: If algorithm was validated only at vendor’s institution, performance at your hospital uncertain

Sensitivity 94% without specificity data: Useless. High sensitivity with low specificity = many false positives

Scanner-agnostic claims: MS lesion AI is notoriously scanner-dependent; claims of universal performance are suspicious

No discussion of clinical utility: Detecting lesions does not equal improving patient outcomes


Check Your Understanding

The following cases are fictional decision exercises. They do not describe actual patients, products, outcomes, legal holdings, or standards of care.

Scenario 1: The LVO Alert at 2 AM

Clinical situation: You’re the interventional neuroradiologist on call. At 2:15 AM, you receive a Viz.ai mobile alert: 68-year-old man, right-sided weakness, NIHSS 14, CT angiography shows left M1 occlusion. Patient is at community hospital 30 minutes away by helicopter.

You open the Viz.ai app and review the images. The M1 occlusion is clear. But you also notice the patient had a large left MCA territory stroke 2 years ago (visible on non-contrast CT as encephalomalacia).

Question 1: Do you mobilize the cath lab team for thrombectomy?

Click to reveal answer

Answer: Yes, but with additional information gathering.

Reasoning: - Acute M1 occlusion is a thrombectomy indication regardless of prior stroke history - Prior stroke in same territory does not automatically exclude thrombectomy, but does require considering: - What is baseline disability? (if mRS 5 before this stroke, thrombectomy less likely to benefit) - What is new deficit vs. baseline? (need to talk to patient’s family or primary physician) - What is stroke onset time? (still within window?)

What to do: 1. Accept the transfer (do not delay getting patient to comprehensive stroke center) 2. Call community hospital ED while patient is en route: - What was baseline functional status? - What is time last known normal? - Any contraindications to thrombectomy (anticoagulation, recent surgery)? 3. Review imaging carefully when patient arrives: - Confirm acute occlusion (not chronic occlusion from prior stroke) - Assess collateral status - Check ASPECTS score for ischemic core 4. Make final decision based on: - Baseline function (if mRS 0-2, proceed) - Deficit is new and severe (NIHSS 14 indicates major deficit) - Time window (if <6 hours from onset, proceed; if 6-24 hours, check perfusion imaging)

Bottom line: LVO alerts are not autonomous treatment decisions. They accelerate evaluation and transfer, but clinical judgment remains essential.

Prior stroke in same territory is NOT an absolute contraindication. Many patients with prior stroke in one territory can have new stroke in another branch and benefit from thrombectomy.

The LVO algorithm did its job: Detected acute occlusion and got patient to you faster. Your job is clinical decision-making.


Scenario 2: The Suicide Risk Algorithm

Clinical situation: Your hospital deployed a suicide risk prediction algorithm that analyzes EHR data. You’re seeing a 32-year-old woman in primary care for diabetes follow-up. She mentions feeling “a bit down lately” and having trouble sleeping.

You check the EHR: The suicide risk algorithm has flagged her as “LOW RISK” (10th percentile).

Question 2: Do you skip the detailed depression and suicidality screening because the algorithm says low risk?

Click to reveal answer

Answer: No. The low-risk output does not replace the clinical assessment prompted by the patient’s symptoms.

Reasoning:

Why suicide risk algorithms fail: 1. Base rate problem: Positive predictive value depends on event prevalence, outcome definition, prediction horizon, threshold, and population. A high discrimination metric can coexist with many false positives. 2. “Low risk” can create false reassurance: A model cannot exclude acute thoughts or stressors that are absent from its inputs 3. Algorithms cannot detect acute stressors: EHR data from last week does not capture what happened this morning (job loss, relationship breakup, eviction notice) 4. Clinical presentation matters: Patient saying “I’m feeling down” is a red flag requiring exploration, regardless of algorithmic score

What you should do: 1. Do not use the algorithm to skip assessment 2. Perform standard depression screening: - PHQ-9 questionnaire - Direct questions about suicidal ideation: “Have you thought about hurting yourself?” - Risk factors: Prior attempts, family history, access to means, substance use 3. Document the information actually assessed: Record the symptoms, direct questions, risk and protective factors, clinical reasoning, safety plan, and follow-up rather than treating the score as dispositive. 4. Review the deployed pathway: If the output is used, the institution should define its intended use, validation evidence, escalation pathway, monitoring, and failure response. A tool without a safe and evaluated response pathway should not influence care.

Why this matters: - Patient is telling you she’s depressed (“feeling down, trouble sleeping”) - Believe the patient, not the algorithm - Direct assessment remains necessary when symptoms or other clinical information raise concern - No algorithm can substitute for the duties defined by the clinical context, institutional policy, and applicable professional standards

Bottom line: Suicide-risk models can cause harm when they create false reassurance or high false-alert burden. Their value depends on the defined population, endpoint, threshold, calibration, clinical response, and demonstrated effect on care.

If a hospital has deployed one, do not use the score as a substitute for assessment. Request evidence and governance review if the model’s role, validation, or response pathway is unclear.


Scenario 3: The MS Lesion Count Discrepancy

Clinical situation: You’re evaluating a 29-year-old woman with optic neuritis and a single brain MRI white matter lesion. Multiple-sclerosis diagnosis requires application of current diagnostic criteria, including evidence relevant to dissemination in space and time and exclusion of better explanations. A lesion count alone is not sufficient. The radiologist’s report says “1 lesion in left periventricular white matter.”

However, the MS lesion detection AI (which your hospital recently deployed) flags 4 lesions total: the one the radiologist saw, plus 3 additional small (3-4mm) lesions in different locations.

Question 3: Do you diagnose MS based on the AI lesion count of 4?

Click to reveal answer

Answer: No. Review the MRI yourself with a neuroradiologist before making MS diagnosis.

Reasoning:

Why lesion count discrepancies happen: 1. AI false positives: Small vessel ischemia, perivascular spaces, artifacts often flagged as MS lesions 2. Size threshold differences: Radiologists may not report 2-3mm lesions; AI flags everything 3. Lesion location and timing matter: Current criteria use specified central-nervous-system regions and evidence of dissemination in time or accepted alternatives. AI may count lesions in non-diagnostic locations or combine lesions that do not satisfy the criteria.

What you should do:

  1. Review the MRI images yourself:
    • Look at the 3 additional lesions the AI detected
    • Are they in McDonald criteria locations?
    • Do they look like MS lesions (ovoid, perpendicular to ventricles, Dawson fingers)?
    • Or do they look like small vessel ischemic disease, perivascular spaces, artifacts?
  2. Get neuroradiology over-read:
    • “AI detected 4 lesions; report lists 1. Can you review and clarify?”
    • Experienced neuroradiologist can distinguish MS plaques from mimics
  3. Consider additional testing:
    • Spinal cord MRI (may show additional lesions supporting MS diagnosis)
    • CSF analysis for oligoclonal bands
    • Evoked potentials
    • Wait for second clinical event (if not urgent to start treatment)
  4. Do not diagnose MS based solely on AI lesion count:
    • MS diagnosis has major implications (lifelong immunosuppression, insurance, disability, prognosis)
    • Requires high confidence, not algorithmic suggestion

Why this matters: - False positive MS diagnosis causes harm: Unnecessary treatment with expensive immunosuppressants that have side effects - McDonald criteria exist for a reason: They require clinical + imaging evidence to prevent overdiagnosis - AI lesion detection is a tool, not a diagnosis: It flags possible lesions for expert review, not autonomous diagnostic conclusions

Fictional cautionary example: A patient is diagnosed with MS after an automated lesion count is accepted without expert review and begins a disease-modifying therapy. Later specialist review concludes that the small lesions are more consistent with a mimic. The example illustrates the downstream consequences of accepting a lesion count as a diagnosis; it is not an actual reported case or measured outcome.

Bottom line: Use MS lesion AI to help detect possible lesions, but never to diagnose MS without expert confirmation.

When AI and expert review disagree, adjudicate the disagreement against the images, current diagnostic criteria, relevant biomarkers, and an appropriately qualified specialist. Neither label should be accepted without review.


Questions About Neurology AI

How much time does stroke AI save?

Viz.ai-specific studies support shorter workflow intervals. A stepped-wedge cluster-randomized trial found adjusted reductions of 11.2 minutes in door-to-groin time and 9.8 minutes in CT-to-treatment time, without a significant improvement in 90-day functional independence. Effects should not be transferred to other products.

What is the sensitivity of AI for intracranial hemorrhage detection?

There is no class-wide sensitivity figure for intracranial-hemorrhage AI. In one workflow study, the evaluated detector had 73% sensitivity, 80% specificity, and AUC 0.85 while worklist reprioritization shortened time to diagnosis for routine outpatient head CT.

Does AI work for suicide risk prediction?

Suicide-risk models have shown low positive predictive value and cannot replace direct assessment or validated clinical pathways. Performance varies by population, outcome, time horizon, and threshold; a low-risk label must never provide false reassurance.

Can AI diagnose Alzheimer’s disease?

AI can quantify imaging features associated with Alzheimer disease, but it cannot establish the diagnosis by itself. Results must be interpreted with symptoms, examination, validated biomarkers, treatment eligibility, and the exact intended use of the tool.

Can AI diagnose ALS from a blood test?

Research classifiers using plasma proteomics, gene expression, and electrophysiology have reported promising discrimination for ALS in defined cohorts. They are not autonomous blood-test diagnoses, and they require prospective multisite validation against relevant mimics before clinical use.

How do brain-computer interfaces help ALS patients communicate?

Implanted brain-computer interfaces can decode attempted speech or movement-related neural activity into text or synthesized speech. A 2024 BrainGate report described 97.5% accuracy at approximately 32 words per minute in one participant with ALS, an important single-participant demonstration rather than population-level effectiveness evidence.

Are there FDA-cleared AI devices for ALS?

No ALS-specific AI device authorization is identified in the primary FDA records cited in this chapter. Company-reported Breakthrough Device Designations are not marketing authorization and should be verified from an exact primary record before clinical interpretation.