Psychiatry and Behavioral Health

Many psychiatric diagnoses have no single confirmatory laboratory or imaging test. Assessment instead integrates symptoms, behavior, function, longitudinal history, collateral information, and clinical judgment. That heterogeneity does not make psychiatric AI impossible, but it makes the target, time horizon, comparator, and intended action unusually important. Suicide risk models illustrate the distinction: a model may stratify recorded risk in a defined population without demonstrating that its use prevents suicide or improves care.

Learning Objectives

After reading this chapter, you will be able to:

  • Distinguish suicide risk stratification from demonstrated clinical utility
  • Evaluate digital phenotyping approaches and their significant privacy and accuracy concerns
  • Assess AI chatbot therapy applications (Woebot, Wysa) with appropriate skepticism
  • Recognize how diagnostic heterogeneity, outcome definition, and context affect psychiatric AI validation
  • Navigate the unique ethical considerations of AI in mental health care
  • Identify the rare psychiatric AI applications that may provide clinical value
  • Apply evidence-based frameworks for evaluating behavioral health AI

The Clinical Context: Psychiatric assessment often integrates subjective symptoms, observed behavior, functional change, longitudinal course, and social context rather than one confirmatory test. The resulting labels and outcomes can be heterogeneous, so validation must specify what a system predicts, for whom, over what period, and what action follows.

What Does Not Work:

Application Outcome Lesson
Vanderbilt Suicide Risk Algorithm Low PPV for return encounters involving suicide attempt Risk ranking is not proof that model use prevents suicide
Meta suicide-prevention systems Company-described detection and escalation workflow No peer-reviewed patient-outcome evaluation identified
Depression signals from digital phenotyping Associations in heterogeneous research cohorts Individual diagnosis, transportability, privacy, and clinical utility remain separate questions
Autonomous psychiatric diagnosis No authorized example is established in the FDA records reviewed here Intended use and authorization must be checked device by device

What Shows Limited Promise:

Application Status Caveats
Structured CBT-oriented chatbots Randomized and meta-analytic evidence of short-term symptom improvement Low-certainty evidence, heterogeneous interventions, and prediction intervals that include no effect
NLP for clinical note analysis Research stage Documentation support, not diagnosis
Treatment response prediction Early research Not ready for clinical use

Critical Insights:

  • Risk classification is not prevention: Discrimination and PPV do not show that displaying a score improves outcomes
  • A “low-risk” label cannot close the assessment: Acute presentation and current clinical information still govern evaluation
  • Digital phenotyping raises profound ethical concerns: Continuous smartphone monitoring without clear benefit
  • Rare outcomes constrain positive predictive value: The acceptable tradeoff depends on the intervention, resource burden, and harm from false classifications
  • Harm-focused chatbot reviews are not risk–benefit verdicts: A 2026 scoping review mapped 119 papers into five clusters; most crisis-response work is vignette-based and AI-psychosis writing is almost entirely conceptual (Diel et al., 2026)

The Bottom Line: Psychiatric AI requires narrow claims and prospective evidence tied to an actionable workflow. No score or chatbot should replace a current suicide assessment. Structured digital interventions, documentation support, and treatment decision support may add value, but their evidence cannot be transferred to general-purpose chatbots or autonomous diagnosis.


Part 1: Why Psychiatric AI Is Difficult to Validate

The Fundamental Problem

Psychiatric diagnosis relies on:

  • Patient self-report of symptoms (variable reliability)
  • Clinician assessment of behavior and affect (subjective, variable inter-rater reliability)
  • Absence of biomarkers (no blood test for depression, no scan for schizophrenia)
  • Heterogeneous presentations (10 patients with depression may have 10 different symptom patterns)

These features make the prediction target less stable than a single laboratory measurement. They also make clinical context indispensable. A model trained to reproduce a diagnostic code may learn documentation and access patterns rather than the full clinical construct.

The Suicide Prediction Algorithm Failures

Vanderbilt Suicide Attempt and Ideation Likelihood (VSAIL) Model:

The Vanderbilt University Medical Center (VUMC) implemented a suicide risk prediction model in their Epic EHR, one of the most rigorously studied implementations.

Prospective validation study (2019-2020): (Walsh et al., JAMA Network Open, 2021)

  • 115,905 predictions for 77,973 patients over 296 days
  • Approximately 392 predictions per day
  • Patient demographics: 54% men, 45% women, 78% White, 16% Black

Performance in the highest risk group:

Outcome Positive Predictive Value Number Needed to Screen
Suicidal ideation 3-4.3% 23
Suicide attempt 0.3-0.4% 271

What this means: In that study setting, approximately one of every 271 patients in the highest-risk group had a recorded return encounter involving a suicide attempt during the evaluated horizon. The remaining patients did not have that recorded outcome during the study window. This does not mean they had no clinical need, and it does not establish whether an intervention triggered by the score would help.

Hybrid approach (2022): (Wilimitis, Walsh et al., JAMA Network Open, 2022)

Combining the VSAIL machine learning model with in-person Columbia Suicide Severity Rating Scale (C-SSRS) screening improved performance:

  • PPV for suicide attempt: 1.3-1.4% (vs. 0.4% for VSAIL alone)
  • Sensitivity for suicide attempt: 77.6-79.5%

Why the improved prediction result does not establish clinical utility:

  1. Base rate constraint: Rare outcomes can produce low PPV even when discrimination or sensitivity appears favorable
  2. Resource implications: A high alert volume may consume clinical capacity unless the response pathway and threshold are designed and tested
  3. Alert fatigue: Clinicians stop responding to flags
  4. False reassurance: “Low risk” predictions give dangerous false confidence

The comparison often cited: Walsh and colleagues noted that the number needed to screen for the recorded attempt outcome was comparable to some accepted screening programs. That comparison does not settle utility, because the downstream intervention, harms, cost, and outcome evidence differ across conditions.

Facebook/Meta suicide-prevention systems:

Meta described machine-learning systems intended to identify potential suicide and self-injury content, route high-priority cases for human review, and connect users or emergency services with support (Meta Engineering, 2018). The company has continued to describe suicide and self-harm safeguards, including a 2026 parent-notification feature for supervised teen accounts (Meta, 2026). These company descriptions establish the workflow, not a reduction in suicide attempts or deaths. No peer-reviewed patient-outcome evaluation was identified for the program.

The lesson: A suicide risk model can be evaluated as a prediction model, but a clinical deployment claim requires more. It must specify a response pathway, measure false-positive and false-negative consequences, and prospectively test whether use improves patient-relevant outcomes.

International Consensus Against Suicide Risk Scales

The failure of suicide prediction algorithms is now reflected in official clinical guidelines. The VA/DoD Clinical Practice Guideline for Assessment and Management of Patients at Risk for Suicide (2024) explicitly notes that “current algorithms will be correct only about 1% of the time” among those classified as at risk. The guideline finds insufficient evidence to recommend for or against suicide risk screening programs.

International bodies have reached similar conclusions. Guidelines from the UK (NICE NG225), Australia, and New Zealand explicitly advise against using risk assessment scales for prediction and treatment allocation (Knipe et al., Lancet, 2022). NICE states directly: “Do not use risk assessment tools and scales to predict future suicide or repetition of self-harm.”

The cited prediction studies do not establish that displaying their scores prevents suicide attempts or deaths. A risk category also cannot substitute for a current assessment of intent, plan, means, protective factors, acute stressors, and available support.


Part 2: Digital Phenotyping: Promise and Peril

The Concept

Passively monitor smartphone use (typing speed, app usage, GPS movement patterns, voice call frequency, accelerometer data) to detect depression, anxiety, or mood changes without patient self-report.

What the Research Shows

A systematic review of digital phenotyping for stress, anxiety, and mild depression found that smartphone sensors can identify behavioral patterns associated with mental health symptoms (Choi et al., 2024).

Sensors used in studies:

  • GPS (location, movement patterns)
  • Accelerometer (physical activity)
  • Bluetooth and Wi-Fi (social proximity)
  • Ambient audio and light sensors
  • Screen usage patterns
  • Typing dynamics

Performance claims require their original context:

  • Individual studies have reported high classification metrics, but the result depends on the sample, label, sensing window, missing-data handling, and validation design
  • Some research evaluates whether signals precede symptom-score changes; that is not equivalent to predicting a clinically diagnosed episode
  • Location, activity, and communication patterns can be associated with mental health measures without functioning as specific diagnostic biomarkers

Why These Claims Require Skepticism

  1. Study populations: Most research involves nonclinical cohorts or self-identified depression via questionnaires, not formally diagnosed patients
  2. Modality and missingness: Combining sensors can add information but also increases data loss, privacy exposure, and opportunities for overfitting
  3. External validation: Performance in one cohort or device environment may not transport to another population, operating system, or care setting
  4. Active vs. passive data: Many “passive sensing” studies still rely on self-report surveys for outcome measurement

The Problems

  1. Consent: Is continuous monitoring with periodic algorithm-generated diagnoses truly informed consent? Can consent be withdrawn without losing access to care?
  2. Privacy: Smartphone data reveals intimate details: where you go, who you talk to, what you search, when you sleep
  3. Accuracy in practice: Correlation between “reduced movement” and depression does not mean an algorithm can diagnose depression in individual patients
  4. Equity: Algorithms trained predominantly on white, Western populations may interpret cultural differences as pathology
  5. Data security: Who owns the data? What happens if it’s breached or sold?
  6. Coercion potential: Could employers, insurers, or courts access mental health inferences from passive data?

Current Status

The evidence base remains dominated by observational and feasibility studies. A 2025 longitudinal study demonstrated that extended data collection in adolescents was feasible, but it did not establish that digital-phenotyping alerts improve diagnosis, treatment, or patient outcomes (Huang et al., 2025). Feasibility, predictive validity, and clinical utility are three separate evidentiary steps.

Ethical Consensus

Digital phenotyping raises unresolved questions about consent, secondary use, group privacy, clinical responsibility, and the interpretation of missing or behaviorally generated data. The gap between “a pattern can be detected” and “this system should guide care” requires technical, clinical, and ethical evaluation.


Part 3: Applications With Emerging Clinical Evidence

Chatbot Therapy Applications

The idea of computer-delivered psychotherapy dates to 1966, when Stanford psychiatrist Kenneth Colby proposed that timesharing computers could scale therapeutic interventions beyond what human therapists could provide (Colby et al., 1966). MIT’s Joseph Weizenbaum rejected this vision as “immoral” in his 1976 book Computer Power and Human Reason, sparking a debate that remains unresolved. See History of AI in Medicine for this foundational controversy.

The evolution of AI-based mental health interventions accelerated in the 2010s, with early work establishing foundational principles for digital psychiatric tools. A 2020 review characterized this development trajectory and identified core design considerations that continue to shape the field (D’Alfonso, 2020).

Woebot, Wysa, and similar apps provide CBT-based interventions via smartphone. These represent the most studied psychiatric AI applications.

Evidence from systematic reviews:

A systematic review examining studies published from 2017 through 2024 found heterogeneous evidence across Woebot, Wysa, and Youper interventions. Designs, populations, comparators, intervention duration, and outcome measures differed, so product names alone do not establish a common effect (Farzan et al., 2025).

Chatbot Effect on Depression Effect on Anxiety Key Populations
Woebot Evaluated in randomized and observational studies Results depend on comparator and population College students and other adult cohorts
Wysa Evaluated in randomized and observational studies Product-specific results should not be transferred to other chatbots Chronic-disease and maternal-health cohorts among studied populations
Youper Limited study representation in the cited review Evidence is less mature than the better-studied interventions Adult users in the identified study

Meta-analysis findings:

  • A 2024 meta-analysis of 176 randomized trials and more than 20,000 participants found small average effects for smartphone mental health interventions broadly, including depression (g=0.28) and generalized anxiety (g=0.26). That review is useful context for digital interventions, but it is not chatbot-specific evidence. (Linardon et al., 2024)
  • A 2026 meta-analysis isolated 29 randomized trials of CBT-oriented psychological chatbots. It found moderate postintervention improvement in depressive symptoms (g=-0.55, 95% CI -0.70 to -0.40) and a small improvement in anxiety (g=-0.26, 95% CI -0.37 to -0.14). Both 95% prediction intervals crossed no effect, and evidence certainty was rated very low to low (Gong et al., 2026)
  • At follow-up in the same review, the pooled depression effect was smaller (g=-0.32, 95% CI -0.55 to -0.09), while the anxiety effect was not statistically significant (g=-0.19, 95% CI -0.43 to 0.04) (Gong et al., 2026)
  • A separate 2026 meta-analysis of 15 randomized trials and 1,737 participants found small-to-moderate effects for depressive symptoms, while generalized anxiety, stress, and positive affect were not significant after publication-bias adjustment (Hang et al., 2026)

Landmark RCT evidence:

  • Fitzpatrick et al. (2017): College students using Woebot for 2 weeks showed significantly greater depression score reduction than control group (self-help e-book) (F=6.47, p=.01)
  • Chaudhry et al. (2024): Patients with chronic diseases using Wysa showed significant reductions in depression and anxiety vs. no-intervention controls (P ≈ .004)

Critical limitations:

  • Engagement varies: Attrition, session completion, and acceptability should be reported for each intervention and comparator rather than assumed from a class-wide percentage
  • Comparison groups matter: Effect sizes smaller when compared to active controls vs. waitlist
  • Long-term data remain limited: Follow-up effects are smaller or uncertain in current meta-analyses
  • Crisis handling: AI chatbots may have difficulty with complex emotional nuance or crisis situations
  • Evidence maturity gap: A 2025 World Psychiatry systematic review found that only 16% of LLM-based chatbot studies underwent clinical efficacy testing, with 77% still in early validation phases. Only 47% of all chatbot studies focused on clinical efficacy, exposing critical gaps in therapeutic benefit validation (Hua et al., World Psychiatry, 2025). A parallel 2025 systematic review of 85 AI mental health studies spanning diagnosis, monitoring, and intervention found methodological quality weakest in the intervention domain, where only 38% of studies were rated good, and concluded that rigorous external validation is needed before deploying pre-trained models (Cruz-Gonzalez et al., 2025)

For OpenAI’s specialist benchmark of how chatbots answer in mental-health conversations, see MentalHealthBench in the evaluation chapter.

Adolescents and young adults: A 2025 systematic review and meta-analysis of 15 randomized trials involving 1,974 participants ages 12-25 found a moderate-to-large adjusted effect on depressive symptoms (Hedges g=0.61), especially in subclinical populations, but no significant adjusted effects for generalized anxiety, stress, affect, or mental well-being (Feng et al., 2025). The signal is narrower than marketing claims: conversational agents may help early depressive symptoms in some young people, but they are not validated as broad mental health substitutes.

Why users turn to AI in crisis: A December 2025 study surveying 53 individuals with lived crisis experience found people use AI chatbots to fill “in-between spaces” of human support, turning to AI when professional help is unavailable (waitlists, after-hours) or when they fear burdening family and friends (Ajmani et al., 2025, preprint). Interviews with 16 mental health experts in the same study emphasized that human connection remains essential during crisis. The authors propose designing AI as a “bridge towards human-human connection rather than an end in itself,” increasing user preparedness for positive action while de-escalating negative intent.

Regulatory status and product claims:

  • Wysa reported receiving FDA Breakthrough Device Designation in 2022. A designation is not marketing authorization, and the exact indication and status should be confirmed in an FDA record before clinical procurement
  • The evidence for a named structured therapeutic chatbot must not be transferred to a general-purpose LLM or to a different product
  • A 2026 Nature Medicine multi-turn audit of 9 general-purpose consumer chatbots found concerning behavior was common, context-dependent, and accumulated over conversation turns (Weilnhammer et al., 2026)

Appropriate vs. Inappropriate Use of Chatbot Therapy

Potentially appropriate after product-specific review:

  • Adjunct to traditional therapy (not replacement)
  • Access expansion for underserved populations
  • Psychoeducation and skill-building between sessions
  • Symptoms and population that match the evaluated intervention, with a defined escalation plan

Do not assume the chatbot is sufficient when the patient needs:

  • Replacement for human therapist
  • Evaluation or treatment of psychosis, severe depression, mania, or another complex presentation
  • Suicide or self-harm assessment and crisis response
  • Patients requiring medication management
  • Crisis intervention

Safety Concern: Chatbots and Suicidal Patients

In a simulated evaluation of 29 agents responding to suicidal-crisis scenarios, the authors concluded that the tested chatbots should be contraindicated for suicidal patients because of inconsistent crisis responses and validation patterns (Pichowicz et al., 2025). That conclusion applies to the tested systems and prompts, not every future system. A separate preprint found inconsistent responses to delusions and suicidal ideation and evidence of stigma toward some psychiatric conditions (Moore et al., 2025, preprint). Prompt studies identify safety failure modes; they do not demonstrate clinical benefit, harm rates, or causal patient outcomes.

Scale and evidence boundaries: A nationally representative survey found that 13.1% of U.S. adolescents and young adults had used generative AI for mental health advice, corresponding to an estimated 5.4 million people in that age group (McBain et al., 2025).

McBain covers adolescents and young adults. The adult U.S. complement is a Prolific stratified survey (August–October 2024; analytic n=1,871) in which 24% (95% CI 22–26%; n=449) reported using LLMs for mental health, mostly ChatGPT among users; users were younger, disproportionately male and Black, had higher psychopathology and treatment history, and more often reported access barriers (41% versus 21%), with stated uses including emotional support, learning therapy skills, and supplementing existing therapy (Stade et al., 2026). This is reported use, not efficacy or clinical benefit. Do not read 24% as national prevalence: the authors note that Prolific oversamples technology adopters (any-LLM use 92% versus Pew’s 23%), and their 14–18 million U.S. adult figure is a Pew-adjusted estimate (Stade et al., 2026). Several authors report developer or industry advising (Stade: paid advising to Sonar Mental Health, Sonia Health, and OpenAI; Eichstaedt: equity and paid advising from Jimini Health plus equity in Sonar Mental Health; Tait: previous Jimini Health employment).

OpenAI has published provider-estimated rates for conversations involving possible psychosis, mania, self-harm, or suicide, but those operational estimates have not been independently validated and cannot be interpreted as diagnoses or incident cases (OpenAI, 2025). A Danish electronic-health-record analysis describing 38 patients remains a preprint and is hypothesis-generating rather than causal evidence (Olsen et al., 2025, preprint).

A 2026 JAMA Psychiatry cross-sectional prompt study evaluated 158 unique psychosis-related prompts across three chatbot versions, producing 474 response pairs. It found that response appropriateness varied with prompt content and symptom type (Shen et al., 2026). The study supports concern about inappropriate responses to psychotic content, not a claim that chatbot use causes psychosis.

A 2026 Nature Medicine audit (SIM-VAIL) ran 810 multi-turn conversations across 9 consumer chatbots and scored 13 clinician-defined risk dimensions (Weilnhammer et al., 2026). Concerning behavior was common, highest for psychosis and mania, accumulated over turns, and could be lowered by rewriting an early high-concern reply; the authors call the pattern a vulnerability-amplifying interaction loop. This is an adversarial API-level simulation, not a real-world harm-rate estimate, it does not evaluate Woebot or Wysa, and two authors report Microsoft relationships.

Novel lawsuits in 2024-2025 have alleged AI chatbots encouraged minors’ suicides, including the case of 14-year-old Sewell Setzer III, who died by suicide in February 2024 after extensive interaction with a Character.ai chatbot (NBC News, October 2024). These cases raise liability concerns for platforms deploying these tools without adequate safeguards.

Platform safeguards vary. Some providers describe classifiers, response-routing systems, crisis-resource banners, or parental controls (Anthropic, 2025; Meta, 2026). Their existence does not establish sensitivity, specificity, equitable performance, or improved patient outcomes. A person in crisis requires access to trained human support and local emergency pathways rather than reliance on a general-purpose chatbot.

A 2026 npj Digital Medicine scoping review mapped 119 harm-focused papers on LLM-based chatbots into five clusters: conceptual work (32), mental-health support use (29), cognitive overreliance (15), AI dependence (36), and AI psychosis (7) (Diel et al., 2026). Most crisis-response research is vignette or benchmark work (18/29 in that cluster); dependence research is mostly correlational; AI-psychosis writing is almost entirely conceptual (6/7). The review is harm-scoped by design, included preprints, and did not assess study quality; it is not a risk–benefit verdict and does not show that chatbots mostly harm.

NLP for Clinical Documentation

Natural language processing can:

  • Extract symptoms from clinical notes
  • Identify patients not receiving guideline-concordant care
  • Support quality improvement
  • Auto-complete portions of psychiatric evaluations

This is documentation support, not diagnostic AI. Performance must be evaluated against the exact extraction, summarization, coding, or drafting task, and any generated text still requires review before entering the clinical record.

Treatment Response Prediction

Emerging research attempts to predict:

  • Which patients will respond to specific antidepressants
  • Optimal medication selection based on clinical and genetic factors
  • Likelihood of treatment-resistant depression

Current status: Many treatment-response models remain in the research phase. A 2026 randomized trial evaluated a web-based evidence and preference decision-support tool, showing that carefully designed support can be tested prospectively without implying autonomous diagnosis or a class-wide effect for psychiatric AI.

PETRUSHKA Trial (JAMA, March 2026):

The PETRUSHKA (Personalising antidEpressant Treatment foR Unipolar depreSsion combining individual cHoices, risKs and big datA) trial randomized 540 adults at 47 sites in Brazil, Canada, and the UK. Of 520 eligible participants, 493 were included in the primary analysis (Cipriani et al., 2026).

The web-based clinical decision-support tool combined clinical and demographic predictors with patient preferences to personalize antidepressant selection. The trial evaluates that complete evidence-and-preference workflow. It should not be generalized to autonomous prescribing or described as proof that any model using psychiatric data will improve outcomes.

Key findings:

  • At 8 weeks, 41 of 241 participants (17%) assigned to PETRUSHKA discontinued the prescribed antidepressant for any reason, compared with 69 of 252 (27%) receiving usual care (adjusted relative risk, 0.62; 95% CI, 0.44 to 0.88; P=.007)
  • Discontinuation due to adverse events at 8 weeks occurred in 22 of 241 participants (9%) versus 39 of 252 (16%) (adjusted relative risk, 0.59; 95% CI, 0.36 to 0.97; P=.04)
  • At 24 weeks, mean PHQ-9 scores were 7.1 versus 9.2 (adjusted mean difference, -1.92; 95% CI, -3.06 to -0.78), and mean GAD-7 scores were 4.6 versus 5.8 (adjusted mean difference, -1.39; 95% CI, -2.26 to -0.52)

PETRUSHKA is important because it links a decision-support intervention to prespecified comparative endpoints rather than reporting retrospective model accuracy alone. The open-label design and substantial missing data, particularly for later symptom outcomes, limit certainty. Replication, implementation evaluation, longer follow-up, and cost-effectiveness evidence are needed before broad adoption.

Separately, Walsh et al. fitted L1-regularized EHR models of treatment-resistant depression at first antidepressant prescription across Mass General Brigham, Vanderbilt University Medical Center, and Geisinger Clinic (2004–2022) and found they did not transport (Walsh et al., 2026). Internal C-statistics were 0.51–0.65 and external C-statistics 0.50–0.58 (AUPRC 0.07–0.15 internal; 0.07–0.11 external); the authors attribute the failure to TRD outcome heterogeneity, EHR source-data bias, and site-level practice variation. This is a transportability-failure result for structured EHR risk models of TRD, not an evaluation of PETRUSHKA and not a warrant to deploy a TRD score.


Part 4: Ethical Considerations

Unique Challenges in Psychiatric AI

  1. Capacity is decision-specific: Some patients may have impaired capacity during an acute episode, but a psychiatric diagnosis alone does not establish incapacity
  2. Stigma: AI-generated psychiatric labels may follow patients permanently
  3. Coercion risk: Predictive algorithms could be used for involuntary commitment
  4. Privacy: Mental health information is especially sensitive
  5. Equity: Psychiatric presentations vary by culture, language, and socioeconomic status

What Clinicians Should Do

  1. Do not substitute a score for assessment: A risk model may inform a workflow only when its population, threshold, response pathway, and limitations are understood
  2. Match evidence to the claim: Prediction performance, workflow effects, symptom outcomes, and suicide prevention are different endpoints
  3. Prioritize the therapeutic relationship: AI cannot replace human connection in mental health care
  4. Advocate for appropriate use: Support research while opposing premature clinical deployment

A 2025 comprehensive review in Current Psychiatry Reports concluded that current evidence only supports AI as a complement to clinical expertise, not a replacement (Jalali et al., Curr Psychiatry Rep, 2025). Similarly, the evolving field of digital mental health requires careful navigation between smartphone apps, generative AI, and virtual reality applications, each with distinct evidence bases and implementation challenges (Torous et al., World Psychiatry, 2025).


Professional Society Guidelines on Psychiatric AI

American Psychiatric Association (APA)

The APA Board of Trustees approved a Position Statement on the Role of Artificial Intelligence in Psychiatry in March 2024. The statement acknowledges both opportunities and significant risks.

APA Position Statement on AI (March 2024)

Opportunities identified:

  • Clinical documentation assistance
  • Care plan suggestions and lifestyle modifications
  • Identification of potential diagnoses and risks from medical records
  • Automation of billing and prior authorization
  • Detection of potential medical errors or systemic quality issues

Risks and concerns:

  • Unacceptable risks of biased or substandard care
  • Violations of privacy and informed consent
  • Lack of oversight and accountability for AI-driven clinical decisions

Key guidance for physicians:

  1. Approach AI technologies with caution, particularly regarding potential biases or inaccuracies
  2. Ensure HIPAA compliance in all uses of AI
  3. Take an active role in oversight of AI-driven clinical decision support
  4. View AI as a tool intended to augment, not replace, clinical decision-making
  5. Remain skeptical of AI output in clinical practice
  6. Preserve meaningful professional oversight and accountability for AI-supported care

For current guidance: APA Position Statement on the Role of Augmented Intelligence in Clinical Practice and Research (PDF)

The position statement supports responsible augmentation while emphasizing bias, privacy, consent, transparency, and oversight. Those principles should be applied to the intended use and workflow of a specific system rather than treated as evidence that a product works.

Critical note: A professional-society statement about responsible AI does not constitute endorsement of autonomous psychiatric diagnosis or a specific suicide prediction product.

APA Resource Document on AI in Psychiatric Practice (July 2026)

The 2024 APA Augmented Intelligence position statement on this page remains the short normative stance (augment, do not replace). In July 2026, APA’s Joint Reference Committee approved a longer Resource Document on Artificial Intelligence in Psychiatric Practice that turns that stance into implementation guidance across documentation, ethics and privacy, scribes, regulation, and clinician competencies (APA Resource Document PDF; APA RD page).

The Resource Document states that there are currently no FDA-approved AI applications for psychiatric diagnosis, treatment, or clinical decision-making. AI outputs are framed as advisory supports for workflow and information tasks; diagnostic and treatment responsibility remains with the clinician. The document includes a practical chapter on AI scribe technologies and urges psychiatrists to ask patients about their own AI use (platforms, frequency, purpose, and any temporal link to symptom change), while noting that so-called “AI psychosis” is an emerging observation rather than an established diagnosis. Named products in the Resource Document are illustrative, not APA endorsements.

Regulatory status can change after July 2026; check current FDA listings before relying on it.

American Medical Association (AMA)

The AMA’s Principles for Augmented Intelligence in Health Care apply to all medical specialties, including psychiatry. The AMA prefers the term “augmented intelligence” to emphasize that these tools should assist, not replace, physician judgment.

AMA Augmented Intelligence Principles
  1. Augmentation over automation: AI should enhance physician decision-making, not replace it
  2. Transparency: Development, validation, and deployment processes must be transparent
  3. Human oversight: Clinical users and organizations need authority to question, override, or stop an AI-supported workflow
  4. Privacy protection: Patient data must be protected throughout AI development and use
  5. Bias mitigation: Algorithmic bias must be identified and addressed
  6. Ongoing monitoring: Continuous performance evaluation required after deployment

For current guidance: AMA Augmented Intelligence in Medicine

American Academy of Child and Adolescent Psychiatry (AACAP)

AACAP has engaged with AI through educational programming and research published in the Journal of the American Academy of Child and Adolescent Psychiatry. The following are pediatric adoption questions for clinicians and institutions, not a representation of an AACAP position statement:

Pediatric-specific concerns:

  • Consent complexity: Minors require parental consent, but adolescents may resist disclosure
  • Developmental considerations: AI may misinterpret normal adolescent behavior as pathological
  • School-based screening: Profound privacy concerns when AI is used in educational settings
  • Evidence requirements: Higher bar for interventions in developing brains
  • Access and age controls: Consumer terms, parental controls, and age-assurance systems vary by provider and change over time. They should be verified directly rather than assumed. A consumer chatbot should not be used as the sole response to a minor experiencing a mental health crisis

A 2024 study in JAACAP compared ChatGPT versions on suicide risk assessment in youth and found that AI estimated higher risk than psychiatrists, particularly in severe cases, which could lead to inappropriate treatment recommendations (Nguyen et al., JAACAP, 2024).

FDA Regulatory Status

FDA-authorized computerized behavioral therapy devices, not necessarily AI:

Device Indication boundary FDA pathway
Rejoyn Adjunct to clinician-managed outpatient care for adults aged 22 years or older with major depressive disorder who are taking antidepressant medication 510(k) cleared March 30, 2024, K231209
reSET Prescription computerized behavioral therapy adjunct for defined adults enrolled in clinician-supervised outpatient substance-use treatment De Novo authorized September 14, 2017, DEN160018
Somryst Prescription digital therapeutic for chronic insomnia within its labeled population and workflow 510(k) cleared, K191716

Critical findings:

  • The records above do not authorize autonomous suicide prediction or autonomous psychiatric diagnosis
  • Each device has a defined intended use, population, prescription status, and clinical context
  • A digital therapeutic is not automatically an AI system, and FDA authorization does not establish current commercial availability.
  • Product ownership and availability can change even while an authorization record remains in the FDA database

Regulatory Boundary

FDA oversight depends on a software function’s intended use and whether it meets the definition of a medical device. A general-purpose chatbot is not transformed into an authorized psychiatric device because a patient or clinician uses it for mental health conversation. Conversely, a developer that makes device claims must evaluate the applicable regulatory pathway. The correct question is not whether FDA review is optional, but whether the specific function falls within FDA jurisdiction and what its authorization actually covers.

State Regulatory Developments (2025)

Three U.S. states enacted laws in 2025 restricting AI in mental health care:

State Law Effective Key Provisions
Utah Utah Code, Title 13, Chapter 72a May 7, 2025 Establishes disclosure and other requirements for regulated mental health chatbots within the statute’s definitions
Nevada AB 406 July 1, 2025 Restricts AI systems specifically programmed to provide professional mental or behavioral health care, with defined exceptions and civil enforcement
Illinois Wellness and Oversight for Psychological Resources Act, Public Act 104-0054 August 1, 2025 Limits unlicensed AI-delivered therapy and certain uses by licensed professionals, including emotion or mental-state inference in therapeutic communications

These laws have different definitions, exceptions, duties, and enforcement mechanisms. A summary table is not a substitute for reading the controlling statute and obtaining jurisdiction-specific legal review. Illinois, for example, addresses both unlicensed therapy and specified uses of AI by licensed professionals. Nevada defines prohibited systems and exceptions differently. Utah regulates mental health chatbots through its own disclosure and conduct provisions.

Federal activity: A December 2025 executive order directed federal agencies to identify and challenge certain state AI laws (White House, 2025). The order does not by itself erase state statutes. The practical effect depends on later federal action, litigation, statutory authority, and the scope of any judicial decision.

FDA activity: The FDA’s Digital Health Advisory Committee convened in November 2025 to discuss generative AI-enabled mental health devices, focusing on a hypothetical LLM therapy chatbot for major depressive disorder. The committee emphasized concerns about confabulation, bias, and the need for human oversight during crisis situations, but issued no binding regulatory changes (FDA DHAC, November 2025).

Professional Consensus on Suicide Prediction

The guidelines cited in this chapter do not support using a suicide-risk score as a substitute for psychosocial assessment or as the sole basis for treatment allocation. The concerns include:

  1. Base rate constraint: Low prevalence can produce low PPV even when other performance measures appear favorable
  2. Threshold consequences: False positives and false negatives carry different harms depending on the downstream intervention
  3. Dangerous false reassurance: “Low risk” predictions discourage proper clinical assessment
  4. Ethical concerns: Predictive algorithms could be misused for involuntary commitment

Absence of a society endorsement is not itself a formal position. The direct evidence is the content of the applicable guideline or statement. NICE explicitly advises against using risk tools to predict future suicide or repetition of self-harm and against using global risk stratification to determine discharge or treatment.


Check Your Understanding

The following scenarios are fictional teaching exercises. They do not describe actual patients, product outputs, legal holdings, or measured outcomes.

Scenario 1: The Suicide Risk Algorithm

Clinical situation: Your hospital deployed a suicide risk prediction algorithm. You’re seeing a 32-year-old woman in primary care for diabetes follow-up. She mentions feeling “a bit down lately.” The algorithm has flagged her as “LOW RISK” (10th percentile).

Question: Do you skip detailed depression and suicidality screening because the algorithm says low risk?

Answer: No. A low-risk output does not remove the need to assess the patient’s current presentation.

Reasoning:

  1. Outcome prevalence matters: The PPV of a score depends on the population, prediction horizon, and defined outcome
  2. A low-risk category is not a clinical disposition: It cannot establish that current suicidal thoughts, intent, or acute stressors are absent
  3. Data recency matters: Historical EHR data may not capture events that occurred today
  4. Clinical presentation matters: Patient saying “I’m feeling down” requires exploration

What you should do:

  1. Follow the indicated clinical assessment pathway: Evaluate current symptoms and ask directly about suicidal thoughts when clinically indicated
  2. Interpret the score within its validated use: Confirm population, outcome, horizon, threshold, missing-data handling, and intended response
  3. Document the clinical reasoning: Record the information considered and the assessment performed, without implying that a model ruled suicide risk in or out
  4. Report workflow hazards: Escalate evidence of alert fatigue, false reassurance, inequitable performance, or an undefined response pathway through institutional governance

Bottom line: A risk score cannot replace a current clinical assessment or determine disposition by itself. If a health system deploys a model, it should prospectively evaluate the complete response workflow and monitor both benefit and harm.

Scenario 2: The Therapy Chatbot Recommendation

Clinical situation: A 24-year-old graduate student with mild-to-moderate depressive symptoms asks about using a structured digital CBT program while awaiting an appointment. She has limited time and funds and is on a three-month waitlist.

Question: Is it appropriate to recommend a therapy chatbot?

Answer: It can be considered as an adjunct or bridge only after reviewing the specific product, evidence, availability, privacy terms, and patient context.

When chatbot therapy may be appropriate:

  • The patient’s symptoms and age match the evaluated population and product label
  • Current assessment does not indicate a need for urgent or higher-intensity care
  • The patient understands what the program can and cannot do
  • As bridge to traditional therapy, not permanent replacement
  • Patient understands limitations

What to tell the patient:

  1. The evidence applies to the evaluated structured intervention, population, comparator, and follow-up period, not to every chatbot
  2. Short-term average symptom improvements do not guarantee individual benefit, and current meta-analytic certainty is low
  3. The program is not a crisis service. The patient needs a clear route to local crisis and emergency support
  4. The follow-up plan should not depend on whether the patient continues using the digital program

What to document:

  • Discussion of chatbot as adjunct/bridge
  • Assessment of suicidality (negative)
  • Crisis plan provided
  • Plan to continue pursuing traditional therapy

Presentations requiring direct clinical evaluation rather than chatbot reliance:

  • Active suicidal ideation
  • Severe or rapidly worsening symptoms
  • Psychotic symptoms
  • Substance use requiring treatment
  • A history or presentation that requires an individualized safety plan and closer follow-up
Scenario 3: The Digital Phenotyping Research Study

Clinical situation: A research coordinator approaches you about enrolling your patients in a study that monitors smartphone use to predict depressive episodes. The study promises to alert clinicians when patients show “digital biomarkers of depression.”

Question: What concerns should you raise before enrolling patients?

Answer: Multiple ethical and practical concerns require clarification.

Questions to ask the research team:

  1. Informed consent: How is continuous monitoring explained to patients? Can they withdraw consent without penalty?
  2. Data ownership: Who owns the collected data? Can it be sold or shared?
  3. Privacy: What if the data is breached? What if it reveals sensitive information (location of therapy visits, AA meetings)?
  4. Clinical utility: What is the evidence that detecting “digital biomarkers” improves outcomes?
  5. Alert response: Who receives an alert, during what hours, under what protocol, and with what escalation and documentation duties?
  6. Bias testing: Has the algorithm been validated in diverse populations?
  7. Patient burden: Will patients feel surveilled? Will this affect their relationship with their smartphone?

Concerns about implementation:

  • Alerts without validated interventions create burden without benefit
  • Continuous monitoring may worsen anxiety in some patients
  • Research ≠ clinical utility: Patterns detected in studies may not generalize
  • Alert fatigue if algorithm has high false positive rate

Preservation-first response: Enrollment should wait until the protocol, consent process, data flow, security controls, alert workflow, and evidence for the proposed clinical action have been reviewed. Research participation must not be represented as established clinical benefit.

Scenario 4: The AI-Induced Psychosis Case

Clinical situation: A 19-year-old college student presents with new-onset paranoid ideation. During the interview, he describes spending 6-8 hours daily conversing with an AI chatbot (character.ai style), which he believes has developed consciousness and is communicating with him through “hidden messages.” He stopped attending classes, believing the AI is teaching him things his professors cannot.

Question: How do you assess the role of AI in this presentation?

Answer: The temporal association warrants assessment, but it does not establish that the chatbot caused the psychosis.

Assessment approach:

  1. Standard psychiatric evaluation: Rule out organic causes (substance use, medical conditions), assess for primary psychotic disorder
  2. Technology history: Duration, intensity, content of AI interactions; isolation from in-person relationships
  3. Pre-existing vulnerabilities: Family history of psychosis, prodromal symptoms before AI use
  4. Reality testing: Does patient recognize AI is not conscious? Can he distinguish AI output from personal beliefs?

What the literature suggests:

  • A published case report described psychosis arising in temporal association with intensive ChatGPT interaction. A case report can identify a plausible safety signal, but cannot estimate incidence or establish causation (Pierre et al., 2025)
  • A Danish electronic-health-record review described 38 patients with potentially harmful chatbot-related experiences, most often involving reinforcement or consolidation of delusional content. It remains a preprint and should be treated as hypothesis-generating (Olsen et al., 2025, preprint)
  • A cross-sectional prompt study found variable appropriateness across 474 chatbot response pairs to psychosis-related prompts. It evaluates responses, not whether chatbot exposure caused illness (Shen et al., 2026)
  • Provider-published operational estimates describe conversations that classifiers judged as potentially involving severe mental-health concerns. They are not diagnoses, incidence estimates, or causal measurements (OpenAI, 2025)
  • The clinically relevant question is whether the interaction is reinforcing symptoms, disrupting sleep, increasing isolation, or delaying contact with human care
  • A 2026 harm-focused scoping review counted 7 AI-psychosis papers, 6 of them conceptual, which supports asking about chatbot use without treating exposure as a proven cause (Diel et al., 2026)

Management:

  • Standard treatment for first-episode psychosis
  • Technology boundaries as part of treatment plan
  • Family education about AI interaction patterns
  • Do not blame AI for the illness, but address its role in symptom maintenance

Key insight: Current evidence supports asking about chatbot use and evaluating whether responses reinforce delusional content; it does not establish that chatbot use causes psychosis or define a validated exposure threshold. Technology history can be included alongside sleep, substance use, medical causes, mood symptoms, trauma, family history, functional decline, and the usual assessment of first-episode psychosis.


Key Takeaways

Clinical Bottom Line for Psychiatric AI
  1. A suicide-risk score cannot replace clinical assessment. Prediction metrics do not show that displaying a score prevents suicide, and a low-risk output cannot rule out current need.

  2. Structured CBT-oriented chatbots have evidence of short-term average symptom improvement. Certainty is low, interventions are heterogeneous, and general-purpose LLMs are not equivalent to the evaluated products.

  3. Digital phenotyping evidence is mainly observational and feasibility-focused. Clinical utility, transportability, privacy, and response pathways require separate evaluation.

  4. FDA records define product-specific intended uses. The cited computerized behavioral therapies are not necessarily AI and do not authorize autonomous diagnosis or suicide prediction.

  5. Professional societies require AI to augment, not replace, clinical judgment. The APA and AMA are explicit on this point.

  6. AI cannot replace the therapeutic relationship. Human connection remains essential to psychiatric care.

  7. Ask about AI chatbot use when it is clinically relevant. Evaluate content, intensity, sleep disruption, isolation, symptom reinforcement, and whether chatbot use is delaying human care without treating the exposure as a proven cause. Diel et al. 2026 map chatbot-harm literature; the AI-psychosis cluster remains conceptual.

  8. Match claims to endpoints. Retrospective accuracy, prompt behavior, symptom change, workflow performance, and patient outcomes answer different questions.

Questions About AI in Psychiatry

Do suicide prediction algorithms work?

They can stratify risk in defined populations, but low positive predictive value and the absence of randomized outcome evidence limit clinical utility. NICE advises against using risk scales to predict suicide or determine treatment. No model should replace a current clinical assessment.

What happened to Facebook’s suicide prevention AI?

Meta has described machine-learning systems that identify potential suicide or self-injury content and route cases for review or support. The company continues to describe related safeguards, but peer-reviewed evidence that the program reduces suicide deaths has not been published.

Is chatbot therapy evidence-based?

Randomized trials and meta-analyses support short-term symptom improvement for some structured CBT-oriented chatbots, but certainty is low, prediction intervals include no effect, and longer-term anxiety benefit is uncertain. General-purpose LLM chatbots should not be assumed equivalent to evaluated therapeutic programs.

Why is psychiatric AI difficult to validate?

Many psychiatric diagnoses have heterogeneous presentations and no single confirmatory biomarker. Labels, outcomes, culture, time horizon, and clinical context therefore require unusually careful definition, external validation, and prospective evaluation.