1. Heinz and colleagues evaluated a generative AI chatbot for adults with clinically significant depression, anxiety, or elevated risk for eating disorders.
2. Users of the chatbot experienced greater symptom reductions compared to controls on the waitlist and reported higher engagement.
Evidence Rating Level: 1 (Excellent)
Study Rundown: Despite the adverse impact of mental health disorders, mental health infrastructure is inadequately resourced to meet the current and growing demand for care. Heinz and colleagues evaluated Therabot, a generative AI chatbot’s effectiveness in treating symptoms of major depressive disorder (MDD), generalized anxiety disorder (GAD), and clinically high-risk feeding and eating disorder (CHR-FED). Participants received either four weeks of Therabot access with daily prompts or were assigned to a control group in which they remained on the waitlist. After four weeks, participants in the Therabot group were not prompted, but were permitted access to Therabot. Primary outcomes were symptom scores measured using validated depression, anxiety, and weight-concern instruments. Secondary outcomes included engagement, satisfaction, and therapeutic alliance with Therabot. Compared with controls, Therabot users experienced significantly greater reductions across all three symptom domains at four and eight weeks of usage. Therapeutic alliance ratings for Therabot were also comparable to published ratings for human therapists. This study demonstrated that an AI chatbot has promise as a scalable mental health intervention.
Click here to read the study in NEJM AI
Relevant Reading: Integrating Artificial Intelligence into Psychological Counseling: A Narrative Review and Governance Framework
In-Depth [randomized controlled trial]: 210 adults across the United States were recruited for this trial. Eligible participants were at least 18 years old and screened positive for clinically significant MDD, GAD, or CHR-FED. Active suicidality, mania, and psychosis were excluded. Participants were stratified by symptom group and randomized to Therabot (n=106) or waitlist control (n=104). Text-based, primarily cognitive behavioral therapy-informed dialogue was delivered by Therabot, and all chatbot responses were monitored by researchers to prevent unsafe or inappropriate content. Primary outcomes were changes in Patient Health Questionnaire-9 (PHQ-9), GAD Questionnaire-IV, and Weight Concerns Scale (WCS) scores at four and eight weeks. Secondary outcomes included therapeutic alliance, engagement with Therabot, and satisfaction with Therabot. Therabot produced greater symptom reductions than control for MDD at four weeks (-6.13 vs. -2.63; Cohen’s d=0.845) and eight weeks (-7.93 vs. -4.22; d=0.903), GAD at four weeks (-2.32 vs. -0.13; d=0.840) and eight weeks (-3.18 vs. -1.11; d=0.794), and CHR-FED at four weeks (-9.83 vs. -1.66; d=0.819) and eight weeks (-10.23 vs. -3.70; d=0.627). Favorable user satisfaction and therapeutic alliance ratings were found, with users rated Therabot as easy to use and similar to a real therapist. This study was limited by the need for human intervention to ensure no unsafe content and only four weeks of post-intervention follow-up. Overall, this study demonstrated that fine-tuned Gen-AI chatbots offer a feasible approach to delivering personalized mental health interventions on a scale.
Image: PD
©2026 2 Minute Medicine, Inc. All rights reserved. No works may be reproduced without expressed written consent from 2 Minute Medicine, Inc. Inquire about licensing here. No article should be construed as medical advice and is not intended as such by the authors or by 2 Minute Medicine, Inc.




