Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR AI

Date Submitted: Nov 19, 2025
Date Accepted: Jun 22, 2026

The final, peer-reviewed published version of this preprint can be found here:

Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

Bleuze C, Fort K, Martin VP, Névéol A

Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

JMIR AI 2026;5:e88082

DOI: 10.2196/88082

PMID: 42594358

Large Language Models for Mental Health Prediction: A Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

  • Clémentine Bleuze; 
  • Karën Fort; 
  • Vincent P. Martin; 
  • Aurélie Névéol

ABSTRACT

Background:

Natural Language Processing methods have the potential to support clinical practice and research. Notably, Large Language Models (LLMs) provide a new paradigm for clinical text analysis and production with many envisioned applications for mental health. However, LLMs are prone to bias, and we lack studies to validate their clinical utility.

Objective:

The aim of this study is to review current research making use of LLMs for mental health prediction, with a focus on bias and clinical utility.

Methods:

This review follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and is registered with OSF (Open Science Framework). The search was conducted using five scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library and ACL Anthology) with queries associating keywords related to ‘Mental Health’, ‘Large Language Models’ and ‘Prediction’. Included papers were published between 2019 and 2024. The exclusion criteria filtered out protocols, reviews, and articles in languages other than English.

Results:

A total of 2,472 articles were identified, of which 263 (10.6%) were assessed for eligibility and 201 (8.1%) were included in the review. Articles address 15 mental health disorder groups, but largely focus on Depressive Disorders. Due to stronger confidentiality restrictions on clinical data, studies mostly rely on social media data in English. Bias and clinical utility of approaches are discussed in 164 (81.8%) of selected papers, however most papers offer a data and model driven approach of bias. The presence of clinicians among authors was associated with significantly higher conditions diversity (P<.05) and lower usage of social media data (P<.001).

Conclusions:

Bias and clinical utility are lightly covered in LLM-based mental health prediction research. In-depth approaches involving interdisciplinary teams of clinicians and NLP specialists are needed to ensure technical soundness, clinical validation and utility.


 Citation

Please cite as:

Bleuze C, Fort K, Martin VP, Névéol A

Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

JMIR AI 2026;5:e88082

DOI: 10.2196/88082

PMID: 42594358

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.