Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Research Protocols

Date Submitted: Jan 18, 2026
Date Accepted: Jul 17, 2026

The final, peer-reviewed published version of this preprint can be found here:

Frameworks, Methodologies, and Tools for Evaluating Large Language Models in Digital Mental Health Interventions: Protocol for a Scoping Review

Salinas-Layana A, JimĂ©nez-Molina Ă, Lira D, VĂ©liz-Montoya FA, Venegas A, Muñoz N, ChandĂ­a M, Rojas R, Martinez V

Frameworks, Methodologies, and Tools for Evaluating Large Language Models in Digital Mental Health Interventions: Protocol for a Scoping Review

JMIR Res Protoc 2026;15:e91677

DOI: 10.2196/91677

PMID: 42735424

Frameworks, methodologies, and tools for evaluating large language models (LLMs) in digital mental health interventions: Protocol for a scoping review

  • Antonio Salinas-Layana; 
  • Álvaro JimĂ©nez-Molina; 
  • Daniela Lira; 
  • FĂ©lix A. VĂ©liz-Montoya; 
  • Alexi Venegas; 
  • NicolĂĄs Muñoz; 
  • Mario ChandĂ­a; 
  • Rigoberto Rojas; 
  • Vania Martinez

ABSTRACT

Background:

Digital mental health interventions (DMHIs) can help close persistent gaps in access to assessment, prevention, and treatment. Recent advances in generative artificial intelligence (AI), particularly large language models (LLMs), further expand this promise by enabling natural language understanding, personalization, and empathic interaction across assessment, support, and therapeutic contexts. However, significant evaluation challenges persist, including a lack of standardized constructs and validated instruments, which limit the comparability, reproducibility, and generalizability of findings. No systematic synthesis currently documents the frameworks, methodologies, and tools used to evaluate LLMs in DMHIs, thereby hampering the development of a comprehensive evidence base to guide future evaluation efforts.

Objective:

This scoping review aims to systematically map and synthesize the available evidence on frameworks, methodologies, and tools used to evaluate LLMs applied to digital mental health interventions. Specifically, it aims to identify the constructs assessed, the instruments employed, and the evaluation procedures and stages addressed.

Methods:

Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews (PRISMA-ScR) guidelines, this scoping review will search five electronic databases: PubMed, Scopus, Web of Science, IEEE Xplore, and ACM Digital Library, from January 2019 to September 15, 2025. Eligibility criteria will encompass both empirical studies and theoretical proposals evaluating LLMs embedded within DMHIs. Studies limited to risk detection or decision-support systems without an intervention component will be excluded. Data extraction will capture information on conceptual frameworks, methodological designs, evaluation procedures and tools, measured constructs, and other relevant contextual information.

Results:

The systematic search was conducted between September 1 and 15, 2025, yielding 4,273 records across the five databases. After duplicate removal, 2,980 unique records remained for screening. A pilot screening exercise involving four independent reviewers achieved high inter-rater reliability (free-marginal Randolph kappa = 0.81), with 76% unanimous agreement, indicating adequate calibration of selection criteria. These figures are interim process indicators rather than final review findings, as title and abstract screening of the remaining records is currently underway.

Conclusions:

This scoping review is expected to provide one of the first systematic syntheses of frameworks, methodologies, and tools used to evaluate LLMs in digital mental health interventions. By identifying prevailing patterns and gaps, the resulting evidence map is intended to serve as a practical reference for researchers, developers, and policymakers working toward a scientifically grounded, safe, ethical, and effective deployment of LLMs in mental health interventions. Clinical Trial: Open Science Framework (OSF); DOI: 10.17605/OSF.IO/PAZYQ


 Citation

Please cite as:

Salinas-Layana A, JimĂ©nez-Molina Ă, Lira D, VĂ©liz-Montoya FA, Venegas A, Muñoz N, ChandĂ­a M, Rojas R, Martinez V

Frameworks, Methodologies, and Tools for Evaluating Large Language Models in Digital Mental Health Interventions: Protocol for a Scoping Review

JMIR Res Protoc 2026;15:e91677

DOI: 10.2196/91677

PMID: 42735424

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.