Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.
Who will be affected?
Readers: No access to all 28 journals. We recommend accessing our articles via PubMed Central
Authors: No access to the submission form or your user account.
Reviewers: No access to your user account. Please download manuscripts you are reviewing for offline reading before Wednesday, July 01, 2020 at 7:00 PM.
Editors: No access to your user account to assign reviewers or make decisions.
Copyeditors: No access to user account. Please download manuscripts you are copyediting before Wednesday, July 01, 2020 at 7:00 PM.
Kamaleddin MA, Mirjalili M, Barzegar R, Trung Le N, Cote Z, Winkler O, Burback L, Zhang Y, Zaiane O, Monson C, Sharma D, Krishnan S, Greenshaw A, Zeifman R, Selby P, Bhat V
Multiagent Large Language Model Framework for Psychotherapy Fidelity Assessment in Motivational Interviewing and Cognitive Behavioral Therapy Training: Cross-Sectional, Simulation-Based Evaluation Study
Multi-Agent Large Language Model Framework for Psychotherapy Fidelity Assessment in Motivational Interviewing and Cognitive Behavioral Therapy Training: A Cross-Sectional, Simulation-Based Evaluation
Mohammad Amin Kamaleddin;
Mina Mirjalili;
Reza Barzegar;
Nghia Trung Le;
Zachary Cote;
Olga Winkler;
Lisa Burback;
Yanbo Zhang;
Osmar Zaiane;
Candice Monson;
Divya Sharma;
Sri Krishnan;
Andrew Greenshaw;
Richard Zeifman;
Peter Selby;
Venkat Bhat
ABSTRACT
Background:
Psychotherapeutic interventions such as motivational interviewing (MI) and cognitive behavioral therapy (CBT) are commonly delivered alongside pharmacotherapy and incorporated into clinical trials in neuropsychiatric care. However, substantial variability in therapist skill and treatment fidelity across clinicians and sites is difficult to measure at scale, introducing uncontrolled behavioral variance and limiting reproducibility.
Objective:
To develop and validate a scalable, automated platform for conducting and evaluating MI and CBT sessions with the aim of quantifying and reducing therapist-related behavioral variability.
Methods:
We developed a multi-agent large language model (LLM) platform that conducts MI or CBT sessions and evaluates session transcripts using established fidelity instruments, generating criterion-referenced narrative feedback. Discriminative validity was tested using scripted novice and expert therapist profiles across 133 MI and 103 CBT sessions. External validity was assessed using expert-labeled MI transcripts. Clinical reliability was examined by comparing model-generated scores with consensus ratings from practicing therapists. We additionally evaluated whether incorporation of model-generated feedback improved subsequent therapist performance.
Results:
The evaluator reliably distinguished expert from novice performance across all MI and CBT criteria. Mean fidelity scores were significantly higher for expert than novice profiles for both MI (4.38 vs 1.13) and CBT (5.23 vs 1.48). On externally labeled MI transcripts, the model identified low-quality sessions with 91.7% accuracy. Agreement between model scores and therapist consensus ratings was high (ICC(2,1)=0.926 for MI; 0.900 for CBT). Incorporation of evaluator feedback led to significant improvements in novice performance across most criteria in subsequent sessions.
Conclusions:
This multi-agent LLM platform provides a scalable translational tool for objectively quantifying psychotherapeutic fidelity and reducing behavioral variability in neuropsychiatric training and research, with implications for improving reproducibility and treatment quality.
Citation
Please cite as:
Kamaleddin MA, Mirjalili M, Barzegar R, Trung Le N, Cote Z, Winkler O, Burback L, Zhang Y, Zaiane O, Monson C, Sharma D, Krishnan S, Greenshaw A, Zeifman R, Selby P, Bhat V
Multiagent Large Language Model Framework for Psychotherapy Fidelity Assessment in Motivational Interviewing and Cognitive Behavioral Therapy Training: Cross-Sectional, Simulation-Based Evaluation Study