Accepted for/Published in: JMIR Mental Health
Date Submitted: Apr 21, 2026
Date Accepted: Sep 17, 2026
Ensemble of Domain-Specific Natural Language Processing and Large Language Models for Detecting Suicidal Ideation and Mental Health Conditions in Social Media Text: Development and Evaluation Study
ABSTRACT
Background:
Background:
Mental illness contributes substantially to global disability, and public adoption of artificial intelligence for mental health support is accelerating without commensurate safety evaluation. General-purpose large language models frequently miss genuine cases of suicidal ideation in naturalistic text, a failure mode with immediate clinical consequences. Domain-specific natural language processing methods, grounded in diagnostic criteria and validated symptom vocabularies, offer a contrasting approach, but the performance and safety characteristics of ensembles that integrate the two have not been formally quantified.
Objective:
Objective:
This study aimed to (1) benchmark five modelling approaches for classifying mental health conditions in social media text; (2) develop and optimise ensemble architectures that integrate a domain-specific natural language processing classifier with a fine-tuned large language model; and (3) derive a closed-form mathematical framework that constrains ensemble weighting in safety-critical classification.
Methods:
Methods:
We evaluated 79,160 mental health-related social media texts spanning nine categories (anxiety, bipolar disorder, depression, a normal baseline, personality disorder, stress, suicidal ideation, attention-deficit/hyperactivity disorder, and autism spectrum disorder), drawn from publicly available Reddit and Twitter corpora. Five architectures were compared: a base GPT-4o-mini model using prompt engineering; a fine-tuned GPT-4o-mini model; a domain-specific natural language processing classifier (a support vector machine with a radial basis function kernel and term frequency–inverse document frequency features informed by symptom vocabularies); and two ensemble strategies integrating the natural language processing classifier with the fine-tuned language model, hybrid probability–indicator fusion and soft probability fusion. Ensemble weights were optimised by stratified grid search. A keyword-informed correction layer addressed negation, temporal displacement, and implicit safety proxies. Closed-form upper bounds on the permissible language model weight were derived under both fusion strategies.
Results:
Results:
The domain-specific natural language processing classifier achieved 91.8% overall accuracy, substantially exceeding the base large language model (58.7%) and the fine-tuned large language model (74.5%). The optimised hybrid ensemble reached 93.6% accuracy and reduced the suicidal ideation miss rate from 32.5% (base large language model) to 1.9%, a 17-fold safety improvement. A sharp performance cliff appeared when language model weight exceeded 50%, matching the derived closed-form bound (approximately 44% at a natural language processing calibrated confidence of 0.90). The keyword-informed correction layer correctly resolved three clinically important edge cases—implicit suicide risk, negation, and temporal displacement—that probabilistic averaging misclassified.
Conclusions:
Conclusions:
Domain-specific clinical grounding is a structural, not optional, requirement for safe mental health text classification. Probabilistic ensembles can outperform either constituent model alone when language model weight is bounded by a derivable, model-agnostic constraint, yielding clinically meaningful safety gains in suicidal ideation detection. These findings generalise to biomedical informatics applications in which model failures carry direct safety consequences and inform the safe integration of large language models into clinical decision-support systems. Clinical Trial: n/a
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.