Previously submitted to: JMIR Mental Health (no longer under consideration since Dec 08, 2025)
Date Submitted: Dec 7, 2025
Open Peer Review Period: Dec 8, 2025 - Dec 8, 2025
(closed for review but you can still tweet)
NOTE: This is an unreviewed Preprint
Warning: This is a unreviewed preprint (What is a preprint?). Readers are warned that the document has not been peer-reviewed by expert/patient reviewers or an academic editor, may contain misleading claims, and is likely to undergo changes before final publication, if accepted, or may have been rejected/withdrawn (a note "no longer under consideration" will appear above).
Peer review me: Readers with interest and expertise are encouraged to sign up as peer-reviewer, if the paper is within an open peer-review period (in this case, a "Peer Review Me" button to sign up as reviewer is displayed above). All preprints currently open for review are listed here. Outside of the formal open peer-review period we encourage you to tweet about the preprint.
Citation: Please cite this preprint only for review purposes or for grant applications and CVs (if you are the author).
Final version: If our system detects a final peer-reviewed "version of record" (VoR) published in any journal, a link to that VoR will appear below. Readers are then encourage to cite the VoR instead of this preprint.
Settings: If you are the author, you can login and change the preprint display settings, but the preprint URL/DOI is supposed to be stable and citable, so it should not be removed once posted.
Submit: To post your own preprint, simply submit to any JMIR journal, and choose the appropriate settings to expose your submitted version as preprint.
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Human-in-the-Loop Machine Learning for Depression Risk Screening among Bangladeshi University Students: An Explainable Decision Support Framework
ABSTRACT
Background:
Depression is a major contributor to the global burden of disease and is particularly prevalent among young people, including university students. In Bangladesh, limited mental health resources, stigma, and gaps in early identification can delay timely support for students who are at risk. At the same time, digital health and machine learning approaches create an opportunity to scale screening and improve triage in resource-constrained settings, but these tools must be clinically aligned, interpretable, and fair to be trusted in real services. Against this backdrop, this study develops a human-in-the-loop, explainable, and calibrated machine-learning decision-support framework that translates predicted depression risk into clear tiers (Normal, Mild, Moderate, Severe) to assist counselors in prioritizing care while also providing understandable feedback for students
Objective:
**Objective of the study** The overarching objective of this study is to develop and evaluate a human-in-the-loop, explainable, and clinically usable machine-learning decision-support framework for depression risk screening among Bangladeshi university students, with emphasis on reliable probability calibration and actionable risk communication. **Specific objectives** 1. To adapt and assess a cross-population depression screening approach by leveraging relevant external survey evidence alongside local Bangladeshi student data where appropriate. 2. To implement and compare probability calibration strategies and identify clinically meaningful thresholds to enhance trust and decision alignment. 3. To convert model probabilities into four intuitive risk tiers (Normal, Mild, Moderate, Severe) to support counselor triage and student-facing understanding. 4. To integrate explainability and subgroup fairness checks to ensure transparent and equitable screening outcomes across demographic groups.
Methods:
This study used a quantitative, cross-population transfer-learning design for depression risk screening among Bangladeshi university students. The pipeline leveraged NHANES 2017–2018 as a source dataset and a Bangladeshi student dataset as the target, retaining only shared variables (e.g., age, gender, marital status, smoking, alcohol use, weekly physical activity). Preprocessing included missing-value imputation, feature standardization, one-hot encoding, and SMOTE to address class imbalance. A model zoo (logistic regression, random forest, and gradient boosting variants) was evaluated, with the best model selected using ROC-AUC and PR-AUC criteria, then calibrated using Platt scaling and isotonic regression. Clinically meaningful thresholds were tuned (F1/Youden), and calibrated probabilities were mapped into four risk tiers (Normal, Mild, Moderate, Severe). Explainability (SHAP), subgroup fairness checks, decision-curve analysis, and bootstrap-based validation were applied to ensure transparent, reliable, and counselor-usable outputs.
Results:
The calibrated Random Forest model showed excellent performance on the Bangladeshi holdout set, achieving **ROC–AUC = 0.993**, **PR–AUC = 0.997**, **Accuracy = 0.977**, **F1 = 0.981**, and strong calibration with a **Brier score = 0.021** (with supportive bootstrap CIs reported). The evaluation was conducted on a clearly defined split of the Bangladeshi student cohort (**1,400** in the calibration set and **600** in the holdout set), where severe depressive symptoms were notably prevalent at about **61.3%** in both subsets. These results suggest the framework can provide reliable, tiered risk outputs suitable for counselor triage and student-facing screening support.
Conclusions:
This study concludes that a calibrated transfer-learning approach can feasibly adapt NHANES-based depression screening signals for use among Bangladeshi university students. The findings highlight three key contributions: maintained discrimination and calibration after transfer and probability scaling, clinically meaningful cut-offs through threshold optimization, and conversion of predicted probabilities into four intuitive risk tiers (Normal, Mild, Moderate, Severe) that are understandable for both counselors and students. Practically, the framework can reduce counselor workload via triage while offering a simple, potentially less stigmatizing self-understanding for students; however, it is intended as decision support rather than diagnosis and requires broader external validation and real-world co-design before routine clinical deployment.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.