Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Previously submitted to: JMIR AI (no longer under consideration since Apr 24, 2026)

Date Submitted: Oct 4, 2024

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Evaluating Machine Learning-Based Ultra-Brief Versions of the PHQ-9 as Major Depression Screening Instruments in Primary Care

  • Darragh Glavin; 
  • Eoin Martino Grua; 
  • Carina Akemi Nakamura; 
  • Bruce Arroll; 
  • Tim J Peters; 
  • Pepijn van de Ven

ABSTRACT

Background:

Empirical evidence supporting the superiority of the Patient Health Questionnaire-2 (PHQ-2), which queries the cardinal symptoms of depression (anhedonia and depressed mood), over alternative Patient Health Questionnaire-9 (PHQ-9) item pairings for major depression screening is limited. Our previous analysis indicated that alternative PHQ-9 item pairings were more effective than the PHQ-2 at screening for depressive symptomatology.

Objective:

To improve the effectiveness of ultra-brief major depression screening by identifying the most effective ultra-brief version of the PHQ-9, using the computerised Composite International Diagnostic Interview (CIDI-Auto) as the reference standard diagnosis, and inputting item pairings into machine learning (ML) models.

Methods:

A data-driven, machine learning (ML)-based framework was applied to the PHQ-9 to derive the most effective ultra-brief version for major depression screening on a primary care data set that had previously been used to validate the PHQ-2. A post-hoc statistical analysis investigated individual item performance and the impact of item interrelationships on pairing performance.

Results:

The phq2&3 (ML-based depressed mood and sleep disturbances item pairing) emerged as the most effective screening instrument in terms of area under the curve (AUC) during the training process (0.904), outperforming the PHQ-2 (0.883). The phq2&3 also had a higher Youden’s index than the PHQ-2 (0.670 vs. 0.637). As single-item screeners, phq2 and phq6 had the highest AUC and Youden’s index. The phq6 had the highest correlation to the CIDI-Auto diagnosis, with phq2 ranking fourth. The phq3 had the sixth highest AUC and Youden’s index, and the eight highest correlation to the CIDI-Auto diagnosis. Of the 36 item pairings, the highest inter-item correlation was between items phq2 and phq6, while items phq2 and phq3 had the seventh lowest. The phq2&3 achieved a higher AUC than the PHQ-2 on the withheld test data (0.898 vs. 0.882) but scored lower in terms of Youden’s index (0.646 vs. 0.653). Notably, the phq1&2 (ML version of the PHQ-2) had a higher AUC (0.890) than the PHQ-2. It also had the highest Youden’s index (0.677) of all instruments on the test data, with a higher specificity than the PHQ-2 (0.805 vs 0.781) and the same sensitivity (0.872).

Conclusions:

Evaluating individual item performance overlooks the inter-relationship between the item pairings in ultra-brief instruments. While highly inter-correlated items indicate good internal reliability, this often harms their predictiveness as a pairing. Though the phq3 item exhibited relatively weak predictive performance individually, it complemented and enhanced screening performance when combined with the phq2, and achieved the best pairing screening performance. The use of ML models with ultra-brief instruments enhanced depression screening performance compared to traditional sum score methods. This was attributed to the flexibility provided by ML models in mapping item pairings to CIDI-Auto diagnosis and the increased number of thresholds available with ML-based screening instruments.


 Citation

Please cite as:

Glavin D, Grua EM, Nakamura CA, Arroll B, Peters TJ, van de Ven P

Evaluating Machine Learning-Based Ultra-Brief Versions of the PHQ-9 as Major Depression Screening Instruments in Primary Care

JMIR Preprints. 04/10/2024:67176

DOI: 10.2196/preprints.67176

URL: https://preprints.jmir.org/preprint/67176

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.