Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Previously submitted to: JMIR Mental Health (no longer under consideration since Jan 26, 2024)

Date Submitted: Jan 26, 2024
Open Peer Review Period: Jan 26, 2024 - Jan 26, 2024
(closed for review but you can still tweet)

NOTE: This is an unreviewed Preprint

Warning: This is a unreviewed preprint (What is a preprint?). Readers are warned that the document has not been peer-reviewed by expert/patient reviewers or an academic editor, may contain misleading claims, and is likely to undergo changes before final publication, if accepted, or may have been rejected/withdrawn (a note "no longer under consideration" will appear above).

Peer review me: Readers with interest and expertise are encouraged to sign up as peer-reviewer, if the paper is within an open peer-review period (in this case, a "Peer Review Me" button to sign up as reviewer is displayed above). All preprints currently open for review are listed here. Outside of the formal open peer-review period we encourage you to tweet about the preprint.

Citation: Please cite this preprint only for review purposes or for grant applications and CVs (if you are the author).

Final version: If our system detects a final peer-reviewed "version of record" (VoR) published in any journal, a link to that VoR will appear below. Readers are then encourage to cite the VoR instead of this preprint.

Settings: If you are the author, you can login and change the preprint display settings, but the preprint URL/DOI is supposed to be stable and citable, so it should not be removed once posted.

Submit: To post your own preprint, simply submit to any JMIR journal, and choose the appropriate settings to expose your submitted version as preprint.

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

WONDER - Waveform-Based Optimal Neurological Depression Evaluation Using Representations via Speaker Identity Invariant Training: An Observational Study

  • Biman Najika Liyanage; 
  • Yunhan Lin; 
  • Zhengwen Zhu; 
  • Jun Yang; 
  • Jun Yang; 
  • Zongfeng Li; 
  • Yutao Sub; 
  • John Wong Chee Meng; 
  • Weihua Yue

ABSTRACT

Background:

Recent progress in identifying depression through speech patterns has accelerated because of the enhanced precision of foundation models. These models, trained on vast unlabeled speech datasets using self-supervised learning (SSL), extract powerful speech characteristics within their transformer encoder layers. The self-supervised representations of speech capture para-linguistic features that contain information about psychomotor retardation symptoms of individuals with major depressive disorder (MDD). Fully automated depression screening methods have been exploring the potential of fine-tuning such foundation models for downstream tasks such as speech emotion recognition and speech-based depression detection. However, it is vital to finetune such models with diverse datasets beyond laboratory-collected ones for the models to be robust and generalized when deployed in real-world applications.

Objective:

This study aims to introduce two novel approaches to improve automated depression screening methods, and improve the system’s robustness. The first objective is to tackle the challenge of limited data, and the second is to train the model in a way that minimizes speaker-specific biases. Achieving these two objectives is vital to ensure that models are robust and generalized when deployed in real world applications.

Methods:

The first objective is achieved by leveraging self-supervised representations and finetuning the model on a diverse real-world dataset with data augmentations. The latter is achieved by using a speaker invariant training architecture enforcing the model to learn para-linguistic features while discarding information related to speaker-specific features. This approach ensures the model discerns speech features indicative of depression, irrespective of the speaker.

Results:

The study validates the methodology across two unique language datasets using only raw acoustic components of speech. The proposed approach employs an adversarial training method, outperforms the baseline model, and achieves a macro F1-Score of 0.83, setting a new state-of-the-art (SOTA) on the publicly available DAIC-WOZ. Similarly, the Oizys Chinese corpus achieves a sensitivity of 0.8 and AUC of 0.82 with a relative improvement of 23% in sensitivity compared to the baseline model.

Conclusions:

The study advances biomedical research, enhancing automated depression screening with an efficient, effective, and simple voice interface-to-screen MDD at scale with high precision and reliability using robust vocal biomarkers. Furthermore, the fine-tuned models capture robust features to detect depression using speech with a minimal speaker bias. These self-supervised representations are objective and consistent across fundamentally different language corpuses; they therefore have the potential of being used for multilingual depression screening using speech. Clinical Trial: Our study was a prospective cohort study rather than a randomized controlled trial. The registry center encouraged but did not mandate the study’s registration. We adhered to ethical standards, meeting the ethical guidelines of and securing approval from the ethics committee of Peking University's Sixth Hospital. All participants provided written informed consent. The study was approved on December 15, 2020 (approval number: 62). The participants' privacy and confidentiality were rigorously upheld.


 Citation

Please cite as:

Liyanage BN, Lin Y, Zhu Z, Yang J, Yang J, Li Z, Sub Y, Meng JWC, Yue W

WONDER - Waveform-Based Optimal Neurological Depression Evaluation Using Representations via Speaker Identity Invariant Training: An Observational Study

JMIR Preprints. 26/01/2024:56710

DOI: 10.2196/preprints.56710

URL: https://preprints.jmir.org/preprint/56710

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.