Currently submitted to: JMIR AI
Date Submitted: Jun 12, 2026
Open Peer Review Period: Jul 8, 2026 - Sep 2, 2026
(closed for review but you can still tweet)
NOTE: This is an unreviewed Preprint
Warning: This is a unreviewed preprint (What is a preprint?). Readers are warned that the document has not been peer-reviewed by expert/patient reviewers or an academic editor, may contain misleading claims, and is likely to undergo changes before final publication, if accepted, or may have been rejected/withdrawn (a note "no longer under consideration" will appear above).
Peer review me: Readers with interest and expertise are encouraged to sign up as peer-reviewer, if the paper is within an open peer-review period (in this case, a "Peer Review Me" button to sign up as reviewer is displayed above). All preprints currently open for review are listed here. Outside of the formal open peer-review period we encourage you to tweet about the preprint.
Citation: Please cite this preprint only for review purposes or for grant applications and CVs (if you are the author).
Final version: If our system detects a final peer-reviewed "version of record" (VoR) published in any journal, a link to that VoR will appear below. Readers are then encourage to cite the VoR instead of this preprint.
Settings: If you are the author, you can login and change the preprint display settings, but the preprint URL/DOI is supposed to be stable and citable, so it should not be removed once posted.
Submit: To post your own preprint, simply submit to any JMIR journal, and choose the appropriate settings to expose your submitted version as preprint.
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Characterizing schizophrenia and major depressive disorder through acoustic and linguistic speech markers: an observational study with machine learning analyses
ABSTRACT
Background:
Automated speech analysis offers a low-burden approach for identifying objective markers of psychiatric assessment. While alterations in speech have been documented in both schizophrenia (SZ) and major depressive disorder (MDD), direct comparisons between these conditions within the same cohort remain limited.
Objective:
This study aimed to examine whether automatically extracted acoustic and linguistic speech features differentiate individuals with SZ, individuals with MDD, and healthy controls (HC), and whether these features are associated with clinical symptom severity and diagnostic classification performance.
Methods:
A total of 66 participants (22 with SZ, 22 with MDD, and 22 HC) completed a tablet-based speech assessment, at two time points. Speech tasks included a semi-structured picture description task and emotional storytelling prompts. Automatic extraction of acoustic, temporal, spectral, and linguistic features was performed. Group differences were tested using Kruskal-Wallis and Mann-Whitney U tests, symptom–speech associations were examined using correlation analyses with correction for multiple testing, and diagnostic classification was evaluated using machine learning models with leave-one-out cross-validation. Speech-based models were compared with demographic and symptom-based baseline models.
Results:
The clearest group differences emerged in the picture description task. Participants with SZ exhibited reduced speech output, lower informational content, and altered acoustic features compared to those with MDD or HC. Omnibus group effects survived correction for multiple testing for certain speech features, such as correct concepts, number of pauses and word count (adjusted P<.05). In contrast, univariate differences in emotional storytelling tasks were limited and mainly observed for the positive storytelling task at the second time point. Symptom-speech associations were selective after correction for multiple testing, including associations between depressive symptoms and lower loudness as well as between PANSS scores and jitter-related measures. Within the present sample, speech-based machine learning models showed the most consistent classification performance for distinguishing SZ from MDD, with area under the curve (AUC) values of approximately 0.91 to 0.96 across tasks, whereas performance for distinguishing HC from SZ was moderate.
Conclusions:
Automated speech features captured diagnostically relevant information across SZ, MDD, and HC, with the clearest differentiation observed between SZ and MDD. Semi-structured picture description appeared particularly sensitive to disorder-related differences in speech output and informational content. These results suggest that automated speech analysis could be used as a complementary tool for differential diagnosis and symptom monitoring, but larger multisite studies with external validation are required before clinical implementation.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.