Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: Journal of Medical Internet Research

Date Submitted: Sep 27, 2026
Open Peer Review Period: Sep 27, 2026 - Nov 22, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Acoustic Speech Features for Mental State Monitoring: Systematic Review

  • Megan Stubbs; 
  • Julian Thrash III; 
  • Shaba Rahavi; 
  • Ann Hartman; 
  • Lovette Ochieng; 
  • Alexandra Hartman; 
  • Sultan Scafi; 
  • Yitzi Devor; 
  • Zvika Shinar

ABSTRACT

Background:

Acoustic and prosodic speech features could allow continuous measurement of mental state in long term care, where monitoring relies on staff observation and periodic self assessment. Existing reviews show these features differ between people with and without a diagnosis, mostly for depression and in conversational speech. Within speaker stability, which tracking change against a person’s own baseline requires, and adherence to repeated recording have not been synthesized.

Objective:

To characterize acoustic features reported against 5 mental states (depression, anxiety, psychological stress, agitation or irritability, and apathy), the speaking tasks and designs behind that evidence, their within speaker repeatability, and adherence to repeated voice collection.

Methods:

Five databases and 3 preprint sources were searched on August 17, 2026, for English language publications from 2005 onward. Eligible studies related a named acoustic feature or feature family to 1 of the 5 states in adults, reported within speaker repeatability across separate occasions, or reported adherence to repeated voice collection. Each feature by state result was an extraction unit. Studies of short fixed content speech (sustained vowels, isolated words or phrases, or read passages) were extracted in full and formed the primary synthesis; conversational and picture description studies received reduced extraction. Risk of bias was appraised per contribution and certainty rated per feature family; synthesis was narrative, following the Synthesis Without Meta-analysis guideline.

Results:

The 125 included records yielded 1,873 results from 79 independent participant samples. Depression accounted for 1,685 results, and no record reported an acoustic feature against agitation. The primary synthesis comprised 341 results for individually named features from 15 studies of short fixed content speech; 573 of the 914 feature level results rested on conversational or pooled speech. Intensity was the only family whose direction held on both kinds of speech, and pitch was unsettled on both. Across 570 detection results on a common scale, median performance was 0.70. Three records from 3 independent samples assessed within speaker repeatability; the 2 reporting an intraclass correlation coefficient (ICC) gave only its average measure form, so the single measure ICC needed to judge change against a person’s baseline is unreported. Seven records reported adherence, which reached 79 percent across 40 weeks with clinical support and 71 percent in an unsupervised protocol. All 27 certainty ratings were very low.

Conclusions:

Acoustic evidence for mental state is concentrated in depression and conversational speech, and a direction found on one speaking task does not reliably transfer to another. The measurement property an individual monitoring system depends on, the reliability of a single recording, has not been reported by any included study, while repeated collection is achievable with human support. The literature supports a measurement design rather than a feature list. Clinical Trial: PROSPERO CRD420261473970


 Citation

Please cite as:

Stubbs M, Thrash J III, Rahavi S, Hartman A, Ochieng L, Hartman A, Scafi S, Devor Y, Shinar Z

Acoustic Speech Features for Mental State Monitoring: Systematic Review

JMIR Preprints. 27/09/2026:113062

DOI: 10.2196/preprints.113062

URL: https://preprints.jmir.org/preprint/113062

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.