Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Mar 19, 2026
Date Accepted: Jul 13, 2026

The final, peer-reviewed published version of this preprint can be found here:

Depictions of Depression in Generative AI Video Models: Mixed Methods Study of OpenAI’s Sora 2

Flathers M, Smith G, Herpertz J, Zhou Z, Torous J

Depictions of Depression in Generative AI Video Models: Mixed Methods Study of OpenAI’s Sora 2

J Med Internet Res 2026;28:e95682

DOI: 10.2196/95682

PMID: 42735385

Depictions of Depression in Generative AI Video Models: A Preliminary Study of OpenAI's Sora 2

  • Matthew Flathers; 
  • Griffin Smith; 
  • Julian Herpertz; 
  • Zhitong Zhou; 
  • John Torous

ABSTRACT

Background:

Generative video models are increasingly capable of producing complex depictions of mental health experiences, yet little is known about how these systems represent conditions like depression. Because AI-generated content may reach people during vulnerable periods, understanding what visual narratives these models produce for sensitive concepts carries clinical relevance.

Objective:

This study aimed to characterize how OpenAI's Sora 2 generative video model depicts depression, and to examine whether depictions differ between the consumer App and developer API access points, which differ in their product-layer mediation

Methods:

We generated 100 videos using the single-word prompt "Depression" across two access points: the consumer App (n=50) and developer API (n=50). Two trained coders independently coded narrative structure, visual environments, objects, figure demographics, and figure states. Inter-rater reliability was assessed using Cohen's kappa, with dimensions showing insufficient agreement excluded from analysis. Computational features (visual aesthetics, audio, semantic content, temporal dynamics) were extracted and compared between modalities using Welch's t-tests with Benjamini-Hochberg false discovery rate correction.

Results:

App-generated videos exhibited a pronounced recovery bias: 78% (39/50) featured narrative arcs progressing from depressive states toward resolution, compared with 14% (7/50) of API outputs. This divergence was reinforced across channels. App videos brightened over time (slope = 2.90 brightness units/second vs. -0.18 for API; d = 1.59, q < .001) and contained three times more motion (d = 2.07, q < .001). Across both modalities, videos converged on a narrow visual vocabulary: predominantly seated figures (93%), downward gaze (96%), and recurring objects including hoodies (n=194), windows (n=148), and rain (n=83). Transcript language in depressive phases emphasized weight and containment ("heavy," "drowning," "room"), while recovery phases reversed these patterns: brightness increased 27% (d = 0.70, p < .001), gaze shifted upward in 68% of recovery videos, and terms like "light" and "breath" emerged. Figures were predominantly young adults (88% aged 20-30) and nearly always alone (98%). Gender varied by access point: App outputs skewed male (68%), API outputs skewed female (59%).

Conclusions:

Sora 2 does not invent new visual grammars for depression but compresses and recombines cultural iconographies, while platform-level constraints substantially shape which narratives reach users. Clinicians should be aware that AI-generated mental health video content reflects training data and platform design rather than clinical knowledge, and that patients may encounter such content during vulnerable periods.


 Citation

Please cite as:

Flathers M, Smith G, Herpertz J, Zhou Z, Torous J

Depictions of Depression in Generative AI Video Models: Mixed Methods Study of OpenAI’s Sora 2

J Med Internet Res 2026;28:e95682

DOI: 10.2196/95682

PMID: 42735385

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.