Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Mar 19, 2026
Date Accepted: Jul 13, 2026
Depictions of Depression in Generative AI Video Models: A Preliminary Study of OpenAI's Sora 2
ABSTRACT
Background:
Generative video models are increasingly capable of producing complex depictions of mental health experiences, yet little is known about how these systems represent conditions like depression. Because AI-generated content may reach people during vulnerable periods, understanding what visual narratives these models produce for sensitive concepts carries clinical relevance.
Objective:
This study aimed to characterize how OpenAI's Sora 2 generative video model depicts depression, and to examine whether depictions differ between the consumer App and developer API access points, which differ in their product-layer mediation
Methods:
We generated 100 videos using the single-word prompt "Depression" across two access points: the consumer App (n=50) and developer API (n=50). Two trained coders independently coded narrative structure, visual environments, objects, figure demographics, and figure states. Inter-rater reliability was assessed using Cohen's kappa, with dimensions showing insufficient agreement excluded from analysis. Computational features (visual aesthetics, audio, semantic content, temporal dynamics) were extracted and compared between modalities using Welch's t-tests with Benjamini-Hochberg false discovery rate correction.
Results:
App-generated videos exhibited a pronounced recovery bias: 78% (39/50) featured narrative arcs progressing from depressive states toward resolution, compared with 14% (7/50) of API outputs. This divergence was reinforced across channels. App videos brightened over time (slope = 2.90 brightness units/second vs. -0.18 for API; d = 1.59, q < .001) and contained three times more motion (d = 2.07, q < .001). Across both modalities, videos converged on a narrow visual vocabulary: predominantly seated figures (93%), downward gaze (96%), and recurring objects including hoodies (n=194), windows (n=148), and rain (n=83). Transcript language in depressive phases emphasized weight and containment ("heavy," "drowning," "room"), while recovery phases reversed these patterns: brightness increased 27% (d = 0.70, p < .001), gaze shifted upward in 68% of recovery videos, and terms like "light" and "breath" emerged. Figures were predominantly young adults (88% aged 20-30) and nearly always alone (98%). Gender varied by access point: App outputs skewed male (68%), API outputs skewed female (59%).
Conclusions:
Sora 2 does not invent new visual grammars for depression but compresses and recombines cultural iconographies, while platform-level constraints substantially shape which narratives reach users. Clinicians should be aware that AI-generated mental health video content reflects training data and platform design rather than clinical knowledge, and that patients may encounter such content during vulnerable periods.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.