Multimodal Dementia Prediction with Large Language Models: Cross-Attention over Text, Audio, and Image
ABSTRACT
Background:
Alzheimer’s disease (AD) is a leading cause of dementia, and there is growing interest in scalable approaches for early screening using speech-based tasks. While prior work has demonstrated promising results using either transcript-based language features or acoustic cues, most approaches remain unimodal or rely on simple fusion strategies that do not explicitly consider interactions across modalities.
Objective:
In this study, we propose an attention-based tri-modal fusion framework that integrates text, audio, and image representations from the Cookie-Theft picture-description task.
Methods:
Our method employs a novel bidirectional cross-attention mechanism to achieve a unified multimodal embedding for downstream tasks. We evaluate the approach on two tasks: AD detection by classifying whether the subject is AD or not, and AD severity assessment by predicting Mini-Mental Status Examination (MMSE) cognitive scores.
Results:
On the AD detection task, tri-modal fusion achieves the best overall performance (F1 = 0.8889, AUC = 0.9032), outperforming unimodal baselines, bimodal fusion, and conventional early/late fusion methods. For AD severity assessment, the proposed multimodal representation reduces prediction error of RMSE to about 4.20, improving over both unimodal and bimodal fusion settings. We further perform the ablation analysis to show that bidirectional cross-attention consistently outperforms conventional unidirectional cross-attention.
Conclusions:
These results demonstrate that attention-based multimodal fusion can enhance dementia prediction from picture-description responses and provide a strong foundation for developing multimodal cognitive screening pipelines.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.