Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Aging

Date Submitted: Feb 10, 2026
Date Accepted: Aug 31, 2026

The final, peer-reviewed published version of this preprint can be found here:

Multimodal Dementia Prediction With Large Language Models: Cross-Attention Over Text, Audio, and Image

Agbavor F, Liang H

Multimodal Dementia Prediction With Large Language Models: Cross-Attention Over Text, Audio, and Image

JMIR Aging 2026;9:e93279

DOI: 10.2196/93279

PMID: 42766599

Multimodal Dementia Prediction with Large Language Models: Cross-Attention over Text, Audio, and Image

  • Felix Agbavor; 
  • Hualou Liang

ABSTRACT

Background:

Alzheimer’s disease (AD) is a leading cause of dementia, and there is growing interest in scalable approaches for early screening using speech-based tasks. While prior work has demonstrated promising results using either transcript-based language features or acoustic cues, most approaches remain unimodal or rely on simple fusion strategies that do not explicitly consider interactions across modalities.

Objective:

In this study, we propose an attention-based tri-modal fusion framework that integrates text, audio, and image representations from the Cookie-Theft picture-description task.

Methods:

Our method employs a novel bidirectional cross-attention mechanism to achieve a unified multimodal embedding for downstream tasks. We evaluate the approach on two tasks: AD detection by classifying whether the subject is AD or not, and AD severity assessment by predicting Mini-Mental Status Examination (MMSE) cognitive scores.

Results:

On the AD detection task, tri-modal fusion achieves the best overall performance (F1 = 0.8889, AUC = 0.9032), outperforming unimodal baselines, bimodal fusion, and conventional early/late fusion methods. For AD severity assessment, the proposed multimodal representation reduces prediction error of RMSE to about 4.20, improving over both unimodal and bimodal fusion settings. We further perform the ablation analysis to show that bidirectional cross-attention consistently outperforms conventional unidirectional cross-attention.

Conclusions:

These results demonstrate that attention-based multimodal fusion can enhance dementia prediction from picture-description responses and provide a strong foundation for developing multimodal cognitive screening pipelines.


 Citation

Please cite as:

Agbavor F, Liang H

Multimodal Dementia Prediction With Large Language Models: Cross-Attention Over Text, Audio, and Image

JMIR Aging 2026;9:e93279

DOI: 10.2196/93279

PMID: 42766599

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.