Enhancing Depression Detection with Multimodal Cross-Attention: A Robust Framework Using the D-Vlog Dataset

Authors

  • Taufeeq Ahmed Manav Rachna International Institute of Research and Studies, Faridabad, India.
  • Ramesh Chandra Sahoo Manav Rachna International Institute of Research and Studies, Faridabad, India.

DOI:

https://doi.org/10.70917/ijcisim-2026-5840

Keywords:

Multimodal Learning, Depression Detection, Psychological Assessment, Cross-Attention Mechanism, Multimodal Dataset

Abstract

Worldwide, depression is a serious public health issue, and struggles around early detection are evident in countless clinical practices. Nearly all traditional methods of depression detection involve interviewing patients and having patients fill out screening questionnaires, a process that is inefficient and tough to reproduce. As depression detection methods become more automated, the ability to identify clues that may suggest depression through behavioral changes in people is capturing more attention. This is especially true as more people share their personal lives through vlogs via social media. In this paper, we present the various techniques that we have built to find depression in vlog videos. In the proposed method, we employ a cross-attention mechanism to integrate acoustic and visual behavioral cues (or “modalties”) that reflect the alignment (or relationship) between the two. Then, we conduct cross-modal fusion of the processed individual streams from the specific encoders (or, ‘modality encoders’). We empirically assessed our proposed model using the D-Vlog dataset, consisting of 961 vlogs from 816 patients with depression annotations. Our proposed model is able to achieve 82% accuracy, 0.79 F1-score, and 0.89 Area Under Curve (AUC) score. Compared with unimodal/depression detection systems and traditional (i.e., early & late) cross fusion models, our proposed model outperformed all these models. The results suggest that explicitly modelling cross-modal relationships improves the ability to capture behavioural patterns associated with depressive states. These findings highlight the potential of multimodal learning for developing automated tools that can assist in large-scale depression screening using real-world video data.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-04

How to Cite

Taufeeq Ahmed, & Ramesh Chandra Sahoo. (2026). Enhancing Depression Detection with Multimodal Cross-Attention: A Robust Framework Using the D-Vlog Dataset. International Journal of Computer Information Systems and Industrial Management Applications, 18(23s), 1816–1829. https://doi.org/10.70917/ijcisim-2026-5840

Issue

Section

Original Articles