An Integrated AI Framework for Depression Detection Using Audio, Visual and Textual Data
DOI:
https://doi.org/10.70917/ijcisim-2026-5562Keywords:
Ensemble Learning, Stacking Method, Machine Learning, Multimodal Depression, Unimodal, BimodalAbstract
Depression, as a mental health condition, often manifests through diverse verbal, non-verbal, and behavioral signals. Early and accurate detection of these indicators is critical to ensuring timely and effective intervention. This work introduces a multimodal Integrated AI framework for depression detection that uses: audio, visual, and textual data to improve classification performance. The proposed framework integrates features through ensemble stacking techniques by combination of depressive indicators such as facial expressions, speech patterns, and linguistic text content to form a holistic representation of depression detection. Conducted experiments demonstrates that the by integration of these modalities significantly enhanced classification performance. In all the conducted experiments, multimodal feature vectors outperformed unimodal and bimodal approaches. This best performance was achieved by this novel framework and reported an accuracy of 88.97%, precision of 91.42%, recall of 89.31%, and F1-score of 92.75%. By utilizing Distress Analysis Interview Corpus (DAIC), an open source dataset and proposed feature extraction methods, the system effectively fuses audio, visual, and textual cues, establishing a more reliable approach for detecting depression.