An Audio-Visual Emotion Recognition Framework for Objective Mental Health Assessment
DOI:
https://doi.org/10.70917/ijcisim-2026-5629Keywords:
Audio-visual emotion recognition, mental health assessment, deep learning, facial emotion recognition, audio emotion recognitionAbstract
Mental health disorders are increasing rapidly due to changes in lifestyle patterns, social pressure and work-related stress. Traditional mental health assessment primarily depends on self-reported questionnaires, which are subjective and does not accurately represent individual psychological condition. To overcome these limitations, this paper proposed an objective behavioural assessment using audio and video emotion analysis for mental health assessment. This proposed framework integrates a deep convolution neural network based VideoEmoNet model and a convolutional neural-network with long short-term memory based AudioEmoNet model. The publically available dataset like FER2013 for the VideoEmoNet model and RAVDESS, TESS, SAVEE, CREMA-D for the AudioEmoNet model are used for experiments. The output of the audio-visual models is integrated using late fusion technique to provide final mental health assessment state. The AudioEmoNet model achieves 97.47% accuracy and VideoEmoNet model achieves an accuracy of 95.01%. The proposed audio-visual framework achieves 98.20% accuracy on benchmark dataset. The proposed framework improves behavioural emotion recognition performance and shows how complementary audio-visual emotion information can be integrated for behavioural evaluation which is important for mental health assessment.