A Hybrid CNN-Transformer Framework For Multimodal Emotion Recognition In Healthcare: A Comparative Study
DOI:
https://doi.org/10.70917/ijcisim-2026-4488Keywords:
Multimodal Emotion Recognition, Hybrid CNN–Transformer, Affective Computing, Intelligent Healthcare Systems, Deep Learning, Emotion-Aware Artificial Intelligence (AI)Abstract
The recognition of emotions is a crucial research field in the domain of intelligent healthcare, as emotional states are related to diagnosis, adherence to therapy, patient safety, mental health, pain perception, and quality of care. In the clinical setting, patients present affect not just by way of words, but also through their facial expression, vocal quality, body movements and reactions. In fact, traditional emotion recognition algorithms—especially unimodal ones—do not capture the complexity of these emotions, as they rely on the information provided from a single source and are therefore sensitive to noise, occlusion, missing information and individual differences. To tackle this drawback, multimodal emotion recognition combines the complementary information from multiple modalities, which increases the robustness and predictive reliability. Meanwhile, deep-learning breakthroughs have ushered in two hugely popular architectural families for this use case: Convolutional Networks, which are adept at local feature extraction, and Transformers, which are very good at the long-range dependency and contextual relationships. This article is the first part of a comparative research paper of a hybrid CNN-Transformer approach to multimodal emotion recognition in healthcare. This study investigates the integration of CNN and Transformer components for enhancing emotional understanding in complex medical contexts like telemedicine, mental health surveillance, pain analysis, rehabilitation, and elderly care. A comparative approach is taken, where hybrid architectures are compared to traditional machine learning approaches, as well as to standalone CNN-based and standalone Transformer-based systems. The main idea is that a combination of local analysis and global reasoning is required for emotion recognition within the healthcare domain, and that such a combination is especially appropriate for the application of emotion recognition in healthcare. The article first presents why the problem of emotion-aware Artificial Intelligence (AI) is relevant in healthcare, and then it summarizes the concepts behind multimodal affective computing. It then elaborates the pros and cons of each modeling paradigm, and suggests a hybrid paradigm that leverages CNNs to extract features from each modality and Transformer layers to fuse features across modalities. The expected advantages of the hybrid model are accuracy, robustness, temporal sensitivity and adaptability to heterogeneous healthcare data sources as compared to the other models, that are detailed in a comparative discussion. Some of the main challenges, such as the availability of annotated clinical data, computational requirements, interpretability, privacy, fairness, and deployment feasibility in real-world healthcare systems, are also discussed. It is concluded that hybrid CNN-Transformer approaches are a promising and practically relevant approach for emotion recognition in healthcare in multimodal settings. Beyond their technical performance, their significance is in being able to effectively support emotionally aware, context-sensitive and patient-oriented intelligent systems. But, to make it successful, it is imperative to have evaluation based on clinical evidence, explainable outputs, and ethical design practices. This paper is a comparative study which serves as a baseline for further studies and development on the reliable, transparent and deployable emotion recognition systems for modern healthcare environments. Multimodal emotion recognition is an essential task in health care, playing a vital role in patient wellbeing. Emotion recognition is an important task in healthcare and has a critical impact on patient well-being, especially in multimodal recognition.