ADAPTIVE CONTEXT-AWARE MULTIMODAL TRANSFORMER FRAMEWORK FOR ROBUST RESPIRATORY DISEASE DIAGNOSIS
DOI:
https://doi.org/10.70917/ijcisim-2026-5422Keywords:
Multimodal learning, transformer architecture, respiratory disease diagnosis, adaptive fusion, medical artificial intelligenceAbstract
Early detection and accurate diagnosis of respiratory diseases are remaining critical for the timely clinical interference and treatment. Many of the conventional diagnostic systems are relying on the unimodal datasets or static multimodal fusion strategies that tends in limiting the ability in capturing the respiratory diseases. Despite the success of recent multimodal learning approaches, where the practical limitations are remaining unresolved. This study proposes an Adaptive Context-Aware Multimodal Transformer Network (AMCT-Net) that is combining the chest X-ray images, respiratory acoustic signals, and structured clinical parameters for the respiratory disease diagnosis. The proposed architecture is employing three modality-specific encoders: a convolutional neural network that tends to extract the radiographic features, where a spectrogram-based audio transformer that is capturing the respiratory sound patterns, and a tabular transformer that is processing clinical variables. The extracted representations are passing through the Adaptive Modality Weighting mechanism that is assigning the dynamic importance scores to each of the modality. A context-aware multimodal transformer fusion module is then modelling the cross-modal dependencies in generating the diagnostic representations. The fused features are subsequently processed through the classification layer that tends to predict the respiratory disease categories. The experimental evaluation is showing that the proposed Adaptive Context-Aware Multimodal Transformer Network (AMCT-Net) is achieving a superior performance compared with the conventional multimodal diagnostic frameworks. The model is reaching 88.72% classification accuracy, 88.10% precision, 88.45% recall, and 88.30% F1-score, which significantly is exceeding the performance of conventional architectures that includes MDT, E-RespiNet, Med-MLLM, FATE, DREAM-OSA, CAMAF, and RMT with the mask attention.