Dual-Attention MobileViT++: A Multi-Scale CNN–Transformer Framework for High-Precision Oral Cancer Classification
DOI:
https://doi.org/10.70917/ijcisim-2026-4093Keywords:
Oral cancer detection, deep learning, MobileViT, dual attention, medical image classification, CNN–transformer hybrid, clinical decision supportAbstract
Oral cancer is one of the life-threatening diseases with a high mortality rate resulting from late diagnosis and limited access to expert screening due to specialized training. The available detection approaches, which are mainly based on visual observations and biopsied analysis, are invasive, time-consuming, and experience dependent. Therefore, a reliable and automated approach is needed to develop an efficient detection system. This study introduced the Dual-Attention MobileViT++ architecture, which integrates the multi-scale depthwise CNN backbone, the hybrid attention models, and the hierarchical MobileViT transformers to characterize the local lesion texture and global contextual patterns. The model was tested on the Kaggle public dataset with no additional modification and enhancement. The oral cancer dataset was modified through medical-grade augmentation to create more balanced datasets comprising about 981 images. Experimental results showed a promising performance, achieving approximately 99.93% accuracy with high precision, recall, and F1-score performed on the validation and test datasets. This result shows that the proposed architecture can generalize well and deliver reliable classification, suitable for oral cancer screening at an early stage.