A Multimodal Deep Learning Framework for Brain Tumor Classification Using EfficientNet and Vision Transformers

Authors

  • Lovenish Sharma Ajeenkya DY Patil University, Pune, India
  • Saroj Kumar Nanda Ajeenkya DY Patil University, Pune, India.

DOI:

https://doi.org/10.70917/ijcisim-2026-4327

Keywords:

Brain Tumor, Multimodal, Deep Learning Framework, EfficientNet-B7, Vision Transformer, CT, MRI

Abstract

Brain tumors are becoming life-threatening issues when not treated at earlier stages. Accurate brain tumor detection is essential in order to support clinical decisions and diagnosis. The study uses the multimodal deep learning framework for brain tumor classification by effectively combining the EfficientNet-B7 network and vision transformer architectures. Modality specific feature extraction has been carried out using the EfficientNet-B7 network for both the Computed Tomography (CT) and Magnetic Resonance Imaging (MRI). Then feature maps are converted into token sequence and passed through the transformer encoder layers in order to effectively model the global dependencies across the image. Then these features are fused and then classified using the fully connected layers whether it is healthy or tumor. Based on the experimental evaluation, the proposed model in this study such as MM-EffiFormer demonstrated the significant classification performance by achieving Accuracy of 0.990, Sensitivity of 0.987, Specificity of 0.995, Dice Similarity Coefficient (DSC) of 0.991, outperforms the existing model FCM-SVM.

Downloads

Download data is not yet available.

Downloads

Published

2026-08-04

How to Cite

Lovenish Sharma, & Saroj Kumar Nanda. (2026). A Multimodal Deep Learning Framework for Brain Tumor Classification Using EfficientNet and Vision Transformers. International Journal of Computer Information Systems and Industrial Management Applications, 18(14s), 1130–1142. https://doi.org/10.70917/ijcisim-2026-4327

Issue

Section

Original Articles