An Explainable Hybrid Deep Learning Framework Integrating SepConv2D, DenseNet121, and Vision Transformer for Automated Skin Lesion Classification

Authors

  • Adesh V. Panchal Gujarat Technological University (GTU), Ahmedabad, Gujarat, India
  • Manish M. Patel Sankalchand Patel College of Engineering, Sankalchand Patel University, Vis Nagar, Gujarat, India

DOI:

https://doi.org/10.70917/ijcisim-2026-4064

Keywords:

Skin lesion classification, Explainable artificial intelligence, Vision Transformer, DenseNet121, Depth wise Separable Convolution, Grad-CAM, Deep learning, medical image analysis

Abstract

Early and accurate diagnosis of skin lesions is essential for reducing the mortality associated with skin cancer and improving patient outcomes. Although deep learning has significantly advanced automated dermatological diagnosis, existing approaches often struggle to simultaneously achieve high classification accuracy, computational efficiency, and model interpretability. Conventional convolutional neural networks effectively learn local texture patterns but are limited in capturing long-range contextual information, whereas transformer-based models provide superior global feature representation at the expense of increased computational complexity. To address these challenges, this paper proposes an explainable hybrid deep learning framework that integrates Depthwise Separable Convolution (SepConv2D), Vision Transformer (ViT), and DenseNet121 for automated skin lesion classification and localization. The proposed architecture exploits SepConv2D for efficient local feature extraction, ViT for modelling global contextual relationships through self-attention, and DenseNet121 for enhanced feature propagation and semantic representation. The complementary features are fused to improve discriminative learning, while Gradient-weighted Class Activation Mapping (Grad-CAM) is incorporated to provide visual explanations by highlighting clinically relevant lesion regions. The framework is evaluated using the ISIC-2019 and HAM10000 benchmark datasets under identical experimental settings. Experimental results demonstrate an overall classification accuracy of 99.94%, with corresponding precision, recall, and F1-score of 99.94%, while maintaining competitive computational complexity in terms of model parameters, MACs, and FLOPs. Comparative evaluation against recent state-of-the-art methods indicates that the proposed framework consistently achieves superior predictive performance and enhanced interpretability. These findings demonstrate that the proposed architecture can serve as an efficient and trustworthy computer-aided diagnostic tool for supporting dermatologists in clinical decision-making and facilitating the adoption of explainable artificial intelligence in healthcare applications.

Downloads

Download data is not yet available.

Downloads

Published

2026-07-31

How to Cite

Adesh V. Panchal, & Manish M. Patel. (2026). An Explainable Hybrid Deep Learning Framework Integrating SepConv2D, DenseNet121, and Vision Transformer for Automated Skin Lesion Classification. International Journal of Computer Information Systems and Industrial Management Applications, 18(13s), 451–471. https://doi.org/10.70917/ijcisim-2026-4064

Issue

Section

Original Articles