A Novel Enhanced ViT-CNN-ATS Architecture with Trigonometric Filters for Breast Cancer Segmentation and Classification in Ultrasound Imaging
DOI:
https://doi.org/10.70917/ijcisim-2026-4441Keywords:
Breast Cancer Detection, Ultrasound Imaging, Vision Transformer (ViT), Convolutional Neural Networks (CNN), Attention Token Selection (ATS), Dice Score, Intersection over Union (IoU)Abstract
This study proposes a novel hybrid deep learning framework, ViT-CNN-ATS, designed for accurate classification and segmentation of breast cancer in ultrasound images. The model integrates the global contextual learning of the Vision Transformer (ViT) with the powerful local feature extraction capabilities of Convolutional Neural Networks (CNN), further refined by Attention Token Selection (ATS) for enhanced focus on the most informative regions. To overcome the inherent challenges of ultrasound imaging such as low contrast, speckle noise, and variability in tumor appearance comprehensive preprocessing techniques including contrast enhancement and noise reduction are employed. A masked ultrasound image dataset, organized into benign, malignant, and normal classes, serves as the foundation for training and evaluation. Data augmentation strategies are incorporated to improve model robustness and generalization. The proposed hybrid model demonstrates significant improvements in classification accuracy, Dice Similarity Coefficient, and Intersection over Union (IoU) scores when compared to baseline models using only CNN or ViT. These results underscore the potential of ViT-CNN-ATS as a reliable, AI-assisted diagnostic tool to support radiologists in early and precise breast cancer detection using ultrasound imaging.