CrackNet-VOS: An Intelligent Vision-Based Framework for Multi-Class Concrete Crack Classification by Type, Orientation, and Severity in Infrastructure Surfaces
DOI:
https://doi.org/10.70917/ijcisim-2026-3005Keywords:
Structural health monitoring, Concrete crack classification, multi-task learning, Vision transformer, Hybrid CNN–Transformer, Infrastructure inspectionAbstract
Concrete surface cracks threaten the safety and service life of civil infrastructure, yet most existing vision-based approaches still address only one crack attribute (typically type or severity) at a time. This work introduces CrackNet-VOS, a vision-based framework that performs joint classification of crack type, orientation, and severity on concrete elements such as beams, slabs, pavements, and tunnels. A multi-source dataset of 7,200 images was compiled from SDNET201, CRACK500, the Concrete Crack Forest Dataset (CCFD), and newly collected field data under varied lighting and environmental conditions. Each image is annotated with three labels: type (longitudinal, transverse, alligator, diagonal), orientation (horizontal, vertical, oblique), and severity (hairline, moderate, severe). The proposed model uses a hybrid feature extractor that combines EfficientNetV2 and Swin Transformer outputs, followed by three attribute-specific classification branches enhanced with attention mechanisms and trained using a weighted multi-task loss. A post-processing stage with calibrated probabilities, cross-attribute consistency checks, and a Composite Crack Risk Score (CCRS) supports maintenance prioritization. On a held-out test set, CrackNet-VOS attains accuracy/F1-score of 94.8%/94.8% for type, 92.3%/92.2% for orientation, and 90.7%/90.7% for severity, with Cohen’s Kappa values of at least 0.87. Compared with strong recent baselines, including EfficientNet B3, YOLOV10–ViT, and transformer-based multi-scale models, the proposed framework achieves gains of approximately 2–5% in F1-score across tasks. These findings indicate that combining a CNN–Transformer hybrid backbone with multi-task, attention-guided heads and consistency-aware post-processing yields a robust and scalable solution for automated, multi-attribute crack assessment in structural health monitoring.