Self-Supervised Multimodal Vision Transformer Framework for Robust Tomato Leaf Disease Detection under Field Conditions
DOI:
https://doi.org/10.70917/ijcisim-2026-3196Keywords:
unsupervised learning, Multimodal fusions of features, Vision Transformer, Plant diseases, Tomato leaf diseases, Domain adaptation, Few-shot learning, Precision agricultureAbstract
Prompt and precise identification of tomato leaf diseases are essential towards enhancing crop production and minimizing losses in precision farming. But current image-based disease detection algorithms are mostly based on fully supervised convolutional neural networks that are trained on controlled datasets and this restricts their resilience and extrapolation in real-field problems with complicated backgrounds, variances of illumination, and scarce labeled samples. In order to overcome these difficulties, this paper introduces a single deep learning model, which incorporates self-supervised representation learning, Vision Transformer (ViT) architecture, attention-based multimodal feature fusion, domain adaptation, and few-shot learning in terms of real-field tomato leaf disease detection.
The offered solution uses self-supervised pretraining to obtain discriminative visual representations using unlabeled field images, thus very minimally reliant on big labeled datasets. Adaptively fused multimodal features such as the visual RGB cues, texture as well as vegetation color indices are combined to improve disease discrimination in harsh environmental conditions through an attention mechanism. Domain adaptation is added to reduce the problems of dataset bias, and few-shot learning can be used to classify the categories of diseases with high precision with a small number of labeled samples.
The framework is tested on tomato leaf images in the real field and compared to the known CNN and transformer-based baselines in accordance with the accuracy, precision, recall, F1-score, and confusion matrix analysis. The experimental findings show that the suggested method performs better as far as robustness and generalization are concerned especially when the data are limited and the tasks are cross-domain. This is also the first research to unite self-supervised learning of Vision Transformers, multimodal fusion, and domain adaptation to detect tomato leaf disease in the real field in a systematic manner, making it well-suited to be deployed in real-life agriculture applications.