Self-Supervised Multimodal Vision Transformer Framework for Robust Tomato Leaf Disease Detection under Field Conditions

Authors

  • Annagiri Amarnath Department of Computer Science, Bharatiya Engineering Science & Technology Innovation University, Gorantla, Andhra Pradesh, India.
  • S. K. Mahaboob Basha Department of Computer Science and Engineering, Sree Dattha Institute of Engineering and Science, Hyderabad, Telangana, India.

DOI:

https://doi.org/10.70917/ijcisim-2026-3196

Keywords:

unsupervised learning, Multimodal fusions of features, Vision Transformer, Plant diseases, Tomato leaf diseases, Domain adaptation, Few-shot learning, Precision agriculture

Abstract

Prompt and precise identification of tomato leaf diseases are essential towards enhancing crop production and minimizing losses in precision farming. But current image-based disease detection algorithms are mostly based on fully supervised convolutional neural networks that are trained on controlled datasets and this restricts their resilience and extrapolation in real-field problems with complicated backgrounds, variances of illumination, and scarce labeled samples. In order to overcome these difficulties, this paper introduces a single deep learning model, which incorporates self-supervised representation learning, Vision Transformer (ViT) architecture, attention-based multimodal feature fusion, domain adaptation, and few-shot learning in terms of real-field tomato leaf disease detection.
The offered solution uses self-supervised pretraining to obtain discriminative visual representations using unlabeled field images, thus very minimally reliant on big labeled datasets. Adaptively fused multimodal features such as the visual RGB cues, texture as well as vegetation color indices are combined to improve disease discrimination in harsh environmental conditions through an attention mechanism. Domain adaptation is added to reduce the problems of dataset bias, and few-shot learning can be used to classify the categories of diseases with high precision with a small number of labeled samples.
The framework is tested on tomato leaf images in the real field and compared to the known CNN and transformer-based baselines in accordance with the accuracy, precision, recall, F1-score, and confusion matrix analysis. The experimental findings show that the suggested method performs better as far as robustness and generalization are concerned especially when the data are limited and the tasks are cross-domain. This is also the first research to unite self-supervised learning of Vision Transformers, multimodal fusion, and domain adaptation to detect tomato leaf disease in the real field in a systematic manner, making it well-suited to be deployed in real-life agriculture applications.

Downloads

Download data is not yet available.

Downloads

Published

2026-07-06

How to Cite

Annagiri Amarnath, & S. K. Mahaboob Basha. (2026). Self-Supervised Multimodal Vision Transformer Framework for Robust Tomato Leaf Disease Detection under Field Conditions. International Journal of Computer Information Systems and Industrial Management Applications, 18(7s), 1284–1300. https://doi.org/10.70917/ijcisim-2026-3196

Issue

Section

Original Articles