A visual–semantic multimodal fusion framework is developed, incorporating a Two-Stage Modality Fusion mechanism that projects semantic and visual features into a unified feature space to optimize cross-modal feature interaction.
Abstract
Identifying plant diseases presents challenges like small sample sizes, imbalanced classes, and intricate background noise, which constrain the generalizability and resilience of conventional deep learning methods that rely on single-modal images. To overcome these limitations, this study introduces a novel approach for recognizing plant diseases with limited examples, utilizing multiple modalities that combine visual and textual data to enhance the model’s discriminative power and generalization performance. First, the official definition of plant diseases is semantically enriched using large language models to automatically generate high-quality textual descriptions that include symptoms, affected plant parts, and agricultural context characteristics. Second, a visual–semantic multimodal fusion framework is developed, incorporating a Two-Stage Modality Fusion mechanism that projects semantic and visual features into a unified feature space to optimize cross-modal feature interaction. The experiments were performed on the PlantVillage dataset and the PlantDoc dataset to validate the N-way K-shot task. Results show that the proposed method achieved accuracies of 79.01% and 90.51% under the 5-way 1-shot and 5-way 5-shot tasks on the PlantVillage dataset, and 59.82% and 70.61% on the PlantDoc dataset. These results outperformed traditional unimodal and existing multimodal approaches. This paper provides an efficient and widely applicable multimodal approach to few-shot plant disease identification in agriculture.
This work proposes EfficientNet-CBAM-Prototype (ECP-Net), a unique end-to-end deep learning architecture that achieves 98.6% test accuracy, 98.59% F1-score, and 98.65% precision on the PlantVillage tomato subset, and proposes a two-phase training strategy wherein CBAM is frozen during phase 1 to allow prototype stabili...
E. Jansi, Kavitha Br· Frontiers in Plant Science· 0 citations
A deep learning model designed to automatically detect grape leaf diseases based on images, using a pretrained ResNet50 which is trained on ImageNet as feature extractor and a Convolutional Block Attention Module to boost its discriminative capacity is introduced.
Maajid Bashir, A. Reshi, Shabana Shafi et al.· International Journal of Mac...· 0 citations
Accurate and timely identification of plant diseases from leaf images remains a critical challenge in precision agriculture, particularly when diseases manifest with spatially disjoint symptoms and subtle textural variations. We propose a hybrid deep learning framework that synergistically combines convolutional neural...
Rishabh Aryan, Anju· Journal of Machine Learning...· 0 citations
Timely and accurate plant disease detection is essential for sustainable agriculture and global food security. However, existing deep learning approaches still face challenges in recognizing diseases across different plant structures and imaging conditions due to variations in scale, appearance, illumination, and backg...
Betty Dewi Puspasari, I-Cheng Chang, Andy Pramono· IEEE Access· 0 citations
A decade of progress across four interconnected frontiers is synthesizes the evolution of deep learning architectures for plant disease detection, the adaptation of foundation models including CLIP, SAM, and DINOv2 to agricultural contexts, and the development of multimodal fusion frameworks integrating imagery, enviro...
A unique, computationally efficient triple-feature block network capable of highly accurate plant disease classification across diverse species and complex imaging environments is proposed.
A. Elkholy, N. Elshennawy, Ahmed M. Gab Allah· Journal of King Saud Univers...· 0 citations