PhytoFormer: A Robust Plant Disease Recognition for Crop Protection Using Multi-Branch Hybrid Attention Model
Abstract
Timely and accurate plant disease detection is essential for sustainable agriculture and global food security. However, existing deep learning approaches still face challenges in recognizing diseases across different plant structures and imaging conditions due to variations in scale, appearance, illumination, and background complexity. While convolutional neural networks (CNNs) effectively capture local visual patterns, they are limited in modeling long-range dependencies. In contrast, Transformer-based models often incur high computational costs and require large-scale training data, which limits their deployment in resource-constrained agricultural environments. To address these limitations, this paper proposes PhytoFormer, an efficient hybrid architecture that integrates a Hierarchical Attention Feature Extractor (HAFE) with a lightweight Hybrid Vision Transformer (HVT) to enhance multi-scale feature representation, suppress irrelevant background information, and capture global contextual relationships with minimal computational overhead. PhytoFormer was evaluated on three datasets, each having unique imaging characteristics: SugarLeaf-IDN (field images), PlantVillage (controlled laboratory images), and PlantDoc (real-world images). The model provided classification accuracies of 94.11%, 99.89%, and 81.15%, respectively, and outperformed the strongest baseline models by 0.55, 0.09, and 1.26 percentage points. Also, with its lightweight architecture, only 3.96 million parameters, 0.28 GFLOPs, PhytoFormer balances classification accuracy, robustness, and computational efficiency. These findings support its potential for practical deployment in intelligent plant disease diagnosis and precision agriculture under resource-constrained conditions.