Skip to content

CropCop: An Auditable 120-Class Plant-Health Model from Benchmark Reconstruction to a Quantised Runtime Artifact

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

This work presents CropCop, a closed-set recognition system spanning 120 operational plant-health classes and an evidence chain from corpus reconstruction to direct execution of the final quantised artifact that establishes strong leakage-controlled internal recognition and software-runtime fidelity.

Abstract

A plant-health score can appear precise while resting on duplicated image families, a long-tailed label space, or a runtime file that was never evaluated. We present CropCop, a closed-set recognition system spanning 120 operational plant-health classes and an evidence chain from corpus reconstruction to direct execution of the final quantised artifact. Starting from 117,546 audited images, we rejected the inherited partition after confirming 3,233 duplicate relationships across split boundaries and froze a 109,107-image benchmark with zero crossings among the audited trusted leakage groups and a 151.7 largest-to-smallest class ratio. A fully fine-tuned DINOv3 ConvNeXt-Tiny reference achieved 98.51% accuracy and 96.87% macro-F1 on the locked internal test. A compact MobileNetV4 Conv-Medium derivative achieved 98.46% accuracy and 96.27% macro-F1 without being presented as evidence for a new distillation method. Validation-only post-training quantisation selected dynamic activations with per-channel weights, and the final 22.60 MiB ExecuTorch/XNNPACK PTE achieved 98.46% accuracy and 96.23% macro-F1 when executed directly. Only six of 16,363 top-1 decisions changed between the converted INT8 graph and the PTE, while paired analysis showed a modest class-balanced loss; an exploratory post hoc fruit-label slice localized a larger recall decline than aggregate accuracy revealed. CropCop establishes strong leakage-controlled internal recognition and software-runtime fidelity; it does not establish performance on unseen farms, camera pipelines, or physical Android hardware.

View source

Similar papers

Open access Aug 2026

A Controlled Evaluation of Dual-Channel Feature Enhancement and Multi-Level Knowledge Distillation for Lightweight Plant Disease Recognition

Plant disease symptoms combine local texture changes with patterns distributed across a leaf, while practical recognition models must remain compact. We introduce DC-FEN, a MobileNetV3-based design that models spatial-token relations and channel interactions in parallel and injects them through gated residual fusion. We also examine output-distribution, direct-feature, and token-relation transfer under same-backbone and heterogeneous teachers. PlantVillage and Plant Pathology 2021 (FGVC8) are evaluated with duplicate-audited, group-aware 70/15/15 splits, an explicit unresolved-leaf sensitivity check, validation-only selection, five training seeds, class-sensitive metrics, and paired seed-wise descriptive summaries. On PlantVillage, the no-additional-attention student, DC-FEN teacher, and DC-FEN joint student obtain macro F1 scores of 96.46±0.91%, 96.90±0.40%, and 96.55±0.25%. On FGVC8, the corresponding scores are 87.29±0.63%, 87.14±0.52%, and 87.20±0.26%. At the prespecified FGVC8 threshold of 0.5, DCAB changed sample-wise F1 by −0.02±0.55 percentage points relative to the unmodified backbone; validation-selected global and label-specific thresholds changed this contrast to +0.28±0.55 and +0.55±0.29 points, while threshold-free macro mAP remained essentially unchanged. A duplicate-audited PlantDoc pressure test reduced frozen-checkpoint accuracy to 30.34±1.10% and 29.57±1.10%, showing that external generalization remains unestablished. A ResNet50 teacher gives logit-only students 97.42±0.51% macro F1 on PlantVillage and 89.82±0.43% sample-wise F1 on FGVC8. After separately weighting the direct and relation terms, the corresponding joint students obtain 97.37±0.56% and 89.94±0.27%, recovering the degradation seen with unit internal weights while remaining close to logit-only transfer. Thus, the study evaluates the benefits and limits of explicit spatial–channel interaction and shows that adding intermediate transfer constraints does not guarantee a stronger student.

Xin Lei, Yonghuai Liu, Ardhendu Behera et al. · 0 citations
Preprint Aug 2026

Shortcut Learning in a Public Grape Disease Dataset: Annotation Granularity as a Modulator, Not a Cause

Public datasets for agricultural disease detection are usually judged fit for use from reported metrics, which say nothing about whether the annotation scheme is internally consistent. On one public grape disease dataset (3288 images, 11995 boxes, 6 classes), varying model capacity, input resolution and detection paradigm yields a test-set mAP50 range comparable to seed-to-seed noise, with the bottleneck at small objects across all five architectures. The finding lies on the data side: one class is annotated at whole-leaf level (median box area 43.16% of the image) while the other five are annotated at lesion level. On 5156 cross-species images containing no grape, 65.7% of the false-positive boxes fall into that one class, an over-representation of 13.41x relative to its share of the training annotations. Counterfactual retraining establishes a causal effect of granularity on the magnitude of the shortcut: shrinking only that class's boxes cuts its cross-species false positives by 66%, and a placebo control confirms the effect is specific to the manipulated class. A manipulation in the opposite direction, with criteria registered in advance, returns a negative result: coarsening the finest class to whole-leaf level (0.57% to 40.37%), matched in box count and share of annotations and with higher in-distribution AP, still leaves its cross-species false positives at zero boxes, while the unmanipulated original class holds 50.0% of them. Annotation granularity is therefore a modulator of this shortcut, not its cause: it can amplify or attenuate a sink that already exists, but cannot create one, and what fixes the destination remains open. We also give a granularity screening statistic requiring neither images nor training, and show airborne lesion-level detection to be optically out of reach. The failure mode is invisible to in-distribution evaluation.

Pu Wang · 0 citations
Conference Jul 2026

Lightweight global-local dual-branch fusion for representative multicrop disease recognition

Multi-crop disease recognition becomes difficult when visually similar lesion patterns must be identified under a tight parameter budget. This paper reports a compact conference-scale study for the computer-vision and machine-learning track of MLES 2026. A representative 14-class subset covering tomato, cucumber, grape, and apple was constructed from 13,205 images, including 10,272 training images and 2,933 held-out evaluation images. The proposed network couples a lightweight global branch, implemented by a shallow CNN stem followed by a Transformer encoder, with a MobileNetV3- Small local branch for texture-sensitive feature extraction. A learned gating head projects and adaptively fuses global and local evidence before classification. On a single RTX 3060 GPU, the model achieved 99.35% Top-1 accuracy and 100.00% Top-5 accuracy, with macro precision, recall, F1-score, and specificity of 99.36%, 99.39%, 99.37%, and 99.95%, respectively. The model uses only 2.17M parameters, indicating that accurate and deployable visual recognition is possible with a compact dual-branch design. To address reviewer concerns on robustness and component attribution, the revised manuscript additionally reports five-fold cross-validation statistics, single-branch baselines, augmentation ablations, and a freezing-strategy study.

Yang Zhang, Rongrong Gu, Chengyuan Li et al. · 0 citations
Open access Aug 2026

Leakage-Free Benchmarking of Electronic Noses for Beef Freshness: A Signal-Richness Criterion for Model Selection

Low-cost metal-oxide-semiconductor (MOS) electronic noses promise rapid, non-destructive meat freshness screening, and published classifiers frequently approach perfect accuracy. Such figures are rarely tested against the two conditions that most inflate them: a target-derived label among the inputs, and random splitting of the correlated samples. Beef freshness is benchmarked here on a public 11-sensor, 12-cut MOS dataset using leakage-free leave-one-cut-out cross-validation in order to predict freshness class and total viable count (TVC) with paired significance tests. A gradient-boosted-tree pipeline is the strongest model (accuracy 0.81±0.10, macro-F1 0.68±0.15, TVC R2=0.77), significantly outperforming a multi-scale attention convolutional network (macro-F1 0.50±0.15; p<0.001). The advantage of this study lies in the representation, not the model family: a network given the same window summaries reaches 0.64±0.17, indistinguishable from the tree. Near-perfect accuracy returns only when TVC is supplied as a feature or samples are split at random (macro-F1 0.97). Under nested, per-fold selection, a five-sensor subset matches the full array. On a rich BME688 heater profile dataset, the network surpasses the tree, an advantage that vanishes as the profile shortens to one step. Evaluation and representation, not architecture, govern reported performance; a signal-richness criterion predicts when a deep temporal model is justified.

Erkan Caner Ozkat · 0 citations
Preprint Aug 2026

BoltNet: An Ultra-Lightweight Convolutional Network for On-Device Plant Species Identification

Automated plant species identification from citizen-science imagery is an established, demanding fine-grained recognition problem: large taxonomic label spaces, visually similar species, and long-tailed observations require real model capacity, while field use constrains memory, latency, and power. Model size is only part of the deployment cost: intermediate activations held in memory during inference and platformdependent execution behavior matter too, so compact recognition must be assessed on target hardware rather than through complexity metrics alone. We present BoltNet, an ultra-lightweight fully convolutional architecture combining a Spatial Redistribution Bottleneck and Logit PreSampling to improve the tradeoff between predictive performance and model size in high-cardinality classification, and report the AccuracyCompression Tradeoff as a complementary diagnostic. On Pl@ntNet300K, BoltNet reaches 0.682 F1-score with 341K parameters (1.37 MB), the highest F1-score among evaluated models below 2 MB and close to substantially larger convolutional backbones. Model-only measurements on a Raspberry Pi 5, Jetson Orin Nano, and Hailo-8 characterize execution across CPU, GPU, and NPU platforms, where BoltNet is the most consistently efficient model, with the best FPS/W on the GPU and NPU and second-best on the CPU. Results on AIDERv2 and CLRS provide secondary evidence of transfer across environmental image-classification tasks. Code available at: https://codeberg.org/danielrossi/BoltNet

Daniel Rossi, G. Borghi, R. Vezzani · 0 citations
Open access Aug 2026

PalmNet: Confidence-Calibrated Edge-Cloud AI for Field Diagnosis of Date Palm Diseases

Date palm (Phoenix dactylifera L.) is a cornerstone crop for Iraq and the wider MENA region, yet reliable in-field diagnosis of leaf disorders remains slow, labour-intensive, and constrained by a limited pool of agronomists. This paper presents PalmNet, a full-stack diagnostic system that classifies nine leaf conditions through a calibrated edge-cloud framework. The system is developed and evaluated on a public dataset of 3,089 field images spanning the nine classes, using a 70/15/15 stratified split. A ShuffleNetV2 student network, distilled from a ConvNeXt-Tiny teacher, is deployed on two complementary edge endpoints: a Raspberry Pi Zero 2 W field station running ONNX Runtime with GPS-tagged capture, and an Android application built in Kotlin with Jetpack Compose and TensorFlow Lite, exposing separate viewer and expert interfaces. The teacher is served from Google Cloud Run and is invoked only for low-confidence predictions. Post-hoc temperature scaling with a single calibration temperature (Tcal = 1.3976) is applied to the student logits, and a calibration-split threshold sweep selects the operating confidence threshold τ = 0.93 for selective offloading. On a held-out test set of 464 images, PalmNet reaches 99.14% top-1 accuracy, a macro-F1 of 0.9789, and a balanced accuracy of 97.33%, on a par with a fully cloud-based baseline while keeping 93.75% of inferences on-device and reducing uplink traffic by approximately sixteenfold (from about 12.3 MB to 0.77 MB across the test pass). Knowledge distillation improves the student macro-F1 by 1.98 percentage points over a non-distilled baseline, and the deployed student carries roughly ten times fewer parameters and FLOPs than the teacher while running in about 93 ms per image on the field station. Bootstrap confidence intervals indicate that PalmNet is statistically on a par with the cloud-only baseline, and a robustness analysis under degraded captures shows that the calibrated router escalates more cases to the cloud as input quality declines. A Firebase-based expert feedback pipeline enables validation and continuous dataset enrichment with real field samples. These results show that coupling calibrated edge inference, selective cloud assistance, and expert-in-the-loop validation yields a practical solution for in-field palm disease diagnosis under the bandwidth and staffing constraints of real-world deployment.

Muntadher Kareem, Raed J. Majeed · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.