Skip to content
Open access

Cross-Domain Evaluation of Modern Deep Learning Architectures for Microscopic Diatom Classification

2026 · IEEE Access · Vol 14, pp. 136634-136656 · 0 citations · 73 references

Abstract

Automatic classification of microscopic images is important for water-quality monitoring and ecosystem-health assessment, but existing deep-learning solutions are usually validated on single-source datasets and overlook the domain shift that arises across laboratories. This study evaluates how five representative deep-learning backbones (ResNet-50, EfficientNetV2-S, ConvNeXt-Tiny, Swin V2-Tiny and MaxViT-Tiny) behave under realistic cross-laboratory conditions in diatom classification. We assembled a Core Dataset of six heterogeneous public databases through a multi-stage curation protocol and evaluated each backbone under three protocols: Single-Source, Standard Mixed, and Leave-One-Source-Out (LOSO). In Single-Source training all five backbones exceeded 94 % accuracy. Under LOSO, accuracy fell by 33 to 47 percentage points (25 to 39 on a matched 27-genus label space) across the same set of models. ConvNeXt-Tiny attained the highest mean LOSO accuracy and the lowest cross-domain variance of all five backbones (standard deviation, SD = 5.80 %; about 64 % accuracy on LOIR, the hardest target on average), whereas MaxViT-Tiny attained the highest mean Macro-F1, while the classical ResNet-50 dropped to single-digit accuracy on the same domain. ConvNeXt-Tiny also had the shortest training time of the five backbones. Illustrative Grad-CAM examples are consistent with the backbones attending to the diatom frustule rather than to scale-bar and text artifacts. Two domain-generalization baselines evaluated under the same protocol show that style-based feature mixing (MixStyle) recovers part of the gap for the two evaluated convolutional backbones, whereas second-order feature alignment (CORAL) shows no consistent benefit. We characterize the cross-laboratory generalization gap, identify the inductive-bias families that are more resilient to it, and release a reproducible LOSO benchmark defined over six publicly available source databases as a common basis for future domain-generalization studies on diatom imagery.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.