Skip to content
Open access

MCSwin-YOLOv8: Multi-Scale Feature Learning for Maritime Ship Detection

Aug 2026 · Applied Sciences · 0 citations · 22 references

Abstract

Maritime ship detection remains challenging because of large scale variations, high inter-class visual similarity, weak target boundaries, and complex maritime backgrounds. This study proposes MCSwin-YOLOv8, an enhanced YOLOv8-based detector that combines three complementary architectural designs. First, a re-parameterizable multi-scale convolutional backbone, named RepMCSwin, is introduced to extract scale-aware semantic information and fine-grained boundary cues. Unlike the standard Swin Transformer, the MCSwin block does not use window-based self-attention but adopts cascaded multi-scale convolutions and residual feature transformation. Second, a Multi-Feature Parallel Convolutional Block Attention Module (MFPCBAM) is developed to compute channel and spatial attention in parallel, thereby preserving weak ship features while suppressing irrelevant background responses. Third, a Modified Generalized Feature Pyramid Network (MGFPN) is constructed to improve cross-level feature interaction and retain high-resolution spatial information through an additional 160 × 160 prediction branch. Experiments were conducted on the public SeaShips dataset and a private infrared maritime ship dataset. MCSwin-YOLOv8 achieved an F1-score of 94.4%, mAP@0.5 of 97.4%, and mAP@0.5:0.95 of 75.1% on SeaShips. On the infrared dataset, the corresponding results were 91.2%, 94.1%, and 66.9%, respectively. Compared with the YOLOv8 baseline, mAP@0.5 increased by 1.4 and 3.0 percentage points on the two datasets. These accuracy gains were accompanied by an increase in model complexity from 11.12 M to 19.29 M parameters and from 28.5 G to 56.7 G FLOPs, indicating an accuracy–complexity trade-off that requires further runtime evaluation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.