Results indicate that combining foundation-model semantic priors with a lightweight detector can improve the reliability of maritime object detection under complex sea-surface conditions.
Abstract
Maritime object detection remains challenging because of complex sea-surface backgrounds, adverse illumination conditions, and severe class imbalance, especially when safety-critical targets such as search-and-rescue vessels are sparsely represented. To address these challenges, we propose a cascaded hybrid DINOv2-YOLOv8n detection framework for maritime scenes. Rather than relying only on supervised learning from raw RGB images, the proposed method introduces semantic priors from a frozen DINOv2 encoder and projects them into a compact representation for a YOLOv8n-based detector. To improve robustness under diverse maritime conditions, the framework uses a sea-state-aware online augmentation strategy and is trained with the standard YOLO detection objective. Experiments on the Maritime Target Data Sharing Project (MTDSP) dataset show that the proposed framework achieves strong detection performance. Specifically, it obtains an overall mAP@0.5 of 89.4% and a precision of 100% on the test set. For sparse search-and-rescue vessels and rigid offshore structures, it achieves mAP@0.5 scores of 99.5% and 96.3%, respectively. These results indicate that combining foundation-model semantic priors with a lightweight detector can improve the reliability of maritime object detection under complex sea-surface conditions.
The proposed framework shows potential to address the gap in practical implementation through reconstruction-based feature learning accompanied by a specified geometric baseline, and shows promise for real-time maritime surveillance applications, though the need for additional thorough verification across many operatio...
M. M. H. Khan, Qi-Wei Hu, Radhakrishna Prabhu et al.· Journal of Imaging· 0 citations
A novel model is introduced by assessing the impacts of several YOLO object detection algorithms with the Convolutional Block Attention Module (CBAM) on aircraft detection from satellite images to demonstrate that attention mechanisms have a significant impact when used with the YOLO architecture for object detection i...
Ibrahim Aruk, Hakan Açıkgöz, Ertuğrul Doğruluk· Konya Journal of Engineering...· 0 citations
Maritime ship detection remains challenging because of large scale variations, high inter-class visual similarity, weak target boundaries, and complex maritime backgrounds. This study proposes MCSwin-YOLOv8, an enhanced YOLOv8-based detector that combines three complementary architectural designs. First, a re-parameter...
Water surface target detection for autonomous rescue USVs faces significant challenges due to complex lighting conditions, reflections, and small partially submerged targets, while edge deployment constraints demand low computational cost. To address these issues, this paper proposes a Lightweight Attention-Enhanced YO...
Yue-Hao Xiong, Bo-Wen Sun, Yu-Ting Yang et al.· 2026 6th International Confe...· 0 citations
Landslide detection is essential for geological disaster mitigation, yet existing deep learning methods still struggle with the high cost of global feature modeling, limited receptive fields in window-based attention, and insufficient fusion of local and global information. To address these challenges, we propose MTTNe...
Jia-Xin Song, Shu-Wen Yang, Hao Zhu et al.· IEEE Geoscience and Remote S...· 0 citations
TRIDEN-YOLO, a lightweight detector built upon YOLOv11n, provides the primary reparameterized contextual representation design through multi-branch training and inference-time fusion, while HFFE and GCD loss are incorporated to enhance hierarchical feature fusion and boundary-aware localization.
Xi Chen, Yuping Sun, Kaibin Zeng· Signal, Image and Video Proc...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.