Aug 2026· International Conference on Multimedia Analysis and Pattern Recognition· pp. 139-144· 0 citations· 23 references
Abstract
Current object detection models face significant challenges when applied to ultra-high-resolution images, as global downscaling frequently destroys crucial details for small objects, while EGC wastes computation on irrelevant regions and causes object truncation at tile boundaries. We propose a two-step reasoning architecture that employs a Guide Model to identify semantically relevant spatial regions, followed by fine-grained detection on the selected high-resolution crops. We formalize EGC with a dynamic partitioning scheme to explicitly define its spatial coverage boundaries and establish an upper bound for spatial recall. Building on this formulation, we introduce a semantic alternative consisting of two Guide Model variants, Qwen-VL and YOLO+LLM, together with a multi-stage hybrid Detector Model integrating GroundingDINO, YOLOv8, and CLIP for verification. Experiments on 3,637 UHR images demonstrate that although grid cropping achieves the highest detection accuracy with an improvement of up to 26.9% mAP over the full-image baseline, our Guide Models recover most of this gain, improving mAP by 19.2% over the baseline while reducing the required number of crops by a factor of 3.7 and effectively mitigating false-positive over-detection.
Detecting small targets in remote sensing imagery remains a critical challenge due to the inherent limitations of low-resolution targets and the frequent clustering of small targets. To address these challenges, we propose an integrated set of improvements to the You Only Look Once (YOLO)v8 framework, leveraging cross-...
Chun-Zhi Li, Sheng-Yan Su, Xiao-Hua Chen et al.· IEEE Transactions on Geoscie...· 0 citations
The rapid development of unmanned aerial vehicle (UAV) technology has made aerial-image object detection increasingly important for natural-resource monitoring, traffic management, and disaster response. Detecting small objects in aerial images remains difficult because objects occupy very few pixels, high-frequency cu...
This work proposes MSGAN, a multi-scale global-local collaborative learning framework that integrates multi-scale mixed convolution and adaptive global-local attention to enhance feature representation and advances robust SOD for complex real-world scenarios and provides insights into attention-guided visual perception...
State-of-the-art vision models process images in their entirety, lacking the ability to selectively zoom in on relevant regions. This limitation is particularly acute in scenarios where processing must be conditioned on a specific task - such as instance detection, which requires localizing a specific object in a high-...
Oleh Kolner, Thomas Ortner, Stanisław Woźniak et al.· 1 citation
This work proposes SmartRes, a framework that performs efficiency optimization in the pixel space via dynamic resolution routing and introduces a margin-regularized routing objective that increases foreground-background logit separation and improves foreground recall.
H. Sun, Wang-Bo Zhao, Fanyue Wei et al.· 0 citations
Small object detection in traffic monitoring suffers from a structural inefficiency in standard detectors: isotropic 3×3 convolutions treat all spatial directions uniformly, yet traffic objects exhibit strong anisotropic geometry—pedestrians are vertically elongated, vehicles are horizontally wide. We propose MSHC-YOLO...
Guoshun Cui, Yi-Cai Zhang, Xin-Yan Huang et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.