Skip to content
Open access

STG: Structured Topology of Gridpoints for Occluded Pedestrian Detection

Aug 2026 · Italian National Conference on Sensors · Vol 26 · 0 citations · 63 references
Medicine

TL;DR

A novel Structured Topology of Gridpoints (STG) framework, operating strictly under standard full-box annotations without any extra visibility supervision, aims to achieve implicit, fine-grained local semantic compensation in pedestrian detection in crowds.

Abstract

Pedestrian detection in crowds is a challenging problem in computer vision. Existing occlusion-handling methods heavily rely on expensive visible-box annotations to locate visible body parts, posing severe limitations in label acquisition cost and open-world generalization. To break through this limitation, we propose a novel Structured Topology of Gridpoints (STG) framework. Operating strictly under standard full-box annotations without any extra visibility supervision, STG aims to achieve implicit, fine-grained local semantic compensation. Specifically, we formulate a coarse-to-fine reasoning paradigm consisting of three interactive stages. To mitigate the high spatial complexity and eliminate background redundancy, we first introduce a Saliency-Aware Feature Filtering (SAFF) mechanism, which leverages gridpoint heatmaps to filter out low-confidence pedestrian candidates. Second, a query-guided Across-Instance Feature Interaction (AIFI) model is designed to utilize inter-instance spatial relationships to propagate missing context from highly visible individuals to their occluded neighbors. Finally, we devise a prior-guided Inner-Instance Gridpoints Interaction (I2GI) model to achieve fine-grained structured part-level feature completion, which dynamically aggregates vital localized cues from diverse human parts to reconstruct holistic pedestrian representations. Extensive experiments on the CityPersons, CrowdHuman, and WiderPerson datasets demonstrate the effectiveness and efficiency of our proposed method. Specifically, STG achieves a log-average miss rate of 7.41% on Reasonable and 32.05% on Heavy Occlusion subsets of CityPersons, while running at up to 16 FPS, outperforming existing part-based methods under full-box supervision.

Read PDF

Similar papers

Preprint Sep 2026

DensePed-Lite: Quality-Aware Adaptive Detection for Dense Pedestrians under Occlusion

Pedestrian detection plays a crucial role in computer vision with applications in autonomous driving, surveillance, and public safety. However, real-world dense scenes bring severe challenges, including heavy occlusion, drastic scale variations, and strict real-time requirements. Existing lightweight detectors struggle...

Zian Wang, Ming-Zhe Liu, Chao-Yi Guo et al. · 0 citations
Open access Oct 2026

WalkScapes: A Pedestrian-View Urban Semantic Segmentation Dataset for Public-Space Monitoring and Autonomous Mobility

Pedestrian-level semantic segmentation supports autonomous mobility, assistive navigation, and safety-oriented perception in public spaces. However, widely used urban segmentation datasets predominantly adopt vehicle-centered viewpoints that differ from those in pedestrian environments in scene geometry, object scale,...

Kacper Chabros, Tomasz Jastrzębski, Patryk Oleś et al. · 0 citations
Preprint Aug 2026

GhostPoint: Self-Supervised Representation Learning by Hallucinating Occluded LiDAR Structure

GhostPoint is proposed, an SSL framework that hallucinates latent features in local neighborhoods around discovered instances, generated via a novel instance voxel dilation, and introduces a predictor-level supervision scheme on sampled voxels from generated neighborhoods.

Mohamed Abdelsamad, Bin Yang, Michael Ulrich et al. · 1 citation
Review Open access Aug 2026

Occlusion-Aware Image and Video Perception for Vulnerable Road User Detection: Methods, Benchmarks, and Deployment Challenges

Occlusion remains one of the main failure points in traffic-scene perception, and the errors it causes do not follow a single, predictable pattern. In a still frame, a camera may capture only a pedestrian’s head or upper torso. In video, a tracker can lose that person for several frames and assign a different identity...

J.-H. Feng, X. Zhang, Z.-L. Liu et al. · 0 citations
Open access Aug 2026

DOLD-Net: Dense occluded livestock detection via global-local feature collaboration

Precise livestock detection is fundamental to smart agriculture; however, real-world environments often present challenges such as high object density, severe occlusion, and ambiguous boundaries. To address feature integrity degradation under such conditions, this paper proposes DOLD-Net: a model for densely occluded l...

Kai-Da Jia, Yang Chai, Bao-Zhou Chen et al. · 0 citations
2026

RDAQ: Real-Time DETR Meets Refinement-Driven Adaptive Querying for Dense Aerial Imagery

Due to extreme density variations and scarce visual features, tiny-object detection in low-altitude uncrewed aerial vehicle (UAV) images remains a challenging task. Although DETR-based detectors have shown promising potential, they rely on a fixed number of object queries, which leads to suboptimal performance in dense...

Ling-Feng Lin, Bo-Xiang Xie, Wen Luo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.