Jun 2026· 2026 Progress in Applied Electrical Engineering (PAEE)· pp. 1-8· 0 citations· 13 references
Computer ScienceEngineering
TL;DR
A two-stage geometry-aware localization pipeline that estimates the projection of the vehicle footprint onto the road plane instead of relying directly on detector geometry and demonstrates clear improvements in localization accuracy compared with naive bounding-box-center-based localization.
Abstract
Accurate vehicle localization from monocular roadside surveillance cameras is an important problem in intelligent transportation systems, traffic monitoring, and traffic conflict analysis. Standard localization approaches typically estimate vehicle position using the center of the detector bounding box, which may lead to large localization errors due to perspective distortion and parallax effects, particularly for elevated roadside cameras and large vehicles.This paper proposes a two-stage geometry-aware localization pipeline that estimates the projection of the vehicle footprint onto the road plane instead of relying directly on detector geometry. In the first stage, vehicles are detected using a YOLO26-based detector. In the second stage, a dedicated ResNet34 regression network predicts four corner points corresponding to the projection of the vehicle base onto the image plane. The final vehicle position is estimated as the geometric center of the predicted quadrilateral.The proposed method was trained using synthetic data generated in the CARLA simulation environment and subsequently fine-tuned on real-world roadside imagery from the DAIR-V2X dataset. Experimental evaluation performed on both synthetic and real-world data demonstrated clear improvements in localization accuracy compared with naive bounding-box-center-based localization. On the DAIR-V2X dataset, the proposed approach reduced the mean image-space localization error from 31.77 px to 15.30 px (51.8% improvement) and the median error to 4.29 px. Median ground-plane localization error for medium-range vehicles decreased from 5.52 m to 0.90 m, while for far-range vehicles it decreased from 8.67 m to 1.84 m.The experiments additionally demonstrated that contextual information surrounding the detector bounding box plays an important role in geometric localization. The largest improvements were observed for distant vehicles and geometrically challenging cases affected by strong perspective distortion and parallax effects.
This paper presents a one-stage learning framework that maps monocular roadside-camera images directly to vehicle states in a ground-fixed coordinate frame. Unlike conventional approaches that first detect vehicles in the image plane and subsequently apply geometric post-processing, the proposed method leverages featur...
Akos T. Kopeczi-Bocz, Tian Mi, Gábor Orosz et al.· 0 citations
The lack of scalable and cost-effective methods for extracting actionable vehicle trajectories from existing traffic closed-circuit television (CCTV) infrastructure limits proactive traffic safety analysis. Traditional trajectory estimation approaches often rely on LiDAR, radar, or calibrated camera systems, which are...
Swaranjit Roy, Ahmed Abdelhadi, Sherif M. Gaweesh· Transportation Research Reco...· 0 citations
A target-free wide-area calibration method for roadside LiDAR-camera systems that estimates extrinsic parameters directly from natural traffic scenes, providing higher calibration accuracy while preserving practical computational efficiency and showing good robustness under challenging conditions.
Findings indicate that decision-level fusion provides scenario-dependent benefits rather than automatic improvement over a strong single-sensor baseline, and AEKF achieves small gains over the LiDAR-only baseline, and object-level connected vehicle observations remain useful when shared at reduced update rates.
Aleksi Pippuri, N. Jayawickrama, Risto Ojala· 0 citations
Accurate multi-object tracking and metric localization support traffic monitoring and cooperative intelligent transportation at complex intersections. This study presents a fixed-camera vision-map fusion framework that addresses two practical difficulties: axis-aligned boxes poorly represent turning vehicles, and uncon...
This paper proposes a calibration-free framework for reliably and effectively estimating vehicle speeds from monocular videos, without relying on roadway features, camera calibration, or roadway-feature-based reference objects. The proposed framework estimates vehicle speeds using a 36-keypoint vehicle template and a h...
Gaofeng Su, Ke-Ya Li, Raja Sengupta et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.