Skip to content
Open access

Unified vision-centric pedestrian crossing action prediction via adaptive patch projection and proactive spatial rectification

Sep 2026 · Pattern Recognition · Vol 183, pp. 114888 · 0 citations · 52 references
Computer Science

TL;DR

ViCross is proposed, a vision-centric pedestrian crossing action prediction framework powered by multimodal large language models, which maintains target-centric reasoning from video frames without additional perception modules beyond first-frame target initialization.

Abstract

Vision cues are available and informative for pedestrian action prediction, but obtaining stable target-centric representations from video frames remains challenging without frame-level external perception cues. Thus, most methods rely on additional perception modules or multi-source information fusion, leaving the reliability of vision-centric setting an open question. To this end, we propose ViCross, a vision-centric pedestrian crossing action prediction framework powered by multimodal large language models, which maintains target-centric reasoning from video frames without additional perception modules beyond first-frame target initialization. While multimodal large language models exhibit strong visual understanding, applying them directly to vision-centric action prediction faces two challenges. First, accurately perceiving target pedestrians often requires high resolution inputs and dense visual tokenization, making full-frame encoding computationally prohibitive. ViCross tackles this with Variable Resolution Patch Mapping module for efficient token allocation while preserving key pedestrian details. Second, missing spatiotemporal priors hinder consistent cross frame reasoning. ViCross mitigates this with a Spatial Constraint Enhancement Strategy that captures past motion, future locations, and action semantics for training-time proactive spatial rectification. Extensive experiments show that ViCross delivers clear gains in vision-centric prediction settings and is competitive with multi-source fusion approaches in several settings. Code is available at https://github.com/2tianyao1/ViCross.git.

Read PDF

Similar papers

Preprint Aug 2026

Mask What Matters: Saliency-Guided Video Self-Supervised Learning for Autonomous Driving

V-JEPA4A is introduced, a domain-specialized variant of V-JEPA for autonomous driving that is pre-trained on publicly available driving videos with a novel saliency-driven masking policy that preserves and predicts context according to semantic importance and temporal relevance, yielding more informative representation...

Christopher Lang, Alexander Braun, Abhinav Valada · 0 citations
Sep 2026

Spatial-Semantic Attention Network With Adaptive Similarity Perception and Memory for Efficient Zero-Shot Object-Goal Visual Navigation.

This article investigates efficient zero-shot object-goal visual navigation, where an agent localizes unseen targets in novel environments using only visual observations and target representations. Existing methods suffer from an inability to establish a unified visual representation, leading to complex and inefficient...

Zi-Hao Dong, Hao Wan, Run-Min Cong et al. · 0 citations
Preprint Aug 2026

Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving

The proposed recurrent cross-frame module aggregates historical context from the previous state to enhance the coarse ground feature of each current frame and facilitates satellite candidate-region classification and hierarchical fine-grained features enable precise local offset estimation.

Jia-Ping Wang, Shaobo Li, Zhen Wang · 0 citations
Preprint Aug 2026

MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model

Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view map from multi-view images. Recent MVPD methods adopt a unified framework that projects 2D image features into a 3D world space and aggregates them into a single feature. Although they are effective, they struggle to gene...

Taiga Yamane, Satoshi Suzuki, Ryo Masumura et al. · 0 citations
May 2025

Object Concepts Emerge from Motion

Object concepts play a foundational role in human visual cognition, enabling perception, memory, and interaction in the physical world. Inspired by findings in developmental neuroscience - where infants are shown to acquire object understanding through observation of motion - we propose a biologically inspired framewor...

Hao Liang, Xiao-Hui Wang, Zhi-Chao Li et al. · 0 citations
Book Open access Sep 2026

Pedestrian Crossing Prediction for Automated Vehicles: A Geometry-aware Method Using Surface Normal

A geometry-aware framework that incorporates surface normal maps encoding pixel-level 3D surface orientation is proposed that incorporates surface normal maps encoding pixel-level 3D surface orientation and highlights the importance of geometry-informed features as an effective complement to conventional visual inputs.

A. Eskandari, Mahdi Rezaei, Mohsen Azarmi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.