ViCross is proposed, a vision-centric pedestrian crossing action prediction framework powered by multimodal large language models, which maintains target-centric reasoning from video frames without additional perception modules beyond first-frame target initialization.
Abstract
Vision cues are available and informative for pedestrian action prediction, but obtaining stable target-centric representations from video frames remains challenging without frame-level external perception cues. Thus, most methods rely on additional perception modules or multi-source information fusion, leaving the reliability of vision-centric setting an open question. To this end, we propose ViCross, a vision-centric pedestrian crossing action prediction framework powered by multimodal large language models, which maintains target-centric reasoning from video frames without additional perception modules beyond first-frame target initialization. While multimodal large language models exhibit strong visual understanding, applying them directly to vision-centric action prediction faces two challenges. First, accurately perceiving target pedestrians often requires high resolution inputs and dense visual tokenization, making full-frame encoding computationally prohibitive. ViCross tackles this with Variable Resolution Patch Mapping module for efficient token allocation while preserving key pedestrian details. Second, missing spatiotemporal priors hinder consistent cross frame reasoning. ViCross mitigates this with a Spatial Constraint Enhancement Strategy that captures past motion, future locations, and action semantics for training-time proactive spatial rectification. Extensive experiments show that ViCross delivers clear gains in vision-centric prediction settings and is competitive with multi-source fusion approaches in several settings. Code is available at https://github.com/2tianyao1/ViCross.git.
V-JEPA4A is introduced, a domain-specialized variant of V-JEPA for autonomous driving that is pre-trained on publicly available driving videos with a novel saliency-driven masking policy that preserves and predicts context according to semantic importance and temporal relevance, yielding more informative representation...
Christopher Lang, Alexander Braun, Abhinav Valada· 0 citations
This article investigates efficient zero-shot object-goal visual navigation, where an agent localizes unseen targets in novel environments using only visual observations and target representations. Existing methods suffer from an inability to establish a unified visual representation, leading to complex and inefficient...
Zi-Hao Dong, Hao Wan, Run-Min Cong et al.· IEEE Transactions on Neural...· 0 citations
The proposed recurrent cross-frame module aggregates historical context from the previous state to enhance the coarse ground feature of each current frame and facilitates satellite candidate-region classification and hierarchical fine-grained features enable precise local offset estimation.
Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view map from multi-view images. Recent MVPD methods adopt a unified framework that projects 2D image features into a 3D world space and aggregates them into a single feature. Although they are effective, they struggle to gene...
Taiga Yamane, Satoshi Suzuki, Ryo Masumura et al.· 0 citations
Object concepts play a foundational role in human visual cognition, enabling perception, memory, and interaction in the physical world. Inspired by findings in developmental neuroscience - where infants are shown to acquire object understanding through observation of motion - we propose a biologically inspired framewor...
Hao Liang, Xiao-Hui Wang, Zhi-Chao Li et al.· Neural Information Processin...· 0 citations
A geometry-aware framework that incorporates surface normal maps encoding pixel-level 3D surface orientation is proposed that incorporates surface normal maps encoding pixel-level 3D surface orientation and highlights the importance of geometry-informed features as an effective complement to conventional visual inputs.
A. Eskandari, Mahdi Rezaei, Mohsen Azarmi et al.· Proceedings of the 18th Inte...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.