Skip to content

Author

Si-Jie Cao

We have 3 of 8 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

DriftingVLA: Native One-Step Vision-Language-Action Generation via Per-Dimension Temporal Drifting

Conventional flow-based vision-language-action (VLA) models support expressive continuous action generation but rely on multi-step refinement to produce each action chunk, increasing latency in online robot control. To address this issue, we introduce DriftingVLA, a native one-step VLA that generates a complete action...

Yuxuan Gao, Shi-Qi Zhang, Yedong Shen et al. · 2 citations
Preprint Sep 2026

Towards Active Cross-View Object Geo-Localization

Cross-view object geo-localization (CVOGL) typically assumes a fixed query image, overlooking the ability of mobile agents to actively acquire more informative observations. To address this limitation, we introduce Active Cross-View Object Geo-Localization (ActiveGeo), where an agent sequentially selects new viewpoints...

Shun-Yu Yao, Xiao-Han Zhang, Zhuo-Ran Yang et al. · 0 citations
Preprint Sep 2026

Multi-View Mixture-of-Experts with Vision-Language Reranking for Cross-View Object Geo-Localization

Cross-view object geo-localization (CVOGL) locates a target in satellite imagery using drone or street-view queries. Existing methods train separate detectors for each viewpoint, leading to parameter redundancy and impeding cross-view knowledge sharing. Moreover, top-ranked satellite candidates are often visually simil...

Xuezhi Fan, Ming Qi, Zhu Han et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.