Skip to content

Author

Qi-Feng Chen

We have 8 of 205 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes

Spatial representation learning for autonomous driving aims to map raw visual signals into structured 3D scene representations, where object-centric bounding boxes and rendering-oriented 3D primitives (\eg, 3D Gaussians) serve as two distinct yet highly complementary levels for scene understanding. Existing methods typ...

Li-Heng Chen, Haokai Pang, Cheng Su et al. · 0 citations
Jul 2026

IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment

This work proposes IQA-T1, a tool-based visual evidence reasoning framework that augments MLLM reasoning with explicit perceptual observations and constructs Q-Tool, a dataset containing 11k multimodal reasoning chains grounded in tool-generated evidence.

Jin-Jian Wu, Jiaqi Tang, Wei Wei et al. · 0 citations
Preprint Aug 2026

Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning

Remember-R1 is proposed, a reinforcement learning framework that mitigates long-context visual forgetting by applying process-level supervision directly on the original reasoning trajectory, demonstrating its effectiveness in mitigating long-context visual forgetting.

Jianmin Chen, Jiaqi Tang, Wei Wei et al. · 0 citations
Preprint Aug 2026

From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation

Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only training do not explicitly supervise the local structure needed for precise, temporally consistent editi...

Zhe-Fan Rao, Bin-Yi Zou, Xuanhua He et al. · 0 citations
Preprint Aug 2026

MSEditor: Toward Consistent Multi-Shot Video Editing

MSEditor is proposed, the first framework designed specifically for consistent multi-shot video editing, which significantly outperforms existing methods on the authors' curated multi-shot video editing benchmark in terms of identity preservation, temporal stability, and overall visual quality.

Kunyu Feng, Yue Ma, Bing-Yuan Wang et al. · 1 citation
Jul 2026

Data Pyramid for Embodied Manipulation

This work organizes the embodied data ecosystem as a pyramidspanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelit...

Yifan Ye, Yankai Fu, Ya-hui Lv et al. · 4 citations
Jul 2026

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

LeapBot-WA establishes a novel Predictive-Latent paradigm for WAMs by operationalizing the Joint-Embedding Predictive Architecture (JEPA) as a World-Anchor and introduces the Isotropic Semantic Autoencoder (ISAE), which reshapes the anchor's latent space into a diffusion-friendly manifold to prevent off-manifold drift.

Pei Liu, Nan Zheng, Lang Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.