Skip to content

Author

Ruqi Huang

We have 6 of 28 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation

Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-sp...

Zhen-Yu Liu, Zhang-Quan Chen, Ke-Yi Chen et al. · 0 citations
Preprint Sep 2026

EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception

Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, while isolated crops can lose the context needed to interpre...

Yao-Xin Niu, Zhangquan Chen, Yang Zhang et al. · 0 citations
Preprint Aug 2026

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only the representation across seven tokenizat...

Junhao Chen, Mingjin Chen, Jingjia Mao et al. · 0 citations
Jul 2026

Engine-Native Editable 3D World Reconstruction with Objects and Lighting

Editable 3D scene creation requires object instances and lights that can be inspected, moved, and imported into standard engines, yet existing single-image methods largely stop at room-scale geometry, baked/global illumination, or text-driven generation. We introduce Lumera (Light-aware Unified Engine-native Reconstruc...

Junhao Chen, Xinghao Chen, Henghaofan Zhang et al. · 2 citations
Jul 2026

Bunraku: Turning a Single Illustration into an Editable Live2D Character

This work presents the first system that, from a single illustration, generates all the structured information a Live2D runtime consumes: ordered RGBA layers, a deformation mesh per layer, and the parameter-driven keypose vertex offsets that make the character move.

Junhao Chen, Jingjia Mao, Dayong Li et al. · 0 citations
Preprint Jul 2026

Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh

D, a reference-guided renderer that extends Wan2.2 camera control from Plucker rays alone to a joint camera-plus-geometry interface and projects a neural 4D G-buffer from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior, supporting tracking+world-po...

Junhao Chen, Mingjin Chen, Henghaofan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.