Skip to content

Author

Xiantao Zhang

We have 8 of 9 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Dense to MoE Adaptation for Compact Vision Language Action Policies

Vision language action (VLA) policies continue to grow in parameter count, making deployment on resource-constrained robot platforms difficult. The central goal is to reduce the number of LLM-side parameters retained in the deployed policy while preserving downstream task performance. Our approach, AdaDE, adapts select...

Mu-Chun Niu, Shuang Chen, Yu-Zhou Wu et al. · 0 citations
Preprint Aug 2026

STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models

On-policy distillation (OPD) has become an effective approach for consolidating multiple task-specialized image generation models into a single student. However, existing OPD methods optimize the student mainly to match the teacher's output velocity, making the teacher the upper limit of the optimization objective. Whi...

Qing-Yan Wei, Guang-Zhao Li, Xiao-Bing Tu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Behavior Quotient Learning for Low-Rank Adaptation of LLM Agents

LLM-based agents rely on heterogeneous interaction capabilities to accomplish complex tasks. Existing approaches often distribute these capabilities across multiple LoRA adapters, which increases adapter storage requirements and introduces routing overhead during inference. A single LoRA avoids this overhead, but learn...

Peng-Yang Zhou, Xiao-Bing Tu, Zheng-Xi Liu et al. · 0 citations
Preprint Aug 2026

Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents

Browser agents perform well on short, clean demonstrations, but real deployment is fundamentally different: agents must sustain dozens of decisions on live websites while recovering from mistakes and navigating complex UIs. We argue that closing this gap requires alignment at every level of the pipeline, including exec...

A. Chen, Yan Cheng, Zhangye Han et al. · 1 citation
Preprint Aug 2026

AViTS: Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation

Diffusion Transformers (DiTs) achieve high-quality generation but are costly due to iterative sampling. Dynamic-resolution sampling reduces early-stage cost by denoising at low resolution; however, uniformly upsampling all latent tokens at resolution transitions incurs redundant computation and may degrade fine-detail...

Haoran Qin, Zhen Yan, Shikang Zheng et al. · 0 citations
#machine learning Preprint Sep 2026

Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache

Diffusion Transformers have become the dominant paradigm in generative AI, but their high computational costs severely hinder real-time applications. Prediction-based feature caching is widely used to accelerate diffusion transformers; however, as the number of steps increases, the deviation between its predictions and...

Zhi-Rong Shen, Rui-Xin Huang, Chang Zou et al. · 0 citations
Preprint Aug 2026

LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching

LinCa decomposes cached features into sub-components with distinct continuity properties via a lightweight invertible network and applies differentiated prediction orders matched to each component, forming a unified Decompose-Predict-Reconstruct pipeline.

Jin-Shan Liu, Hao-Ran Qin, Xiao-Bing Tu et al. · 0 citations
Jul 2026

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

Method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data, is presented, demonstrating that the synthesized data substantially improve downstream WAM generalization.

Zexuan Yan, Yuzhou Wu, Yue Ma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.