DriveCache is proposed, a training-free, action-aware controller that uses planned motion to allocate reuse across scenes and dynamic programming to place it across denoising steps under a calibrated response budget, which improves the overall fidelity-efficiency trade-off over evaluated cache methods.
Jianchun Yang, Jian Liang, Xian-Da Guo et al.· 0 citations
OnPoKD is the first framework that applies on-policy distillation to vision-language model adaptation by learning target construction as a policy decision, and is the first framework that applies on-policy distillation to vision-language model adaptation by learning target construction as a policy decision.
Hong-Yuan Zhang, Xian-Da Guo, Yan-Lun Peng et al.· Information Fusion· 0 citations
VGGD, a visual geometry foundation-aware 3D Gaussian Splatting framework for feed-forward surround-view driving reconstruction, which shifts geometric modeling to the frontend and adapts foundation priors to the driving camera setting and achieves the best overall rendering quality among the compared methods and improv...
Junhong Lin, Jinlong Wang, Xianda Guo et al.· 0 citations
InstructVVT is proposed, an instruction-driven and reference-guided video virtual try-on framework based on a Diffusion Transformer that operates without inference-time spatial priors that outperforms state-of-the-art open-source methods in garment fidelity, structural preservation, and temporal consistency, despite re...
Di Shao, Song-Han Wu, Xin-Yu Chen et al.· 0 citations
The core of SSVAL is Visual Anchor Prompt Injection (VAPI), which introduces prompts that absorb rich knowledge from external VFMs during training, enabling them to serve as stable visual anchors that mitigate representation deviation during inference.
Qian-Long Yang, Bowen Ye, Xianda Guo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.