Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have d...
Yuheng Shi, Xiao-Huan Pei, Min-Jing Dong et al.· 0 citations
SlackDrive is proposed, a pre-inference compute allocator that reuses realized latency to select the compute budget of each control step before model execution, complementing existing profiling and resource scheduling while preserving the driving backbone and its compute actuator.
Xiao-Huan Pei, Heng-Guang Zhou, Yuan-Hao Ban et al.· 0 citations
This work replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards, and adds a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts.
Yuan-Hao Ban, Jia-Qi Feng, Heng-Guang Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.