Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior-guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic sema...
Yi Shi, Huichao Xie, Yuqing Wang et al.· 0 citations
This work introduces Universal Referring, a generalized UAV referring task that jointly expands the query modality and the output cardinality, and presents UAV-URNet, a detection-style baseline that maps heterogeneous queries into a shared query space and predicts variable-size target sets through set prediction.
The sample efficiency and scalability of RL post-training for video MLLMs and introduces OraRL, a decoupled advantage estimator that scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts.
Yunheng Li, Guo-Hong Mu, Hao Li et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.