While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, they frequently suffer from structural hallucinations in human-centric scenarios. Accurately modeling human subjects is foundational for critical downstream applications, such as accurate avatar/video/image genera...
Bozhou Li, Jia-Hang Zhang, Yue Ding et al.· 0 citations
Flux-OPD is proposed, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains and outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propos...
Tengfei Liu, Yang Shi, Yuran Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.