Rapid advances in AI-generated video (AIGV) have increased the risks posed by deceptive video manipulation. Unlike fully synthetic videos, manipulated videos retain most source content and alter only localized regions, making forensic analysis particularly challenging. Existing video forgery research faces two limitati...
Heng-Rui Kang, Zhong-Hao Yan, Yuxuan Yang et al.· 0 citations
This work introduces Spatially-Grounded Thought Tokenization to explicitly align special tokens with spatial locations for clear spatial correspondence and visual interpretability, and proposes Mixture-of-Thought-Tokens, a new free-form multimodal grounding method that bridges the perception-reasoning gap.
Tianyi Gao, Han Fang, Tianyi Ding et al.· arXiv.org· 0 citations
MMBench-Live is presented, a continuously evolving multimodal benchmark built by a multi-agent-driven automated pipeline that preserves stable model rankings, maintains semantic alignment with the original benchmark, and exhibits weaker contamination-related memorization signals, suggesting a practical and scalable par...
Yuanzhi Liu, Shousheng Zhao, Bo Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.