Diffusion Transformers (DiTs) have become a dominant architecture for video generation, but their efficiency is limited by the quadratic complexity of full attention. Sparse attention reduces this cost by retrieving important blocks and computing attention only within them, but inaccurate retrieval can either degrade g...
Yun-Wei Dai, Jia-Rui Wen, Hui-Ping Zhuang et al.· 0 citations
X-SG$^2$S is the first framework to unify 1D-to-3D watermarking and enable simultaneous multi-modal watermark embedding in 3DGS, achieving this with minimal rendering interference and zero modifications to parameters or pipelines.
EgoSafe-Bench is introduced, a benchmark specifically designed to probe forensic reasoning in egocentric safety scenarios, generated by pairing each of the 3,000 video clips with a QA chain governed by the proposed Hierarchical Reasoning Evaluation (HRE) protocol.