While radiance field representations have achieved remarkable success in photo-realistic novel view synthesis with densely captured images, their performance sharply degrades under sparse input conditions, resulting in floater artifacts and missing regions caused by its inherent shape-radiance ambiguities and lack of i...
Hao-Yu Zhang, Shuai-Feng Zhi, Zhen-Hua Du et al.· IEEE Transactions on Visuali...· 0 citations
Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and how the human body will move to realize them, providing an important capability for assistive robotics and human-computer interaction. Existing methods struggle to translate semantic understanding into precise c...
Qiao-Hui Chu, Haoyu Zhang, Meng Liu et al.· 0 citations
Although text-to-image diffusion models exhibit remarkable generative power, concept erasure techniques are essential for preventing harmful content. Existing adversarial probes evaluate these methods by testing whether erased concepts can still be recovered. However, existing erasure and probe methods remain largely t...
Qian-Long Xiang, Miao Zhang, Kun Wang et al.· 0 citations
UAV-MAS is proposed, a training-free multi-agent system for MLLM-based UAV aerial image understanding and reasoning, comprising a Domain-Specific Perception Engine that routes queries to task-appropriate visual tools, a Context-Aware Iterative Refinement module (CAIR) that validates intermediate reasoning to curb error...
Hao-Yu Zhang, Shuoxun Zhang, Peng Ye et al.· 0 citations
By systematically revealing the strengths and limitations of existing models in aerial-ground collaborative reasoning, AeroGround provides a foundation for developing more capable aerial-ground collaborative embodied intelligence systems.
Shenghong Yi, Lin Zhang, Muzian Li et al.· 0 citations
This work proposes a novel information overloading method that is equipped with both extensive text and multi-dimensional image attacks, underscoring the need for stronger defenses against complex multimodal jailbreak inputs.
Haoyu Zhang, Yangyang Guo, Mohan S. Kankanhalli· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.