Agentic recognition requires visual perception to move beyond static scene understanding and produce structured scene representations that support the perception--reasoning--action loop. Existing single-image 3D generation methods, however, mainly produce visually plausible object assets rather than simulation-ready sc...
Lin-Tao Wang, Ming-Yang Sun, Yang Liu et al.· 0 citations
Generative models have advanced image-conditioned 3D content creation, yet generating controllable and executable 3D scenes from a single image remains challenging. Existing 3D generative approaches can synthesize visually plausible objects and scenes, but their spatial layout estimation is coupled with specific asset...
Ding-Kang Yang, Yi-Zhou Liu, Wen-Dong Cheng et al.· 0 citations
Omni-modal models have expanded multimodal interaction across vision, audio, speech, and language. However, their training is predominantly organized around semantic descriptions and general-purpose objectives, leaving physical attributes, interaction states, and causal mechanisms only partially specified. This gap is...
Yi-Zhou Liu, Jing-Hang Han, Kai Qiu et al.· 0 citations
PFAdapter is proposed, a communication-efficient framework introducing hierarchical LoRA decomposition to explicitly separate adapter parameters into global-shared and local-private components, establishing an efficient solution for agentic AI deployment in resource-constrained communication networks.
Jing Liu, Kun Yang, Yan Wang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.