We propose ECHO for dyadic 3D facial motion generation under a strict dual-stream audio-only setting, formulating the problem as an asymmetric task involving speech-constrained articulation and one-to-many listener reactions. To address this asymmetry, ECHO decomposes motion into a deterministic anchor that captures st...
Zhuo-Qiang Cai, Yu-Jie Sun, Chao-Yue Niu et al.· 0 citations
On-device large language model (LLM) serving is a cornerstone of local-first personal intelligence, offering users data sovereignty, strong privacy guarantees, and freedom from cloud API latency and cost. Although KV caching is widely used to reduce latency in long-context inference, existing designs were primarily opt...
Zheng-Xiang Huang, Sheng-Heng Chen, Chao-Yue Niu et al.· 0 citations
Agents tend to optimize, select, or constrain execution structures before decisive runtime outcomes are observed. However, such pre-execution commitment creates an orchestration bottleneck: when intermediate evidence invalidates the pending continuation, agents must either execute stale steps or replan broadly, compoun...
Tian-Xing Wang, Ming-Ming Zhao, Shuai Huang et al.· 1 citation
C2KV is proposed, a unified framework for non-prefix KV reuse that jointly optimizes KV cache compression and concatenation that significantly reduces KV cache storage and transfer costs.
Chuheng Du, Jun-Yi Chen, Hanlin Tang et al.· Proceedings of the 32nd ACM...· 2 citations
Evidence-First Reflection (EFR), a two-stage reflector that explicitly decouples action-induced visual differences extraction from outcome verification, makes reflection better grounded in screen transitions, while reducing both visual search complexity and reasoning burden.
Yijie Ma, Chao-Yue Niu, Fan Wu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.