Experience reuse can reduce repeated exploration in embodied agents, but a trajectory that succeeded previously may be unsuitable for the current execution context. Existing memory systems pri marily optimize construction and retrieval; semantic relevance and historical success therefore remain insufficient when retrie...
Quan-Quan Li, Hong-Bo Zhang, Yi-He Chi et al.· 0 citations
This work proposes NOSA, a trainable sparse attention mechanism natively designed for KV cache offloading that explicitly constrains the volume of CPU-GPU KV transfers, thereby achieving low communication overhead and high decoding throughput.
Yu-Xiang Huang, Peng-Jie Wang, Ji-Cheng Han et al.· arXiv.org· 4 citations· ⚡1
FutureBridge is presented, which ranks joint LLM-SLM token candidates according to how well they support the SLM's subsequent reasoning, and indicates that token selection benefits from modeling whether the receiving SLM can use each candidate to continue reasoning, rather than relying on the LLM's local preference alo...
Quanquan Li, Hongbo Zhang, Yihe Chi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.