It is proved the matching lower bound $\Omega(n+\sqrt n\,\Delta L_{\max}/\varepsilon^2)$ for randomized IFO algorithms, including those that choose component indices and query points from the full preceding history, and PAGE and SPIDER are minimax optimal up to universal constants under individual and mean-squared smoo...
Large language models (LLMs) are increasingly deployed on edge nodes to support edge intelligence applications. To overcome limited GPU memory, offloading-based methods partition model parameters between the GPU and host memory, enabling inference on commodity hardware. However, deploying a single model instance using...
Zhen-Zheng Li, Zhi-Qing Tang, Jian-Xiong Guo et al.· IEEE Internet of Things Jour...· 0 citations
Under individual smoothness, the optimal incremental first-order oracle (IFO) complexity of nonconvex finite-sum optimization has remained open. Known algorithms use $O(n+\sqrt{n}\,\Delta L_{\max}/\varepsilon^2)$ calls, while prior lower bounds miss a factor of $\sqrt{n}$. We prove the matching lower bound for randomiz...
FeatFix is introduced, a local exact-feature correction method for cached diffusion inference that replaces the complete draft block output with the exact output computed from the same incoming state, avoiding token- or channel-level partial replacement and full-timestep recomputation.
Han-Shuai Cui, Zhiqing Tang, Z. Yao et al.· arXiv.org· 0 citations
MemTxn is a governance layer outside the answer model that verifies whether an update is supported by its source and restores the application-visible state after a fault, and achieves the highest average F1 across all twelve answer-model configurations.
Han-Shuai Cui, Zhiqing Tang, Z. Yao et al.· arXiv.org· 2 citations
This work introduces a diffusion model as a generative prior to produce high-quality global deployment plans, effectively avoiding the local-optima problem common in conventional reinforcement learning.
Jie Gao, Xing-Dan Wang, Zhi-Qing Tang et al.· Tsinghua Science and Technol...· 0 citations
Advancements in edge computing and container technology have made it increasingly popular and convenient to deploy Large Language Models (LLMs) through containers at the edge. However, the limited GPU resources of edge servers make it impractical to retain the model in GPU memory for long periods due to the high memory...
Zhenzheng Li, Zhiqing Tang, Jianxiong Guo et al.· IEEE Transactions on Mobile...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.