Large language models (LLMs) are increasingly deployed on edge nodes to support edge intelligence applications. To overcome limited GPU memory, offloading-based methods partition model parameters between the GPU and host memory, enabling inference on commodity hardware. However, deploying a single model instance using...
Zhen-Zheng Li, Zhi-Qing Tang, Jian-Xiong Guo et al.· IEEE Internet of Things Jour...· 0 citations
This work introduces a diffusion model as a generative prior to produce high-quality global deployment plans, effectively avoiding the local-optima problem common in conventional reinforcement learning.
Jie Gao, Xing-Dan Wang, Zhi-Qing Tang et al.· Tsinghua Science and Technol...· 0 citations
Advancements in edge computing and container technology have made it increasingly popular and convenient to deploy Large Language Models (LLMs) through containers at the edge. However, the limited GPU resources of edge servers make it impractical to retain the model in GPU memory for long periods due to the high memory...
Zhenzheng Li, Zhiqing Tang, Jianxiong Guo et al.· IEEE Transactions on Mobile...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.