Preprint
Jul 2026
Voltron: Enabling Elastic Multi-Device Execution of LLM Inference for Empowered Edge Intelligence
Voltron is proposed, a novel on-device LLM inference framework that elastically utilizes multiple user-end devices for LLM inference execution while adapting to diverse real-world edge environments, satisfying user QoS requirements.
Chanwoo Cho, Wooseok Kim, Yonglak Son et al.
· 0 citations