DRLM, a Deep Reinforcement Learning-based LLM query orchestration framework in edge environments shows robust and stable orchestration, and improves latency under increasing workloads up to 61.4%, demonstrating robust and stable orchestration.
Abstract
Large language model (LLM) services increasingly process heterogeneous queries with diverse latency, accuracy, and resource requirements. While edge deployment reduces response time, the heterogeneity of devices and the diversity of model families, parameter scales, and quantization levels make efficient LLM query orchestration challenging. This paper introduces DRLM, a Deep Reinforcement Learning-based LLM query orchestration framework in edge environments. DRLM integrates two lightweight predictors: (i) a class-conditioned quality estimator that maps queries to semantic categories and infers model performance, and (ii) a feature-driven latency predictor that estimates inference time across model-device configurations. These predictions, combined with system state, feed a factorized Proximal Policy Optimization (PPO) agent that performs state-aware orchestration decisions. To enable data-driven orchestration, we construct a large-scale benchmarking dataset with 223 835 measurements spanning 1258 queries, 6 query classes, 8 model families (32 deployed instances), 5 quantization levels, and heterogeneous edge devices. Evaluation on a 64-node edge cluster and comparison with three baselines and two state-of-the-art methods show that DRLM reduces inference latency by up to 51% and queuing delay by up to 67 %, while incurring at most 8% accuracy loss. It improves latency under increasing workloads up to 61.4%, demonstrating robust and stable orchestration.
Large language models (LLMs) and agentic AI systems are creating rapidly growing inference energy demands as model sizes grow and reasoning trajectories extend. While in practice, many queries do not require the capabilities of the largest available model, and routinely directing such queries to a high-capability model...
Muhammad Abdur Rab Siddiqui, Daniel Rojas, Chen Yang et al.· 0 citations
Large language models (LLMs) are increasingly deployed to solve complex scientific and practical problems via iterative optimization. However, dynamically coordinating diverse search mechanisms as candidate quality, failure modes, and resource budgets evolve remains a critical open challenge. Targeted empirical diagnos...
Chen-Xing Wei, Si-Chen Liu, Lizzie Liu et al.· 0 citations
The growth of cloud-based data infrastructure has skyrocketed and the complexities in managing and optimizing data pipelines at scale has never been higher. Conventional pipeline orchestration techniques are based on models that use rigid configurations and rule-based scheduling systems, which are inherently non-adapti...
Pranay Kumar Raini· International Conference Com...· 0 citations
Learning-based Structured Query Language (SQL) optimizers often face low sample efficiency and high training costs. To address these challenges, this study proposes a query optimization framework, GT- MuZero, which integrates a Graph Transformer (GT) with the model-based reinforcement learning (RL) algorithm MuZero. Th...
Fu-Ming Ye, Wen-Ting Li, Qiong Zhou et al.· Informatica· 0 citations
The growing deployment of applications in real-world edge environments is driving an increasing demand for large language model (LLM) training in distributed systems. This demand creates significant challenges for resource management, particularly in resource-constrained edge cloud environments. Federated learning (F...