Skip to content
#edge computing Book Open access

QMScaler: A QoS-Constrained Resource-Efficient Microservice Autoscaling Framework for Edge Environments

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 10 references

TL;DR

QMScaler is proposed, a QoS-driven microservice horizontal scaling framework based on Monte Carlo Tree Search (MCTS) that aims to minimize the number of container instances while meeting the QoS requirements of multiple application functions.

Abstract

Edge computing has become a key paradigm for supporting low-latency and high-concurrency services deployed with microservice architectures. However, limited device resources, highly dynamic workloads, and complex service dependencies make it difficult for existing autoscaling approaches—often relying on static thresholds or simple workload prediction—to ensure both QoS guarantees and efficient resource utilization. In particular, most prior studies overlook the QoS impact of instance migration paths during scaling and fail to account for the heterogeneous contributions of microservices to end-to-end latency, resulting in unbalanced resource allocation and degraded overall performance. To address these challenges, we propose QMScaler, a QoS-driven microservice horizontal scaling framework based on Monte Carlo Tree Search (MCTS). QMScaler aims to minimize the number of container instances while meeting the QoS requirements of multiple application functions. It incorporates a fine-grained performance model that captures the effects of scaling actions and migration paths, a microservice importance metric that prioritizes resources for critical services, and a domain-knowledge-guided heuristic MCTS algorithm that improves search efficiency and decision stability. Experiments on real edge clusters demonstrate that QMScaler achieves required QoS with fewer instances and exhibits superior adaptability and robustness under dynamic and bursty workloads compared to existing methods.

Read PDF

Similar papers

Open access 2026

Counterfactual Autoscaling for Resource-Efficient Service Orchestration in the Cloud–Edge Continuum

CARSO (Counterfactual Autoscaling and Resource-efficient Service Orchestration), a proactive and interpretable framework that integrates eXplainable Artificial Intelligence (XAI) into the autoscaling process that outperforms state-of-the-art proactive autoscaling frameworks in both QoS compliance and overall resource u...

Lazaros Liatsas, Godfrey M. Kibalya, Angelos Antonopoulos · 0 citations
Preprint Aug 2026

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving

OpScale is presented, a practical operator-level orchestration framework of profiling, provisioning, placement, and runtime serving that attains SLOs with up to 36.3% fewer GPUs and 28% less power, or achieves 44% higher throughput under fixed cost budgets.

Xingqi Cui, Chieh-Jan Mike Liang, Ziang T. Tang et al. · 0 citations
Open access Aug 2026

Hierarchical Scheduler with Adaptive Time-Budget Reallocation for Time-Triggered Edge-Fog-Cloud Architectures

The lack of determinism restricts the integration of safety-critical applications into Edge–Fog–Cloud (EFC) architectures. Existing EFC schedulers are typically designed for dynamic, best-effort operation based on unmanaged resource allocation and elastic virtualization. This paradigm introduces unbounded queueing, res...

Omar Hekal, Josepaul Paulachan, Daniel Onwuchekwa et al. · 0 citations
Preprint Aug 2026

Beyond the Limits: Flexible and Congestion-Aware Cluster Scheduling for the Cloud

The results show that soft SLO limits reduce corrective rescheduling actions by 49% compared to hard-limit approaches while maintaining acceptable performance guarantees, and resource-aware scheduling decreases node-level congestion and further mitigates SLO violations, demonstrating the effectiveness of incorporating...

Oliver Larsson, Thijs Metsch, Cristian Klein et al. · 0 citations
Open access Sep 2026

EASE: Resource-aware Query Scheduling across Heterogeneous Cloud Compute Services

Modern cloud platforms provide diverse compute services, including virtual machines, Function-as-a-Service, and Query-as-a-Service, each offering unique trade-offs in performance, elasticity, and cost. While these services collectively cover the diverse needs of OLAP workloads, existing systems typically rely on a sing...

Wen-Bo Li, Hao-Qiong Bian, Chao Zhang et al. · 0 citations
Book Open access Sep 2026

RDPart: A reuse-based OS-level cache-partitioning policy for fairness optimization in cloud data centers

RDPart is proposed, an OS-level Reuse-Driven LLC Partitioning policy designed to improve fairness while preserving the QoS of cloud workloads, and adopts a black-box design, making it well-suited for public cloud environments where real-time QoS feedback from applications is unavailable.

Javier Aznal, J. C. Saez, Carlos Bilbao · 1 citation

Related blog posts

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.