Sep 2026· Proceedings of the International Conference on Parallel Processing· 0 citations· 10 references
TL;DR
QMScaler is proposed, a QoS-driven microservice horizontal scaling framework based on Monte Carlo Tree Search (MCTS) that aims to minimize the number of container instances while meeting the QoS requirements of multiple application functions.
Abstract
Edge computing has become a key paradigm for supporting low-latency and high-concurrency services deployed with microservice architectures. However, limited device resources, highly dynamic workloads, and complex service dependencies make it difficult for existing autoscaling approaches—often relying on static thresholds or simple workload prediction—to ensure both QoS guarantees and efficient resource utilization. In particular, most prior studies overlook the QoS impact of instance migration paths during scaling and fail to account for the heterogeneous contributions of microservices to end-to-end latency, resulting in unbalanced resource allocation and degraded overall performance. To address these challenges, we propose QMScaler, a QoS-driven microservice horizontal scaling framework based on Monte Carlo Tree Search (MCTS). QMScaler aims to minimize the number of container instances while meeting the QoS requirements of multiple application functions. It incorporates a fine-grained performance model that captures the effects of scaling actions and migration paths, a microservice importance metric that prioritizes resources for critical services, and a domain-knowledge-guided heuristic MCTS algorithm that improves search efficiency and decision stability. Experiments on real edge clusters demonstrate that QMScaler achieves required QoS with fewer instances and exhibits superior adaptability and robustness under dynamic and bursty workloads compared to existing methods.
CARSO (Counterfactual Autoscaling and Resource-efficient Service Orchestration), a proactive and interpretable framework that integrates eXplainable Artificial Intelligence (XAI) into the autoscaling process that outperforms state-of-the-art proactive autoscaling frameworks in both QoS compliance and overall resource u...
Lazaros Liatsas, Godfrey M. Kibalya, Angelos Antonopoulos· IEEE Transactions on Network...· 0 citations
OpScale is presented, a practical operator-level orchestration framework of profiling, provisioning, placement, and runtime serving that attains SLOs with up to 36.3% fewer GPUs and 28% less power, or achieves 44% higher throughput under fixed cost budgets.
Xingqi Cui, Chieh-Jan Mike Liang, Ziang T. Tang et al.· 0 citations
The lack of determinism restricts the integration of safety-critical applications into Edge–Fog–Cloud (EFC) architectures. Existing EFC schedulers are typically designed for dynamic, best-effort operation based on unmanaged resource allocation and elastic virtualization. This paradigm introduces unbounded queueing, res...
Omar Hekal, Josepaul Paulachan, Daniel Onwuchekwa et al.· Future Internet· 0 citations
The results show that soft SLO limits reduce corrective rescheduling actions by 49% compared to hard-limit approaches while maintaining acceptable performance guarantees, and resource-aware scheduling decreases node-level congestion and further mitigates SLO violations, demonstrating the effectiveness of incorporating...
Oliver Larsson, Thijs Metsch, Cristian Klein et al.· 0 citations
Modern cloud platforms provide diverse compute services, including virtual machines, Function-as-a-Service, and Query-as-a-Service, each offering unique trade-offs in performance, elasticity, and cost. While these services collectively cover the diverse needs of OLAP workloads, existing systems typically rely on a sing...
Wen-Bo Li, Hao-Qiong Bian, Chao Zhang et al.· Proceedings of the ACM on Ma...· 0 citations
RDPart is proposed, an OS-level Reuse-Driven LLC Partitioning policy designed to improve fairness while preserving the QoS of cloud workloads, and adopts a black-box design, making it well-suited for public cloud environments where real-time QoS feedback from applications is unavailable.
Javier Aznal, J. C. Saez, Carlos Bilbao· Proceedings of the Internati...· 1 citation
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…