Skip to content

Similar papers

Review Open access Jul 2026

Enhancing the Kubernetes Scheduler: A State-of-the-Art Review from Cloud to Edge

A comprehensive review of Kubernetes scheduling strategies published between January 2023 and January 2026 is presented and a multi-dimensional taxonomy is established that categorizes scheduling approaches based on common objectives, modification methods, optimization methodologies, targeted workloads, evaluation methods, scheduling scopes, and performance metrics.

Mohammed Alhakimi, R. Latip · 0 citations
Open access Aug 2026

CELLServe: An SLO-Aware and Cost Efficient LLMs Serving System for Serverless Computing Environments

CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.

Zejian Wang, Nan Lin, Zinuo Cai et al. · 0 citations
Preprint Aug 2026

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving

OpScale is presented, a practical operator-level orchestration framework of profiling, provisioning, placement, and runtime serving that attains SLOs with up to 36.3% fewer GPUs and 28% less power, or achieves 44% higher throughput under fixed cost budgets.

Xingqi Cui, Chieh-Jan Mike Liang, Ziang Tang et al. · 0 citations
Conference Jul 2026

LLM Inference Performance Optimization in Limited-Resource Environments

With the widespread adoption of LLM-based chatbots, cloud-dependent solutions have come to dominate the market. However, open-source pre-trained LLMs are enabling implementation of local solutions. Achieving competitive performance locally requires the ability to run high-parameter models. Here, the primary bottleneck is GPU VRAM capacity, which limits model parameter size. Furthermore, the efficiency of inference optimizations such as kv caching depends directly on the amount of available VRAM remaining after the model is loaded. In such a scenario with hardware constraints, we conduct user tests by applying kv cache quantization. As a result, we identify distinct performance trends in critical metrics such as Time-to-First-Token (TTFT) and Total Generation Time. Additionally, we evaluate model accuracy results using the LLM-as-a-judge paradigm.

E. Yilmaz, Muhammet Furkan Coşkun, M. Aydoğdu et al. · 0 citations
Preprint Aug 2026

LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations

This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers, using an LLM to predict key metrics such as execution time and energy consumption from source code.

Hanzhao Wang, Jingxuan Wu, Yumeng Li et al. · 0 citations