Skip to content

A Holistic Remote Fork Strategy Toward Fast Function Scaling in Serverless Edge Computing

Oct 2026 · IEEE Transactions on Mobile Computing · Vol 25, pp. 17732-17743 · 0 citations · 37 references

Abstract

Serverless edge computing, despite its flexibility and efficiency, is hindered by high startup latency during peak load. Remote fork, employing either Checkpoint/Restore (C/R) or Remote Direct Memory Access (RDMA), offers a potential solution for function scaling acceleration. Although RDMA fork is faster, the opportunities are limited, whereas C/R fork is more common but slower. Moreover, the regeneration capability that a forked function can further fork new instances complicates the remote fork decisions for fast scaling. Therefore, in this paper, we are motivated to address the problem on how to holistically exploit C/R fork and RDMA fork with the consideration of the underlying infrastructure features (e.g., topology, resources, etc.) to realize fast function scaling. We first formulate it to an Integer Linear Programming (ILP) problem. We further introduce a Heat metric to assess the potential of an edge server as a fork destination according to the topology and resource availability, and propose a Heat-based fork strategy (HEAT) for both the fork destination site and the corresponding fork mode decisions. Experiment results demonstrate that HEAT improves the function scaling speed by 46% and RDMA resource utilization by 51%, compared to state-of-the-art scaling solutions.

View source

Similar papers

Book Open access Sep 2026

RSC: Enabling High-Performance RDMA for Serverless Services with Adaptive Network Multiplexing

Serverless functions are short-lived, yet RDMA Reliable Connections (RCs) are expensive to create and assume long-lived endpoints. Even when compute sandboxes are warm, an invocation can still suffer a network-cold start if it establishes a new RC or accesses a remote endpoint for the first time. We present RSC (Reliab...

Wen-Di Song, Guang-Ping Xu, Yang-Yang Fan et al. · 0 citations
Preprint Aug 2026

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack,...

Shuo Yang, Xiao-yun Fan, Melissa Z. Pan et al. · 2 citations
Open access Aug 2026

CELLServe: An SLO-Aware and Cost Efficient LLMs Serving System for Serverless Computing Environments

CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.

Ze-Jian Wang, Nan Lin, Zi-Nuo Cai et al. · 0 citations
Preprint Sep 2026

Gutenberg: Taming Latency-Critical Cloud Services with Near-Data-Processing

Latency-critical cloud services place growing pressure on memory while requiring isolation, fairness, and predictable QoS. Near-data processing (NDP) reduces data movement by executing requests close to memory, and prior systems further improve locality through caching and replication. However, writes make replica main...

Qi Lin, Phillip B. Gibbons, Jovan Stojkovic et al. · 0 citations
Book Open access Sep 2026

QMScaler: A QoS-Constrained Resource-Efficient Microservice Autoscaling Framework for Edge Environments

Edge computing has become a key paradigm for supporting low-latency and high-concurrency services deployed with microservice architectures. However, limited device resources, highly dynamic workloads, and complex service dependencies make it difficult for existing autoscaling approaches—often relying on static threshol...

Tian-Yang Zheng, Peng-Fei Yang, Zhe Xu et al. · 0 citations
Preprint Sep 2026

Darpan: A Digital Twin Framework for the Next-Generation Computing Continuum

Computing-continuum applications distribute work across devices, edge systems, fog resources, and clouds. While a placement, scheduling, or recovery decision is being made, resource availability, network conditions, and application progress may change, so the decision can be invalid by the time it is executed. Existing...

Zhi-Yu Wang, Rajkummar Buyya · 0 citations

Related blog posts

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.