Skip to content
Book Open access

Anchor: Mitigating GPU Shallow Disruptions with Decoupled Memory

Sep 2026 · Proceedings of the ACM SIGOPS 32nd Symposium on Operating Systems Principles · 0 citations · 38 references

Abstract

Job failures are frequent in large-scale GPU clusters for LLM workloads, leading to significant resource wastage. The vast majority of these are shallow disruptions (e.g., software errors or updates), where only the worker process crashes while the underlying GPU and OS kernel remain intact. Existing recovery systems, however, are failure-agnostic, relying on heavyweight checkpointing that forces a full, slow state reload even when the data could have survived on the GPU. In this paper, we propose Anchor, a system that avoids this overhead by decoupling memory ownership from the failure-prone worker process. Anchor uses a daemon process to manage GPU memory. Leveraging the GPU's native IPC memory mechanism, Anchor enables a restarted worker to remap and reuse the daemon-managed memory with no overhead on the normal execution path. To make this idea practical for real-world ML applications, Anchor introduces an access-path-based identity resolution mechanism to remap logical tensors to their physical memory objects after restart. Anchor also provides lightweight, application-aware consistency protocols to ensure data integrity with little overhead. We integrate Anchor into DeepSpeed and vLLM. Our evaluation demonstrates recovery acceleration, with recovery time reduced by 60.7% for inference and 26.6% for training on average. For Qwen3-235B inference, Anchor reduces the recovery time from 143.78 s to 27.85 s.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.