Skip to content
Book Open access

MEDO: Adaptive Multi-Stream Data Offloading for Efficient Management of Disaggregated Memory Pool

Sep 2026 · Proceedings of the International Conference on Parallel Processing · pp. 67-77 · 0 citations · 21 references

Abstract

The scale of data-intensive workloads in intelligent datacenters has grown rapidly in recent years, intensifying the need for efficient data movement within disaggregated memory pools across memory and storage subsystems. However, most existing solutions focus primarily on optimizing communication between compute nodes and data nodes, overlooking fine-grained management of in-pool data flows. Furthermore, prevailing data swap strategies often suffer from high overhead, limited parallelism, and a lack of adaptivity to dynamic workloads, especially for bursty data access patterns in heterogeneous memory pools. To address these issues, this paper introduces MEDO, a high-parallelism data offloading system for disaggregated memory pools. MEDO leverages a novel multi-stream data offloading architecture, featuring parallel data streams and approximate LRU queues, to maximize throughput and efficiently handle diverse workloads. Additionally, MEDO incorporates a lightweight, adaptive offloading agent that dynamically optimizes data placement decisions and fine-grained system configurations. Our prototype achieves up to 3.6 × latency reduction on real-world data services compared with baselines and can reduce in-pool memory usage by up to 50% on state-of-the-art disaggregated memory systems.

Read PDF

Similar papers

Book Open access Sep 2026

Spatiotemporal Load Balancing for Near-Memory Accelerated Databases by Partial Resharding

Near-memory acceleration, where a large number of compute nodes with limited memory process data in parallel, is a promising approach for in-memory databases. Therein, partitioning is required to balance data and queries for skewed workloads. However, existing load balancing methods lack efficient support for dynamic w...

Takato Hideshima, Shigeyuki Sato, Tomoharu Ugawa · 1 citation
Open access Aug 2026

Transfer-Efficient Data Processing in Disaggregated Systems

Disaggregated analytics systems separate compute, memory, and storage to improve elasticity and resource utilisation, making network transfer a central bottleneck. Existing systems typically access remote data at fixed coarse granularities, transferring entire files, columns, or column-chunks even when queries proces...

David Loughlin, Holger Pirk · 0 citations
Jul 2026

CrocSort: Resource-Efficient, Skew-Resilient Parallel External Merge Sort

CrocSort is presented, a byte-balanced parallel external merge sort with configurable memory and per-phase thread settings with practical resource-configuration rules for selecting these settings from input size, memory budget, and thread cap.

Riki Otaki, Charles Benello, Fuheng Zhao et al. · 0 citations
Open access Aug 2026

HARMONI: Heterogeneity-Aware I/O Scheduling for Mixed Workloads in SSD-Based HPC Systems

Experimental results across diverse HPC and AI workloads demonstrate that HARMONI significantly reduces average makespan by up to 90% compared to state-of-the-art schedulers, effectively bridging the gap between evolving storage hardware and the dynamic I/O demands of modern HPC systems.

Ze-Xi Cai, Tong Zhao, Shadi Ibrahim et al. · 0 citations
Book Open access Aug 2026

DynamoServe: A Distributed Tiered Memory System for Multi-tenant LLM Serving

DynamoServe is presented, a multi-tenant LLM serving framework that addresses challenges through three key innovations: leveraging stranded GPU memory to offload model weights and KV caches, mitigating resource fragmentation in multi-workload environments, and improving memory locality through coordinated data placemen...

Diman Zad Tootaghaj, Khaled Diab, Bob Lantz et al. · 0 citations
Book Open access Aug 2026

STORM: Enabling Traffic Scheduling for RDMA

STORM is presented, a NIC-level scheduler for all types of RDMA workloads using NIC-only information: the known RDMA request size, and per-queue-pair backlog, and converts these signals into a small number of extra priority levels on the wire and prioritizes requests that are either near completion or blocking queued d...

Jichun Wu, Ran Shu, Gianni Antichi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.