Skip to content
Book Open access

ActiveRDMA: SmartNIC-offloaded, Multi-Step RDMA operations

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 53 references

Abstract

Modern distributed systems rely on one-sided RDMA for low-latency, high-throughput data movement without CPU intervention. However, one-sided RDMA provides only basic primitives (read, write, atomics) where the target side is passive. Applications with complex communication patterns, such as remote data structure traversal, require the host CPU to orchestrate multiple steps of communication, resulting in several network round-trips and thus diminishing the benefits of one-sided RDMA. We present ActiveRDMA, an active RDMA model that extends traditional one-sided semantics by leveraging on-NIC programmability. This enables complex, multi-step RDMA patterns to execute directly on the target side, eliminating host CPU involvement and reducing communication round-trips. We implement and evaluate ActiveRDMA on the NVIDIA DPA, an on-path SmartNIC architecture. Our evaluation demonstrates substantial improvements for fine-grained, latency-sensitive communication patterns: completion time reduces by up to 43% for messages up to 4 KiB that cannot benefit from batching. For operation rate, ActiveRDMA trades single-QP efficiency for scale-out throughput, surpassing conventional one-sided RDMA from 4+ connections onward, achieving up to 1.8× higher operation rate as the number of QPs scales. Bandwidth-dominated communications, however, see no advantage over conventional one-sided RDMA. In a distributed graph traversal application, we achieve up to 33% runtime reduction, with complete overlap of communication and computation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.