Benchmarking Parallel Programming Models for High-Throughput Security Telemetry Analytics: OpenMP, MPI, and CUDA
Abstract
This paper compares three parallel programming paradigms—shared memory (OpenMP), distributed memory (MPI), and heterogeneous GPU computing (CUDA)—for a telemetry-inspired Map-Filter-Reduce-Sort analytics pipeline. The study evaluates datasets ranging from 100 million to 1 billion elements and focuses on end-to-end behavior, memory pressure, communication overhead, and out-of-core execution. The pipeline is an intentionally simplified proxy for common stages in security telemetry analytics: Map represents feature normalization or enrichment; Filter represents threshold- or rule-based selection; Reduce represents global aggregation; and Sort represents event ranking or prioritization. To strengthen workload coverage and provide a limited external validity check, the study includes a CICIDS2017-derived replicated flow workload, compute-intensity variants of the Map stage, detailed overhead decomposition for OpenMP, MPI, and CUDA, and a model-based sensitivity projection of MPI communication cost. Experimental results on a virtualized 32-core CPU platform with one legacy dual-GPU NVIDIA Tesla K10 accelerator board show a platform-specific crossover: CUDA is fastest at 100M–250M elements, whereas MPI is fastest at 500M–1B after the CUDA implementation enters the out-of-core regime. These quantitative thresholds are specific to the evaluated hardware and implementations and should not be generalized directly to modern accelerators. All MPI measurements were obtained on a single host using intra-node shared-memory transport; the multi-node analysis is a model-based sensitivity projection rather than a measured cluster experiment. Beyond execution time, we discuss deployment-relevant risks for cybersecurity analytics: host-device transfers, collective data redistribution, temporary buffers, and out-of-core spill mechanisms as potential sources of availability degradation under high-retention workloads and other adverse input conditions.