Skip to content
Open access

Autonomous Data Pipeline Optimization Using Reinforcement Learning Techniques

2025 · International Journal of Applied Data Science & Modern Computing · 0 citations

Abstract

Modern large-scale data pipelines support analytics, AI, ML, and real-time applications but face challenges related to scalability, resource utilization, reliability, and changing workloads. This paper proposes a reinforcement learning (RL)-based autonomous optimization framework that integrates RL agents with data orchestration platforms to continuously monitor pipeline states and optimize operations. The framework uses system metrics such as workload patterns, queue lengths, execution delays, resource consumption, and failure rates to make intelligent decisions on task scheduling, resource allocation, workload balancing, and fault recovery. Three RL algorithms—Q-learning, Deep Q-Networks (DQN), and Proximal Policy Optimization (PPO)—are evaluated. Experimental results demonstrate improved throughput, reduced latency, enhanced fault tolerance, and better resource efficiency compared to traditional rule-based approaches. The proposed framework enables adaptive, self-managing data pipelines that improve scalability, resilience, and operational efficiency across enterprise, cloud, and edge environments.

Read PDF