Skip to content
Review Open access

TD-Learning and Q-Learning: A Survey of Theory, Analysis, and Trends

Aug 2026 · International Journal of Control, Automation and Systems · 0 citations · 23 references

Abstract

This paper provides a comprehensive survey of the convergence properties of temporal-difference (TD) learning and Q-learning, which are two fundamental algorithms in reinforcement learning (RL). We systematically categorize the existing literature into the tabular setting and two function approximation regimes: linear and nonlinear (deep neural networks). In the tabular setting, we review the foundational stochastic-approximation theory that ensures asymptotic convergence and discuss recent non-asymptotic results that provide explicit sample-complexity bounds under various coverage and sampling conditions. For linear function approximation, we address the stability challenges inherent in off-policy learning and examine stabilization mechanisms, such as gradient-based methods, regularization, and target networks. Furthermore, we explore the recent theoretical advancements in deep RL, focusing on finite-time analysis within the overparameterized regime. By synthesizing these diverse perspectives, this survey highlights the theoretical evolution from asymptotic stability to non-asymptotic efficiency and identifies the remaining gaps and provides a coherent roadmap for future research toward a unified theoretical understanding of RL dynamics.

Read PDF