Skip to content
Preprint

Deep reinforcement learning for separation control in turbulent wind-tunnel flow

Aug 2026 · 0 citations · 47 references
Physics

Abstract

This work investigates Deep Reinforcement Learning (DRL) as a tool for model-free closed-loop active separation control in a fully turbulent wind tunnel flow over a one-sided diffuser. The agent controls an array of magnetic valves (on/off) that eject compressed air into the boundary layer, while the environmental state is reduced to the signal from a single wall-shear-stress sensor placed near the natural transitory detachment point. The control law is learned in real time using Proximal Policy Optimization. Compared to the standard learning design based on the weighted sum of all rewards following an action, we demonstrate that a horizon aligned with the convective time of the flow leads to faster convergence and a more robust control strategy. The resulting control law corresponds to a low-duty-cycle actuation pattern that yields a forward-flow fraction of approximately $53\%$. This compares favorably with conventional and optimized periodic open-loop control ($\sim 40\%$ and $\sim 51\%$, respectively). The findings of this article indicate that, when embedded into an online experiment, DRL represents an efficient tool to identify robust and interpretable active separation control strategies.

View source

Similar papers

Aug 2026

Synthetic jet control of the airfoil based on deep reinforcement learning

This article proposes a closed-loop active flow control framework based on proximal policy optimization (PPO) algorithm and a synthetic jet to suppress severe flow separation of EH1590 airfoil at high angles of attack. Aerodynamic characteristics of airfoils under different jet parameters are obtained through computational fluid dynamics (CFD) simulation, and a deep neural network surrogate model is trained to achieve rapid prediction. Based on this alternative model, the PPO algorithm is used to train the agent, taking the pressure distribution of the flow field around the airfoil as the state, to optimize and improve the weighted objectives of lift-to-drag ratio and jet energy consumption. Finally, the agent is tested in both fixed and variable angle of attack tasks, and its adaptability and control effectiveness in a higher-resolution CFD environment are verified by coupling with CFD. The results show that at a fixed angle of attack of 15°, the lift-to-drag ratio increased from 4.56 to 7.34. Under variable angles of attack, the agent can automatically adapt and maintain a high lift-to-drag ratio. The error between the lift-to-drag ratio calculated directly by coupling with CFD and the test results of the surrogate model is within 5%, which verifies the effectiveness of the training strategy based on the surrogate model in transferring to the real flow field. This study provides an efficient and feasible technical path for the reinforcement learning application of airfoil flow control under high Reynolds number and high angle of attack conditions.

Yuhang Su, Yanping Song, Fu Chen et al. · 0 citations
Open access Jul 2026

Flow separation control of an infinite wing section via multi-agent reinforcement learning

This is the first study in which a fully three-dimensional computational fluid dynamics simulation has been employed to train a DRL model for active flow control in wings, suggesting promising directions for extending DRL-based strategies to higher Reynolds numbers and more complex wing configurations where prior physical knowledge may be limited.

R. Montalà, B. Font, Pol Suárez et al. · 0 citations
Preprint Jul 2026

Gradient-free learning of a closed-loop wall controller for turbulent drag reduction

Closed-loop wall controllers learnt by multi-agent reinforcement learning are usually trained on periodic boxes far smaller than the flows they are meant to drive, and a large part of their drag reduction is lost when they are carried across. Retraining on the target domain is not an affordable remedy: the centralised critic that assigns credit to each wall patch degrades as patches are added, the zero-net-mass constraint couples the patches it is asked to separate, and the episodes must be collected in sequence at a cost that grows with the domain. We propose instead a short gradient-free refinement stage, applied to the transferred policy on the domain it will drive. An evolution strategy scores whole flow episodes against a regularised objective, so no credit has to be assigned to individual patches, and the candidates in a generation are independent and run in parallel. Applied to a recurrent multi-agent policy trained on a minimal flow unit at $\Retau\simeq180$ and evaluated on a channel sixteen times larger in wall-parallel area, a few generations raise the drag reduction from $19.0\%$ to $25.7\%$, above the $22.5\%$ of opposition control. Because the policies before and after refinement share an architecture and a training history, the difference between the controlled flows follows from the refinement alone. It shows in the friction decomposition, in the Reynolds stresses and in the near-wall spectra, and the actuation moves from a weak coupling to the wall-normal velocity towards a strong coupling to the streamwise fluctuation.

Giorgio Maria Cavallazzi, Miguel Pérez Cuadrado, Alfredo Pinelli · 0 citations
#protein folding Open access Aug 2026

The HydroGym reinforcement learning platform for fluid dynamics.

HydroGym is introduced, a solver-independent reinforcement learning platform providing more than 60 validated, openly available flow control environments spanning from canonical laminar flows to complex turbulent flows, with systematic progression in the Reynolds number up to Re = 4 × 105, and Mach number variations in two and three dimensions.

Christian Lagemann, Sajeda Mokbel, Miro Gondrum et al. · 1 citation
Open access Aug 2026

Neural network-based prediction of actively controlled sweep events in a fully turbulent channel flow

This study investigates the use of a feed-forward neural network (FNN) and a hybrid approach combining an FNN with proper orthogonal decomposition (POD) to estimate the effect of control on sweep events within an opposition control framework. Conditionally averaged sweep events (CASEs) were experimentally sampled in a fully turbulent channel flow at a friction Reynolds number of 350. The two architectures were then trained to reconstruct the controlled CASEs using the control parameters (frequency, voltage amplitude, and delay time) as input. The results indicate that both approaches are able to reproduce the control signature across different actuation parameters with comparable accuracy. At the same time, the hybrid POD–FNN architecture requires over two orders of magnitude fewer trainable parameters. Finally, both local and global sensitivity analyses were performed, leveraging the differentiability of the trained models and their rapid inference times. The Jacobian matrix and Sobol indices were computed to conduct a local and a global sensitivity analysis, respectively. The findings revealed a high sensitivity to the actuation frequency and delay time parameters, whereas a lower sensitivity was observed with respect to the voltage amplitude.

E. Saccaggi, A. Nuvoloni, G. D. Di Cicca · 0 citations