Skip to content

Reinforcement Learning-Guided Transmitter Optimization for Short-Reach IM/DD FSO Systems

2026 · IEEE Photonics Technology Letters · Vol 38, pp. 1733-1736 · 0 citations · 19 references

Abstract

We propose a reinforcement-learning (RL) guided transmitter optimization framework for short-reach IM/DD free-space optical (FSO) links that jointly tunes probabilistic shaping (PS), geometric shaping (GS), and pre-equalization (Pre-EQ) filter coefficients using the soft actor-critic (SAC) algorithm. The agent adapts transmitter parameters in a closed-loop manner from measured link-quality feedback, removing reliance on explicit channel models. In a 30-m PS-PAM4 FSO testbed with a 7-tap FIR Pre-EQ at the transmitter and standard offline DSP at the receiver, the agent uses generalized mutual information (GMI) as the reward to coordinate PS/GS with Pre-EQ. Experiments show a peak achievable information rate (AIR) of 194.3 Gb/s at 110 Gbaud and receiver-sensitivity gains of ~1.5 dB over traditional Pre-EQ and ~3 dB over no Pre-EQ, demonstrating the efficacy of closed-loop, model-free transmitter learning for next-generation optical wireless access.

View source

Similar papers

2026

Advancing End-to-End Communication Systems With Physics-Aware Learning

The advent of neural network (NN)-based autoencoders has enabled joint transceiver optimization via end-to-end (E2E) learning. However, most existing approaches rely on differentiable surrogate channel models, leading to performance mismatch when deployed over real-world channels. Emerging model-free variants alleviate this issue but still lack principled means to obtain reliable and informative gradients from physical channels. In this work, we propose a novel E2E optimization framework, termed physics-aware learning (PAL), that eliminates the need for channel modeling by executing forward propagation directly over real channels. To support gradient-based training under such model-free settings, a fused gradient proxy is designed that combines gradient jumps and local Jacobian estimations to derive transmitter-side gradients. We instantiate this framework through the development of a physics-aware learning autoencoder (PALAE), which jointly integrates probabilistic shaping, geometric shaping, and neural equalization within a carrier-less amplitude and phase (CAP) modulation system. While retaining compatibility with conventional DSP-based transceivers, PALAE enables online, interpretable, and modular E2E optimization across heterogeneous components. Simulation and experimental evaluations demonstrate the effectiveness of PALAE, achieving a net bit rate exceeding 400 Gbps in a practical 0.5-km intensity modulation-direct detection (IMDD) fiber system, representing the state-of-the-art rate achieved through CAP modulation with single wavelength.

Yuan Wei, Chaoxu Chen, Li Yao et al. · 0 citations
Preprint Jul 2026

System-Aware Adaptive CSI Feedback via RL-Guided Autoencoder Switching in Multi-User MIMO System

This paper proposes a system-aware adaptive channel state information (CSI) feedback framework for massive multiple-input multiple-output (mMIMO) systems, aiming to dynamically optimize the trade-off between reconstruction fidelity and signaling overhead. While deep learning-based autoencoders (AEs) have enabled significant CSI compression, conventional fixed-ratio schemes fail to adapt effectively to non-stationary channel conditions. To address this limitation, we develop a reinforcement learning (RL)-driven control framework that operates over a bank of pretrained multi-rate AEs, each corresponding to a distinct compression ratio (CR). At each time step, a centralized RL agent selects the most suitable CR for each user based on observed channel conditions and system performance indicators. Distinct from conventional mean squared error (MSE)-centric designs, we introduce a system-aware reward formulation that jointly accounts for spectral efficiency via signal-to-interference-plus-noise ratio (SINR), feedback overhead constraints, and the computational cost of model adaptation. Simulation results on high-dimensional delay-domain CSI datasets demonstrate that the proposed RL-guided framework effectively balances the overhead-accuracy tradeoff and adapts to dynamic channel environments. The proposed method improves spectral efficiency and feedback efficiency compared with fixed compression schemes and adaptive baselines, while maintaining a modest computational and memory footprint. Averaged over different numbers of users and across all considered baselines, the proposed RL framework reduces the CSI feedback cost by more than 53.4%, improves the average downlink sum rate by 53.64%, and reduces the NMSE by 22.38%. These results demonstrate its ability to achieve a more efficient rate-accuracy-feedback tradeoff under dynamic wireless conditions.

Maryam Ansarifard, M. Sharma, George Exarchakos et al. · 0 citations
2026

End-to-End Learning With EM-Aware Differentiable-Ready Discrete-Phase RIS for MIMO–OFDM via Ray Tracing

Most learning-based reconfigurable intelligent surface (RIS) designs assume continuous control or closed-box channel models, limiting physically grounded end-to-end (E2E) training and neglecting practical 1-bit hardware constraints. We propose a physics-informed framework for RIS-assisted MIMO–OFDM that embeds an optics-consistent differentiable ray tracer (RT) in the training loop while enforcing strictly discrete, frequency-flat 1-bit RIS control. Per-element phases are injected into the RT transition matrices and optimized via quantization-aware training (QAT) with hard binary forward passes and surrogate gradients. The same pipeline also acts as a digital twin, enabling controlled sweeps over geometry, materials, and LOS/NLOS conditions to generate labeled CIRs and OFDM channels. We study both staged optimization, where the RIS is trained using a pilot-energy proxy before neural receiver (NRX) training, and full E2E co-optimization by backpropagating NRX loss through the RT–RIS block. In the E2E setting, the RIS is optimized as part of the electromagnetic propagation environment using receiver-side bitwise loss, rather than through an intermediate channel-quality proxy alone. We compare QAT with straight-through estimator (STE), straight-through Gumbel-softmax (ST-Gumbel), and a non-differentiable dueling Double-DQN (DDQN) bit-flip baseline. In fully NLOS scenarios, QAT-optimized RIS with NRX consistently outperforms least-squares and unoptimized baselines, and narrows the gap to a perfect-CSI reference under both staged and E2E training.

Ahmad Faisal Mirza, Messaoud Ahmed Ouameur, Mohammed Ali Dou et al. · 0 citations
Open access 2026

Digital Predistortion for Non-Differentiable Nonlinear Systems: A Reinforcement Learning Framework With Transfer Learning

The analysis demonstrates that the RL-based approach, enabled by an effective neural network initialization strategy, surpasses traditional methods and ML-based DPD schemes such as DLA and ILA and provides a scalable and efficient solution for compensating pattern-dependent nonlinearities in high-speed optical communications.

Arash Rabiepoor, L. Rusch, Ming Zeng · 0 citations
2026

A Reinforcement Learning-Based Scheduling Scheme for FSO and RF Hybrid Satellite-to-Ground Transmission Systems

Low Earth orbit (LEO) satellite-terrestrial communication systems grapple with significant challenges posed by their inherent dynamism and substantial transmission delays. To address these critical issues, this paper proposes a novel hybrid-medium transmission optimization framework that leverages high-altitude platforms (HAPs) as relays. Our primary objective is to minimize end-to-end system delay through the joint optimization of transmission mode selection and wireless communication resource allocation. The resulting joint optimization problem is formulated as a computationally intractable mixed-integer nonlinear programming (MINLP). We present a hierarchical solution strategy to tackle this complexity. Firstly, Lagrangian optimization is employed to analytically derive the intrinsic coupling between resource allocation and transmission mode selection, thereby simplifying the problem into a sequential decision-making process. This sequential problem is subsequently framed as a Markov decision process (MDP), enabling the design of a deep reinforcement learning (DRL) agent tasked with dynamically learning the optimal transmission mode selection policy. By maximizing cumulative long-term rewards, our DRL-based approach effectively reduces overall system delay, unlocking enhanced performance potential for future 6G networks.

Yi Huang, Jin Li, Yanwen Zhu et al. · 0 citations
Preprint Aug 2026

SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning

Reinforcement fine-tuning strategies for adapting Qwen2.5-3B-Base to graduate-level signal mathematical problems from WirelessMATHBench-XL are investigated and whether GSPO or GMPO offer advantages in stability or accuracy over GRPO for signal reasoning tasks is assessed.

Guozheng Sun · 0 citations