Enhancing Stable Behavioral Imitation through Adaptive Reward Weighting in TD3-SAC-GAIL
The results demonstrate the potential of adaptive reward weighting to provide a systematic mechanism for controlling the exploration–imitation trade-off and enhancing the stability and robustness of GAIL-based policy learning while retaining the exploration advantages of the TD3-SAC hybrid framework.
Mehran Ali, Zia Ullah, Aliza Ashfaq
· Journal of Engineering and C... · 0 citations