Temporal-Synergistic Policy Optimization for Unsupervised Low-Light Image Enhancement
This work proposes an unsupervised low-light enhancement framework based on Group Relative Policy Optimization (GRPO), which utilizes perceptual preferences to directly optimize the diffusion policy and introduces a Sliding Window Hybrid ODE-SDE Sampling strategy that confines stochasticity to dynamic sub-intervals, thereby achieving efficient coarse-to-fine exploration and precise advantage attribution.
Yuanfei Bao, Dong Li, Jie Huang et al.
· 0 citations