Saturation-Insensitive Dueling Bandits with General Function Approximation
SI-CDB is introduced, an algorithm that selects opponent arms using a carefully designed heuristic for arm selection that enables saturation-insensitive reward learning and recovers the near-optimal dependence for linear reward classes, eliminating the unfavorable $1/\sigma'(\cdot)$ factor.
Cheng-Gong Zhang, Xu-Heng Li, Qi-Wei Di et al.
· 0 citations