Preprint
Jul 2026
In-Context Learning as Implicit Policy Gradient
It is shown that score-conditioned In-Context Learning (ICL) admits a structural correspondence to policy gradient optimization, and an exact upper bound on the distribution shift induced by a bounded attention update is derived, yielding a trust-region-like analogy to KL-constrained policy optimization.
Masahiro Kaneko, Timothy Baldwin
· 0 citations