Preprint
Jul 2026
Personalizing Large Language Model Agents with Small Policy Models
This work forms personalization of a frozen agent as online learning of a per-user execution policy from scalar feedback observed only for the executed action, and proposes FABLE (Factorized Adaptive Bandit Layer for Execution), a lightweight policy layer outside a potentially black-box host agent.
Dian Jin, Zhi Zhang, Huichao Li et al.
· 0 citations