Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is...
Next-token prediction has driven remarkable progress in large language models, yet a growing body of evidence suggests that they can struggle on problems governed by complex global constraints. In this work, we focus on this regime and ask whether some of these limitations arise from the inference interface induced by...
Woosang Jeon, Jaeyeon Kim, S. Kakade et al.· 0 citations
The key to the proof is a new framework for quantifying post-measurement damage, based on the quantum Efron-Stein decomposition, which improves all three exponents even in the Offline Shadow Tomography setting.
Si-Tan Chen, R. O'Donnell, Angelos Pelecanos et al.· arXiv.org· 2 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.