Learning from the Gap Between Pass@K and Pass@1
GapFT is introduced, which selects training evidence by the source checkpoint's single-sample outcome and fine-tunes on the Pass@K-Pass@1 gap: problems the policy fails on one sample but solves within K samples, and its analysis relates available gains to transferable failure support.
Xuan-Wen Liu, Jing Qian, Hao-Sheng Chen
· 0 citations