Preprint
Jul 2026
PARTREP: Learning What to Repeat for Decoder-only LLMs
A lightweight gate is trained that predicts high-NLL tokens from early-layer hidden states, enabling token selection during mid-prefill via early exit, motivated by the hypothesis that less predictable tokens are less recoverable from surrounding context and therefore benefit more from late-position repetition.
Andikawati P Widjaja, Yongjun Kim, Hyounghun Kim et al.
· 0 citations