A pool of language models can collaborate and improve collectively by learning from one another's responses. These interactions depend on the instructions used during training. Existing methods typically sample instructions uniformly, even though their usefulness may change as the models improve: an instruction on whic...
Christina Hahn, Shang-Bin Feng, Dean Light et al.· 0 citations
It is argued that it is difficult to predict the alignment state of a model after training by inspecting the training data alone, which highlights the importance of comprehensive post-training alignment evaluations.
Yavuz Faruk Bakman, D. Yaldiz, Baris Askin et al.· 0 citations
In Mixture-of-Experts language models, the router typically selects and weights experts based on the token's hidden state, utilizing limited contextual information. We propose Attention-Aware Routing (AAR), which augments the router with temporal and spectral features extracted from a sliding window of attention weight...
Despoina Kosmopoulou, Anastasios Tsetsilas, Efthymios Georgiou et al.· 0 citations
ScalePRM, which scales verification compute as an alternative to ground-truth supervision for training process reward models, generates multiple independent verifications of each reasoning step and aggregate their judgments to produce synthetic step-level labels without ground truth.
Salman Rahman, Sruthi Gorantla, Arpit Gupta et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.