Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advanta...
Mu-Yang Li, Jie Yang, Zheng-Yu Fang et al.· 1 citation· ⚡1
It is shown that preference alignment preserves the human response distribution only under a restrictive condition, and no consistent evidence that real human preferences satisfy it, and human-likeness is established as an explicit dimension of alignment rather than something assumed to follow from preference alignment...
Su-Qin Yuan, Runqi Lin, Mu-Yang Li et al.· 0 citations
Label Wave is proposed, which does not require validation data for selecting the desired model across various weakly supervised learning paradigms, including learning with noisy labels (LNL), positive-unlabeled learning, and unlabeled-unlabeled learning.
Suqin Yuan, Muyang Li, Lei Feng et al.· IEEE Transactions on Pattern...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.