Group Relative Policy Optimization (GRPO) has shown strong effectiveness in reinforcement learning from verifiable feedback, where sampled rollouts can be compared within a group using task-provided correctness signals. However, extending group-relative optimization beyond verifiable settings is challenging because suc...
This empirical study investigates whether procedural LLM skills can be represented as directions in activation space and whether vector-space operations over these directions can express skill-level behaviors and finds that procedural skills admit a vector-space representation.
Xun-Yi Jiang, Junda Wu, Yuxin Xiong et al.· 0 citations
This work treats each reasoning trace as a sequence of latent states rather than an unstructured texts, and investigates whether inference time interventions can provide fine-grained control over the self-looping reasoning process.
The first MuseCP evaluation framework is introduced that covers four categories of music facets with fine-grained and well-tailored metrics to capture nuanced changes in music attributes and hopes it can offer practical guidance for developing more effective and reliable music editing strategies with strong MuseCP capa...
Yash Vishe, Eric Xue, Xunyi Jiang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.