Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint May 2026

Diversifying RLVR Rollouts via First-Token Exploration

This work identifies the first token of the response as a structurally distinct target for diversification, largely overlooked in prior work, and finds that the first-token distribution is sharply concentrated and only weakly related to downstream correctness, as lower-probability candidates can yield similarly accurat...

Soeun Kim, Albert No · 2 citations
#natural language process... Preprint Sep 2026

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

This work shows that paraphrase robustness is a core requirement for reliable VLM-based reward modeling in robotics and introduces ROBORMBENCH, a benchmark with 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases spanning lexical, syntactic, and action-goal rewrites.

Wonje Jeung, Sangyeon Yoon, Hyesoo Hong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.