Skip to content

Author

Sophia Xiao Pu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning

Inference-time decoding methods improve LLM reasoning by exploring multiple candidate trajectories, yet treat each trajectory as atomic: either retaining it whole or discarding it irreversibly. This wastes computation on partially promising candidates whose high-quality prefixes are abandoned alongside degraded suffixe...

Sophia Xiao Pu, Yumo Xu, Sailik Sengupta et al. · 0 citations
#machine learning Preprint Sep 2026

CompassPlay: Rewarding the Proposer for Where It Moves the Solver

In self-play, a proposer generates verifiable tasks to train a solver. Proposer rewards often depend on the solver's success rate, but equally difficult tasks can differ in their training value. We introduce CompassPlay, a self-play method that rewards the proposer through gradient alignment. The reward favors tasks wh...

Sophia Xiao Pu, Xi-Meng Sun, Jiang Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.