Preprint
Jul 2026
Selective Left-Shift: Turning Test-Time Compute and Difficulty-based Curation into Training Data for Low-Resource Code Generation
This work proposes a three-phase pipeline that resolves this trilemma by decoupling syntax acquisition from algorithmic reasoning and applies Reinforcement Learning with Verifiable Reward~(RLVR) grounded by language-agnostic Input/Output tests.
D. Samaraweera, Anjana Supun, Srinath Perera
· 0 citations