In self-play, a proposer generates verifiable tasks to train a solver. Proposer rewards often depend on the solver's success rate, but equally difficult tasks can differ in their training value. We introduce CompassPlay, a self-play method that rewards the proposer through gradient alignment. The reward favors tasks wh...
Sophia Xiao Pu, Xi-Meng Sun, Jiang Liu et al.· 0 citations
We introduce AMDKernelVault, an open HIP and Triton kernel corpus and training framework for recent AMD CDNA GPUs. Existing LLM-based kernel agents are largely CUDA/NVIDIA-centric and often depend on repeated frontier-LLM calls for generation, reflection, and optimization. To address this gap, we develop HIPKernelGen a...
Ji Liu, S. Majumder, Yi-Qing Huang et al.· 0 citations
Instella-MoE is introduced, a fully open Mixture-of-Experts (MoE) language model with 16 billion total parameters and 2.8 billion active parameters per token trained entirely from scratch on AMD Instinct MI300X and MI325X GPUs, establishing a strong, fully open foundation for efficient, high-performing MoE models and r...
Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.