Skip to content

Author

Saiyong Yang

We have 4 of 21 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR

Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges...

Kun Liang, Chenming Tang, Clive Bai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Hunyuan-A13B Technical Report

We present Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost. The model is pretrained on a rigorously filtered 20T...

Tencent Hunyuan Team, Ao Liu, Bo Zhou et al. · 0 citations
Preprint Aug 2026

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not charac...

Xinming Wang, Wei-Nong Wang, Hongming Yang et al. · 1 citation · ⚡1
#natural language process... Preprint Aug 2026

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

This work compares three fusion paradigms by the artefacts they reuse and suggests that Merge when experts already exist and cheap fusion is paramount; Mix RL when training a unified model without experts, with domain proportions adjusted for cross-domain transfer; and MOPD when preserving domain-specific gains matters...

Sicheng Wu, Kai Yang, Yuchen Cai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.