We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separates global planning from detailed investigation and coordinates revision at the section level. Multiple planning agents f...
Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and s...
Chia-yuan Chang, Ren-Yuan Cheng, Rui Feng et al.· 0 citations
We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evalua...
Li-Qin Ye, Hao-Rui Wang, Fardin Ahmed et al.· 0 citations
QUBRIC, a framework that co-designs queries and rubrics can make rubric-based RL a practical complement to RLVR beyond strictly verifiable tasks, provides evidence that co-designing queries and rubrics can make rubrics a practical complement to RLVR beyond strictly verifiable tasks.