Preprint
Aug 2026
Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
P-Bench is built, a benchmark comprising 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine and introduces Fisher-R1, an open-weight LLM agent trained for rigorous hypothesis testing using synthetic tasks and reinforcement learning.
Jia-Cheng Miao, Jin Mu, Guanhua Chen et al.
· 0 citations