Preprint
Jul 2026
Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models
Empirical results demonstrate that the Reinforcement Learning with Verifiable Rewards framework enables LLMs to transition from passive responders to autonomous assistants, and demonstrates comparable performance to larger and reasoning-enhanced baselines, while RAGES proves superior to vanilla LLMs in generating biologically plausible clinical feedback.
Shengyi Hua, Kangzhe Hu, Conghui He et al.
· 1 citation