Aug 2026
Evaluating LLMs in ROS robotic software code generation
This work constructs ROSDevEval, a specialized benchmark comprising 240 real-world ROS programming tasks, and conducts a comprehensive empirical study to evaluate four state-of-the-art LLMs alongside a specialized code assistant (GitHub Copilot), revealing a severe domain capability gap.
Yuxin Zhao, Xinjun Mao, Tanghaoran Zhang et al.
· Empirical Software Engineeri... · 0 citations