#artificial intelligence
Apr 2026
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
This work introduces CodeRQ-Bench, the first benchmark for evaluating LLM reasoning quality across three coding task categories: generation, summarization, and classification, and proposes VERA, a two-stage evaluator that combines evidence-grounded verification with ambiguity-aware score correction.
Yuangang Li, Justin Tian Jin Chen, Ethan Yu et al.
· arXiv.org · 1 citation