Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

This work introduces CodeRQ-Bench, the first benchmark for evaluating LLM reasoning quality across three coding task categories: generation, summarization, and classification, and proposes VERA, a two-stage evaluator that combines evidence-grounded verification with ambiguity-aware score correction.

Yuangang Li, Justin Tian Jin Chen, Ethan Yu et al. · 1 citation