This work extends PubHealthBench, a question answering benchmark of 7,929 questions derived from UK Government public health guidance, into a retrieval-augmented setting and systematically evaluates retrieval and generation choices, and introduces a rubric-based LLM-as-a-judge covering faithfulness, completeness, clarity, and factual consistency.
Felix Feldman, Joshua Harris, Timothy Laurence et al.· 0 citations