When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study
A supervised ensembling framework that trains a classifier over heterogeneous UQ-based scorer outputs on a small, domain-specific dataset of labeled LLM responses, then applies it to out-of-sample hallucination classification without retrieval, tools, or reference documents is studied.
Mohit Singh Chauhan, Vipin Gyanchandani, Dylan Bouchard
· 0 citations