A generative search answer can cite a supported passage yet omit a source relationship that changes its interpretation. We specify a claim-gated audit of the query-source-answer tuple. An omission is resolved only when relationship evidence, answer adoption, materiality, and disclosure are all observed; incomplete evid...
Kai-Nan Zhou, Chu-Hong Xu, Gang-Zhen Qian et al.· 0 citations
The benchmark, failure-preserving contract, incident provenance, and governance controls needed to prevent infrastructure behavior from being misreported as model behavior and the benchmark, failure-preserving contract, incident provenance, and governance controls needed to prevent infrastructure behavior from being mi...
Hang Xiao, Chu-Hong Xu, Kai-Nan Zhou et al.· 0 citations
Retrieval-augmented generation (RAG) pipelines may omit a source's material relationship to the query. We study a pre-generation triage layer that treats this relationship as query dependent. The method routes canonical query families for enhanced review and assigns retrieved pages to pass, contextualize, exclude, or r...
Kai-Nan Zhou, Gang-Zhen Qian, Chu-Hong Xu et al.· 0 citations
A bounded IHEval comparison uses the same SmolLM2 checkpoint and output budget while preserving its published instruction roles and scorer, and the benchmark measures conditional task and output-contract success.
Kai-Nan Zhou, Zhao-Yi Li, Janet Sung et al.· 0 citations
Five host families are too few for a population claim, and the experiment says nothing about transfer to a new cell library or an industrial design, which supports a narrower conclusion: sibling benchmark variants can inflate apparent transfer.
Hang Xiao, Chu-Hong Xu, Kai-Nan Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.