SureRoute is introduced, a chemical verifier-anchored retrosynthesis platform that suppresses Chemical Hallucination and shows that reliable scientific AI requires not only strong generation, but executable verification.
Abstract
AI models, including large language models, are increasingly integrated into scientific discovery workflows, yet they remain prone to hallucination. In experimental sciences, such errors translate directly into failed wet-lab validations and wasted resources; in self-improving agentic systems, confident errors risk being reinforced rather than corrected. Retrosynthesis provides a representative example of this failure mode: existing models can generate chemically plausible routes, but cannot reliably determine which routes are experimentally feasible. We define \textbf{Chemical Hallucination} as a route that appears valid yet fails under competing reactive sites, unresolved selectivity, or missing mechanistic support, a failure largely invisible to the Recall@$K$ metric. We introduce \textbf{SureRoute}, a chemical verifier-anchored retrosynthesis platform that suppresses Chemical Hallucination. SureRoute combines a multi-model ensemble, data asset retrieval, and \textbf{ChemHarness}, an executable chemical intuition engine for route verification and reliability-first ranking. On a benchmark of 350 real-world industrial targets, SureRoute reaches 74.3\% recall@1, 2.2--3.5$\times$ that of seven single-step models and three frontier LLMs, while cutting top-1 Chemical Hallucination to 4.6\%, a 4--6$\times$ reduction relative to frontier LLMs. As a model-agnostic reranker, ChemHarness drives detectable hallucination toward near-zero across arbitrary backbone candidates. SureRoute shows that reliable scientific AI requires not only strong generation, but executable verification.
A self-reflective framework in which an LLM generates an answer, identifies claims that may be uncertain, performs an internal verification stage, and revises the response before delivery is proposed.
Priti Sharma, Sachin Sharma· Iconic research and engineer...· 0 citations
A concise two-axis framework that integrates an “intrinsic-extrinsic” distinction in source attribution introduced by Ji et al. with a “faithfulness-factuality” distinction in contextual grounding surveyed is presented, yielding four clearly defined hallucination types applicable across tasks, modalities and architectu...
Misbah Khan, Preston Billion-Polak, T. Khoshgoftaar· IEEE Access· 0 citations
HallDetect, a lightweight, reference-free, and black-box framework for hallucination detection, is presented, a lightweight, reference-free, and black-box framework for hallucination detection that is evaluated not only on summarization but across a broader range of source-grounded generation settings.
Achir Oukelmoun, N. Semmar, Gaël de Chalendar· 0 citations
The Latent Critic is introduced, a lightweight low-rank adapter that operates concurrently with a frozen base LLM's generation to actively restructure the transformer's residual stream---amplifying latent grounding signals and translating them into localized, natural language feedback within a single sequence.
This work introduces a two-stage keyword-perturbation method for hallucination detection and extends the same probabilistic framework to four hallucination regimes: knowledge deficit, wrong knowledge, context distraction, and unstable inference.
Xu-Han Tong, Hao-Yue Bai, Da-Wei Zhou et al.· 0 citations
Interpretable machine learning for Large Language Models (LLMs) increasingly relies on sparse probing methods that identify small sets of neurons claimed to detect and causally influence behaviors such as factuality recall, safety alignment, and hallucination. These claims have important implications for model auditing...
Huseyin Cavus, Sebin Sabu, J. Spear et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.