Lara, a machine-checkable language and protocol for checking and revising support for research claims, is introduced, and the metatheory of claim checking and cross-context argument transport is established, and semantic guarantees in Lean 4 are mechanized.
Abstract
As autonomous AI agents take on every stage of scientific inquiry, research output is expanding far beyond human review capacity. Yet scientific communication still relies on natural-language prose: an informal medium prone to ambiguity, hidden assumptions, and untracked limitations that machines cannot reliably audit. We introduce Lara, a machine-checkable language and protocol for checking and revising support for research claims. By turning research arguments into executable artifacts, Lara provides an epistemic kernel for autonomous science: it enables automated validation pipelines for research agents, lets declared bridges connect arguments across papers into an auditable network, and allows both humans and machines to recheck the standing of an encoded claim in milliseconds. In a Lara program, authors explicitly declare their claims, supporting evidence and assumptions, and known objections or limitations. A lightweight, deterministic checker adjudicates these interactions, assigning each claim a reproducible status:"justified","defeated","contested", or"gap", which marks a claim whose support is incomplete and locates the unanswered question. Case studies cover empirical review, a philosophical debate without measurements, and the loss of support when an assumed axiom is withdrawn. We establish the metatheory of claim checking and cross-context argument transport, and mechanize the semantic guarantees in Lean 4 (roughly 117,000 lines), leaving three arguments on paper. The audited public metatheory is"sorry"-free and uses only Lean's three standard axioms; some executable examples additionally trust native evaluation.
Scientific Agent Skills, an open library of 163 such procedures in 16 areas of practice, including genomics, cheminformatics, medical imaging, study design and scientific communication, is presented.
T. Kassis, Vinayak Agarwal, Yu-Huan He et al.· 2 citations· ⚡1
It is argued that computational law can be used as a governance tool and that a desirable goal would be to formalize the law that can and ought to be programmatically executable.
AgentGuardUtil is presented, the authors' entry to CAR-bench Track~1, which treats the AI planer (LLM) as a fallible proposer inside a grounded verify-and-revise loop, and its core novelty is a runtime policy compiler.
R. Bouchekir, D. Safin, Tomas Bueno Momcilovic· 0 citations
AI Agents are increasingly deployed in real-world settings, where they interact with external tools and make sequential decisions with limited human oversight. This creates a pressing need for reliable and auditable explanations of what an agent did and why. However, traditional Explainable AI (XAI) methods fall short...
Vittoria Vineis, Fabiano Veglianti, Lorenzo Antonelli et al.· 0 citations
Large language model agents increasingly rely on compound programs for retrieval, tool use, reasoning, and verification, yet their failures often arise from local procedural decisions. Existing reinforcement-learning and prompt-optimization approaches typically rely on scalar rewards or repeatedly modify entire prompts...
Xu Liu, Wen-Zhang Wei, Jun Cao et al.· 0 citations
The bottleneck is scientific judgment rather than coding, and genuine discovery remains out of reach, so TruthInsightBench makes this gap a measurable target.
Zhi-Bo Yang, Chen Zhang, Yue-Wei Zhang et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.