Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the model, scaling poorly on problems the model is unlikely to sol...
Joseph Rance, Fabio Pizzati, Juil Sock et al.· 0 citations
We present PartiCam, a training-free Particle filtering rooted method for improved Camera controlled video generation. Generating videos that follow a precisely specified camera trajectory remains challenging for large video diffusion models. Training-free approaches are backbone-agnostic and avoid the need to construc...
Amine Ouasfi, Runjia Li, Junlin Han et al.· 0 citations
LatticeMind is presented, a conflict-aware structured memory that handles contradiction at write time, which maintains explicit item status, applies cheap symbolic conflict checks, and invokes LLM reconciliation only for unresolved semantic cases.
Heng Zhou, Lian Zhang, Yutao Fan et al.· 0 citations
Policymakers should urgently obtain more visibility into the automation of AI R&D, develop ways to steer and constrain an intelligence explosion, and prepare society to adapt to an intelligence explosion's impacts.
Alan Chan, Christoph Winter, A. Barto et al.· 0 citations
StraTA is a simple framework that introduces an explicit trajectory-level strategy into agentic reinforcement learning (RL) and trains strategy generation and action execution jointly with a hierarchical GRPO-style rollout design, further enhanced by diverse strategy rollout and critical self-judgment.
Xiangyuan Xue, Yifan Zhou, Zidong Wang et al.· arXiv.org· 1 citation
This work introduces WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories, and builds Blender generators that randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's c...
Hao-Tian Zhang, Feng-Yuan Yu, Dezhi Luo et al.· 0 citations
Adversarial Probing for Implicit VulnErabilities (AdvPIE), a multimodal agentic framework to expose implicit vulnerabilities without requiring access to the parameters of target models, is proposed.
Chang-Sha Ma, Junlin Han, Shuo Chen et al.· 0 citations
It is found that procedural pretraining can improve molecular property prediction even after subsequent molecular pretraining, and procedural data can provide transferable structure for molecular learning and offer a complementary route to improving performance when labeled molecular data are limited.
M. Friedemann, Zachary Shinnick, Philip H. S. Torr et al.· 0 citations
Autoresearch agents are reshaping the research ecosystem, but they can also let flawed claims enter the literature at scale. Human advisors catch such issues in drafts through careful, traceable feedback, yet advisor-style assessment requires extensive manual effort and does not scale. To shift automated paper assessme...
Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is insufficient. However, existing reward models are largely optimised for general preferences rather than contextual grounding, limiting their a...
Rilton Franzone, Valentin Noël, Pu-Yu Wang et al.· 0 citations
A taxonomy of medical hallucination types and a clinician-validated error-injection pipeline that creates matched correct and error-injected responses are developed, finding that more specific rubrics better distinguish correct from hallucinated responses.
Griffin Farrow, Lily Sijia Li, Jack Johnson et al.· 0 citations
GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), is introduced, revealing the substantial gap between current agent capabilities and those required for complex...
Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.