Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an agent may fail. An agent reconciling a long ledger, for example, must repeatedly read its...
Jeffrey Willette, Krishna C. Puvvada, Boris Ginsburg· 0 citations
While autonomous software engineering (SWE) agents achieve high benchmark resolution rates, these scores can mask exploitative behaviors---such as leveraging local Git histories, accessing upstream repositories, or recalling memorized solutions---rather than demonstrating genuine problem solving. We systematize and aud...
Nikolai Ludwig, W. Ahmad, Somshubra Majumdar et al.· 0 citations
Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best, and transfers across corrector families and adds zero parameters to the inference path.
Chao-Han Huck Yang, Zih-Ching Chen, Piotr Żelasko et al.· arXiv.org· 0 citations
This work proposes an efficient method to add streaming ASR capabilities to an existing duplex S2S model by introducing a lightweight ASR head in parallel to the agent text head, enabling real-time user transcription while preserving full-duplex conversational capabilities including turn-taking and barge-in handling.
Ke Hu, Nourchene Ferchichi, Edresson Casanova et al.· 1 citation
A practical training recipe for the normalized Transformer and its evaluation on modern hybrid Mamba-2--Transformer Mixture-of-Experts models shows that the 30B-total-parameter nGPT model reaches the same validation loss using approximately half as many training tokens.
This work introduces SHERLOC (Structured Hypothesis-driven Exploration and Reasoning for Localization), a training-free framework pairing a reasoning LLM with compact repository tools and self-recovery, without fine-tuning or multi-agent orchestration.
Hovhannes Tamoyan, Sean Narenthiran, Erik Arakelyan et al.· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.