Overall, it is found that probing is an effective means to catch a range of different tool-calling errors, including errors arising from using an argument that has the wrong value but the correct type, which might not be recorded by standard logging frameworks.
Eric C. Yeats, Brendan Kennedy, Loc Truong et al.· 0 citations
This work compares Turkish document question answering across three chunking strategies, five embedding models, and two LLMs, over three documents with contrasting layouts, finding the faster LLM is not the more accurate one.
Mustafa Sertac Turkel, Fatma Nur Korkmaz, Ahmet Tugrul Bayrak· 0 citations
Evaluating 13-21 models across six presentation operationalizations and four task-domain operationalizations suggests that, despite confounds, some models possess practical SGTR capabilities, and that SGTR should be monitored and considered in the design of safety-critical AI applications.
J. St-Amand, Callum Canavan, S. Imran et al.· 0 citations
This study explores and evaluates the ability of LLMs to follow and enhance human mental trajectories during semantic memory search and demonstrates that an LLM's abilities to track and predict human memory trajectories in this task exceed those of other humans.
Eric Lacosse, Mariana Duarte, Graham Todd et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
A consolidated memory that states a decision constraint and whose source record has since been superseded by a record that withdraws it is studied: provenance is immutable, the current record has changed, and the memory is stale.
This work introduces MathAdv, a diagnostic benchmark spanning 13 domains across undergraduate- and graduate-level mathematics, and shows how component-wise evaluation can reveal model capabilities and failure modes that aggregate theorem-proving accuracy obscures.
Jiajie Yuan, Connor Martinez Lockhart, Xiao-Yun Liu et al.· 0 citations
JuryProbe is introduced, an empirical consensus-risk diagnostic for reference-free factuality judge panels, paired with a calibration-based routing policy, which estimates consensus risk from a labeled calibration probe using false-negative-only (FN-only) judge correlation and false-consensus lift.
This work introduces Sieve, a search-inspect-fetch strategy driven by a Boolean Query Language (BQL): it searches webpage fields to filter candidates, uses an interchangeable ranker to order them, presents structure-rich result cards for inspection, and fetches only selected sections.
Shuai Wang, Haodong Chen, Yu Yin et al.· 2 citations
Tail subtraction is introduced, which removes shared prompt and continuation semantics from boundary states and yields cleaner, more stable steering signals, and suggests that steering depends on representations of what the model is about to do, not merely on what has already appeared.
Jiaran Ye, Lingxu Ran, Zijun Yao et al.· arXiv.org· 2 citations
This work investigates the ability of Large Language Models to generate structurally valid and constraint-compliant network topologies through a constraint-driven pipeline combining hierarchical modeling and systematic validation, and provides a systematic benchmark for understanding how LLMs handle structural and resilience constraints in topology synthesis.
This work borrows definitions from narratology to analyze eight intricate dimensions of character, such as stylization and wholeness, which consider more than just basic characteristics of characters within LLM and human-written stories.
A. Brei, Abhisheik Sharma, Nicholas Sanaie et al.· Annual Meeting of the Associ...· 0 citations
TokenPilot is presented, a dual-granularity context management framework that reduces costs by 61% and 56% in isolated mode, and 61% and 87% in continuous mode, while maintaining competitive performance compared to prior systems.
Buqiang Xu, Z. Xue, Dian Chen et al.· arXiv.org· 1 citation
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.