Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an important tension. The same operation can rescue a failing trajectory or disrupt one that woul...
Shu-Yao Xiao, Sheng-Ling Wang, Xuan Chen et al.· 0 citations
Deep Research agents synthesize evidence into cited reports, yet a well-cited report can still reach a misleading conclusion. Citation correctness checks whether cited sources support individual claims. It does not show whether adaptive search exposed a representative view of all documents made available for evaluation...
Shu-Yao Xiao, Sheng-Ling Wang, Xuan Chen et al.· 0 citations
An updated model can improve an aggregate metric while degrading a slice that matters to a downstream user. We study checkpoint selection subject to non-regression tolerances relative to a retained incumbent. The central distinction is between failing to detect harm and certifying non-inferiority: the former can releas...
A structural causal model (SCM)-based framework for cross-turn error propagation in memory-augmented LLMs is proposed, and experiments show that error influence generally decays with interaction distance, while the memory-update pathway contributes more persistent effects than question feedback.
Shu-Yao Xiao, Sheng-Ling Wang, Xuan Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.