Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an important tension. The same operation can rescue a failing trajectory or disrupt one that woul...
Shu-Yao Xiao, Sheng-Ling Wang, Xuan Chen et al.· 0 citations
Zeroth-order (ZO) optimization with SGD in random subspaces enables memory-efficient fine-tuning of large language models without backpropagation. However, high gradient estimation noise fundamentally undermines adaptive optimizers like Adam. We propose SubZero+, which achieves practical adaptive ZO optimization throug...
Zi-Ming Yu, Shu-Yao Xiao, Xingyu Zhao et al.· 0 citations
Deep Research agents synthesize evidence into cited reports, yet a well-cited report can still reach a misleading conclusion. Citation correctness checks whether cited sources support individual claims. It does not show whether adaptive search exposed a representative view of all documents made available for evaluation...
Shu-Yao Xiao, Sheng-Ling Wang, Xuan Chen et al.· 0 citations
Experiments show that SubZero+, an improved SubZero framework that improves stability in three complementary ways, consistently outperforms prior ZO baselines, enlarges the stable learning-rate range, and narrows the gap to first-order methods with minimal extra memory overhead.
Ziming Yu, Shu-Yao Xiao, Xingyu Zhao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.