Modern state space models (SSMs), such as Mamba2, provide a compelling alternative to transformers by combining linear-time sequence modeling with recurrent state-space dynamics. However, the behavior of SSMs in distributed learning settings remains poorly understood. In particular, the existing standard federated lear...
Adam Piaseczny, Md Kamran Chowdhury Shisher, Shi-Qiang Wang et al.· 0 citations
A large language model (LLM) agent that inherits a plan through shared memory can hold the latest requirement yet act on a plan derived from an older one: fresh memory, stale plan. Freshness checks miss this failure because they compare local copies with current state (observation currency) rather than the inputs the p...
Evan Chen, Shi-Qiang Wang, Christopher G. Brinton· 0 citations
Fire (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning, is proposed, which provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where s...
Seohyun Lee, Dong-Jun Han, Seyyedali Hosseinalipour et al.· 0 citations
Tory of Scene (ToS), a training-free reasoning schema in which each agent reads its public role, the only difference between the agents, and the task context they all observe, is proposed, which outperforms all six baselines on every benchmark, and each baseline falls far behind it in at least one setting.
Liang-Qi Yuan, Wen-Zhi Fang, Shi-Qiang Wang et al.· 0 citations
Task-wise federated multi-objective optimization (FedMOO) trains a shared model for competing prediction objectives under heterogeneous data, partial participation, and communication constraints. Existing methods commonly derive task weights from gradient or update geometry. This requires task-specific information or i...
Amirardalan Dehghanpour, Seyed Mohammad Azimi-Abarghouyi, Christopher G. Brinton· 0 citations
This work proposes a two-stage generator-in-the-loop alignment framework that consistently outperforms rank-order, random, and REPLUG-style likelihood baselines under various alignment losses and pool size settings, suggesting that answer-level generator feedback is an effective supervision signal for preference alignm...
Zhang-Yu Chang, Dong-Jun Han, Seyyedali Hosseinalipour et al.· 0 citations
P LAN F ENCE is introduced, a dependency-scoped action-validation protocol for distributed LLM-agent teams that avoids repeated update-path coordination as churn grows and avoids validating unrelated state as the shared keyspace grows.
Evan Chen, Shi-Qiang Wang, Christopher G. Brinton· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.