This work proposes AIPO, an enhanced reinforcement learning framework that improves LLM reasoning through active multi-agent interaction during exploration and consistently improves reasoning performance and generalizes across different policy models, collaborator backbones, and RLVR algorithms.
Junnan Liu, Linhao Luo, Thuy-Trang Vu et al.· arXiv.org· 0 citations
Conformal Privacy Auditing is introduced, a distribution-free calibration framework that provides a statistical certificate of re-identification risk for each released document against LLM-empowered adversaries and enables audits of open-source models and proprietary API models in a unified framework.
Shuo Huang, G. Haffari, Xing-Liang Yuan et al.· 0 citations
Compilable Academic Document Parsing (CADP) is proposed, a paradigm that reconstructs a full page as contextual \LaTeX{} plus executable Python, so that structure-preserving elements and executable chart representations can be reconstructed, recompiled, and directly verified against the source page.
Rihui Jin, Jun Wang, Chen Zhu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.