Using frontier agent harnesses to tackle mathematical research problems has emerged as an effective means of advancing mathematics. However, solving frontier problems in mathematics may require a massive number of agents working in parallel for extended periods to construct proofs, thereby generating an enormous volume...
Jun-Yi Zhang, Jin-Xi Yu, E. Jiang et al.· 0 citations
Planned Test-Time Scaling (PTTS) provides a general framework for improving test-time scaling by coordinating reasoning branches, with zero-shot and trainable instantiations that yield substantial performance gains.
Xue-Qing Wu, Lang-Xing Bai, Hritik Bansal et al.· 0 citations
MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements, is presented and PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop is introduced.
Xue-Qing Wu, Ashwin Balasubramanian, Bingxuan Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.