Muon motivates designing optimizer geometry around the function of each parameter block and uses the spectral norm for hidden linear layers. For the language-model head, the spectral norm is not a faithful measure of functional change. Softmax removes shared logit shifts, whereas the spectral norm can assign arbitraril...
Aditya Somasundaram, Charles Guille-Escuret, Alexander Moreno et al.· 0 citations
Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models (VLMs) reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, we introduce WM-VLM,...
Yuheng Zha, Yi-Lei Wang, Qiyue Gao et al.· 0 citations
We introduce Chronosphere, a spatio-temporal neural field that learns representations of climate. A central challenge in geographic representation learning is modeling environmental processes whose spatial and temporal complexity varies widely. Yet existing location encoders typically fix a single level of detail every...
Dan Cher, Eric P. Xing, Ke-Xing Li et al.· 0 citations
Robots operating in open environments act under partial observability, physical constraints, and dynamic task contexts. Beyond mapping observations and language instructions to actions, they must anticipate how candidate actions may affect future states and task-relevant outcomes. Recent advances in world models, video...
Zu-Xing Lu, Hong-Jia Zhai, Guan-Zhi Wang et al.· 0 citations
Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affect hypothesis quality...
Jie-Yuan Liu, Meng-Zhou Hu, Jefferson Chen et al.· 0 citations
Most of mathematical knowledge has been communicated through so-called informal use of mathematics and natural language. With large language models (LLMs) being highly adept in using natural language, they achieve strong performance, yet not perfect, in informal mathematical reasoning. Restraining LLMs to informal reas...
Joshua Ong Jun Leang, Haonan Li, Zheng-Yang Zhao et al.· 0 citations
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion...
S. Sahoo, Ling-Jie Chen, Khiem Pham et al.· 1 citation
Interpolating between these models, Raven is introduced, a linear-time sequence model that maintains a fixed set of memory slots and, at each step, decays and updates only a selected subset via learned, input-dependent routing, thereby preserving long-range content much more effectively.
Arshia Afzal, Aviv Bick, Eric P. Xing et al.· arXiv.org· 5 citations
GB.GeneUnet, an 837M-parameter transformer-based U-Net pretrained on 6 trillion tokens from multi-species genomes in OpenGenome2 is introduced, extending genomic context to 1 Mb with up to 100× inference speedup over GeneMoE, a preliminary MoE transformer baseline of similar model size pretrained on the same data.
Ning Sun, William de Vazelhes, Pan Li et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.