Muon improves large-scale training by applying a spectral-norm steepest-descent update to matrix parameters, but practical models also contain parameter blocks that do not fit dense-matrix geometry. One important case is the tied vocabulary table, which appears in language models and other token generators and can rece...
Arash Lagzian, Paniz Halvachi, Jun-Ming Zhang et al.· 0 citations
Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings.
Lei Bai, Jiaqi Cao, Chiyu Chen et al.· 3 citations
This work presents Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens, demonstrating that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.
Ru-Bin Wei, Jiaqi Cao, Jiarui Wang et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.