Next-token prediction has driven remarkable progress in large language models, yet a growing body of evidence suggests that they can struggle on problems governed by complex global constraints. In this work, we focus on this regime and ask whether some of these limitations arise from the inference interface induced by...
Woosang Jeon, Jaeyeon Kim, S. Kakade et al.· 0 citations
The theoretical results show that GN's optimality for linear regression no longer holds under a poorly chosen basis under both population and stochastic updates, or a move from linear regression to a non-convex variant of logistic regression, and Adam-style methods offer genuine advantages over curvature-inspired preco...
Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is, and how accurately it describes the true optimization dynamics. In this work, we stress test the simplest possible model...
Alexandru Meterez, Pranav Ajit Nair, Depen Morwani et al.· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.