Skip to content

Thinking with Looped Flows

Sep 2026 · 4 citations · 61 references
Computer Science

TL;DR

Looped flows are proposed, an approach that allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples.

Abstract

Humans and machines often solve harder problems by spending more time on computation. In deep learning, looped models implement this idea during inference by recurrently updating a hidden state. In practice, however, their training backpropagates through only one or a few updates, making it hard to train early updates to support future ones. We propose looped flows, an approach that sidesteps this issue by training the recurrence with local denoising objectives. By imposing temporal association across denoising objectives through progressively decreasing noise levels and shared noise, the model is incentivized to learn recurrent states that transfer useful computation over time, even when gradients cover only a few updates. We then formulate inference as integrating the velocity of a probability flow parameterized by the learned denoiser, coupled with recurrent states. This allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples. Across six reasoning benchmarks including two multi-solution benchmarks, looped flows outperform prior state-of-the-art looped models overall, achieving 58.8% test accuracy on ARC-AGI-1 and 12.2% on ARC-AGI-2.

View source

Similar papers

#machine learning Preprint Sep 2026

RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory

RecurTrace introduces Loop Memory Attention, which lets each looped layer attend to its own states from previous iterations along the loop-time axis, so the model can revisit earlier computations instead of relying on the latest state alone.

Yu-Xiang Wang, Kun-Yu Feng, Ying-Da Shen et al. · 3 citations
Preprint Aug 2026

Steering Recurrent Reasoners at Inference Time with Readout Feedback

Readout Feedback (RoFB), a test-time intervention that converts intermediate predictions into token-wise pairwise coupling forces injected into the latent dynamics, is introduced, suggesting that closed-loop steering of latent dynamics can serve as a complementary inference-time control mechanism for recurrent reasonin...

Shunsuke Kamiya, Masanori Koyama, Seongcheol Jeong et al. · 0 citations
#machine learning Preprint Sep 2026

Looped Actor: Depth-Recurrent Reasoning Models for Reinforcement Learning

Looped reasoning models repeatedly apply a shared set of parameters, enabling more computation without increasing the model size. These models also support input-dependent computation by dynamically deciding when to stop looping. Motivated by the recent success of looped transformers in language modeling and reasoning,...

T. Konstantin Rusch, T. Seyde, J. Boyer et al. · 0 citations
Preprint Aug 2026

Think Shallow, Solve Deep: Controlling Recurrent Dynamics for Reliable Test-Time Depth

This work gives a sufficient condition for depth-safety: once an operator's per-step displacement is small relative to the decoder margin, the decoded answer cannot change under further iterations, and gives four operational criteria for useful test-time depth.

Ivan Viakhirev, Kirill Borodin, Amirah Almutairi et al. · 5 citations
#machine learning Preprint Sep 2026

Blackboard Intelligence Can Surpass Autoregressive on Globally Constrained Problems

Next-token prediction has driven remarkable progress in large language models, yet a growing body of evidence suggests that they can struggle on problems governed by complex global constraints. In this work, we focus on this regime and ask whether some of these limitations arise from the inference interface induced by...

Woosang Jeon, Jaeyeon Kim, S. Kakade et al. · 0 citations
#machine learning Preprint Sep 2026

Scheduling Recursive Reasoning in Looped Transformers

Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with shared parameters. However, these models apply each learned update with a fixed unit scale, which can be conservative when updates make persistent progress and overly aggre...

Bo-Yuan Wang, Cheng-Yao Yu, Jia-Xi Ren et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.