Stochastic Autoregressive Learning
It is shown that stochastic autoregressive learning fundamentally differs from the deterministic theory, and that CoT learning at scale $\varepsilon$ is upper-bounded by base learning at scale $\varepsilon/M^2$, whereas e2e learning at scale $\varepsilon$ is upper-bounded, up to logarithmic factors, by $(M/\varepsilon)...