Skip to content

Author

Chris Russell

5 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

SEES: A Self-Evolving Embodied System via Failure-Guided VLA Policy Adaptation

Recent vision-language-action (VLA) policies demonstrate promising generalization across diverse short-horizon tasks. However, they remain unreliable on long-horizon tasks, partly because the large-scale training data is biased toward single-stage manipulation tasks that are cheaper to demonstrate. A single weak atomic...

Zi-Wen Li, Hanlue Zhang, Zhen-Yang Ren et al. · 0 citations
#natural language process... Preprint Sep 2026

Settle: Learning When to Stop Reasoning

Settle extends the accuracy-token-count Pareto frontier of the evaluated stopping methods and predicts whether a correct answer will remain correct in completed traces.

Ryan Brown, Zi-Hao Fu, Chris Russell · 0 citations
#artificial intelligence Preprint Sep 2026

Decoupling Token Roles in Autoregressive Pretraining

Autoregressive pretraining increasingly draws on heterogeneous data, making it important to understand how a model learns from an individual token. The next-token prediction objective naturally identifies a token's contribution with its own loss. However, each token is not only a prediction target but also context for...

Su-Qin Yuan, Runqi Lin, Ke-Yu Lin et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment

It is shown that preference alignment preserves the human response distribution only under a restrictive condition, and no consistent evidence that real human preferences satisfy it, and human-likeness is established as an explicit dimension of alignment rather than something assumed to follow from preference alignment...

Su-Qin Yuan, Runqi Lin, Mu-Yang Li et al. · 0 citations
#machine learning Preprint Jun 2026

Running the Gauntlet: Hard Agentic Tasks

GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), is introduced, revealing the substantial gap between current agent capabilities and those required for complex...

Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.