Skip to content

Author

Grant Wilson

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

A Controlled Replication of a Self-Model Loop for Language-Model Agents: Controls, Confounds, and Finite-Horizon Path Dependence

An independent, clean-room reimplementation of a self-model loop architecture for language-model agents (the "AC1 loop" described in Lark Laflamme's 2026 AC1-LLM / Laflamme-3T essays). The architecture wraps an LLM with a Bayesian belief over the agent's own interaction stance, re-injected into generation each turn alongside a hidden deliberative monologue. The originating write-ups report strong effects but omit the controls needed to separate the contribution of the state's content from the mere presence of an instruction, an honest task baseline, or the intrinsic inertia of the estimator. This work supplies those controls across six experiments: a structured-placebo ablation (E1), a clamped partial factorial over state x gate x monologue x strategy (E2), hidden-stance inference against an honest single-shot LLM baseline (E3), and closed-loop dynamics driven against an OpenAI-compatible chat endpoint with decay-free and state-decoupled null controls (E4–E6). In small synthetic experiments on one or a few model endpoints, mode-labelled prompt additions reliably changed output length and question use, most strongly when the posterior was confident and with the private monologue as the leading but not isolated factor. A classifier-plus-accumulator inferred synthetic modes well above chance without an accuracy advantage over an honest single-shot baseline. Under a cyclic numerical drive the closed loop produced substantial finite-horizon path dependence, most of it accumulator arithmetic, with the size and mechanism of any additional feedback contribution left uncertain; a 100-turn basin test found no evidence of bistability, with relaxation still in progress. Every tested claim reproduces in a "true-but-softer" form once the missing controls are added. Scope is strictly the loop's measurable behaviour; no claim is made about consciousness or any Psi-threshold. Code and raw result data are released (MIT). Version 1.1 revises v1.0 after an external validity review: corrections of fact (E2 is 30 configurations / 180 replies and a partial factorial; the "numbers-only" control is renamed strategy-stripped; the gate channel is substantial, not minor; "loop area" is a mean vertical separation), an E1 reproducibility caveat, and claims softened to what the designs support. See the paper's Revision history. No data were re-run.

Grant Wilson · 0 citations
#small language model Open access Sep 2026

A Controlled Replication of a Self-Model Loop for Language-Model Agents: Controls, Confounds, and Finite-Horizon Path Dependence

An independent, clean-room reimplementation of a self-model loop architecture for language-model agents (the "AC1 loop" described in Lark Laflamme's 2026 AC1-LLM / Laflamme-3T essays). The architecture wraps an LLM with a Bayesian belief over the agent's own interaction stance, re-injected into generation each turn alongside a hidden deliberative monologue. The originating write-ups report strong effects but omit the controls needed to separate the contribution of the state's content from the mere presence of an instruction, an honest task baseline, or the intrinsic inertia of the estimator. This work supplies those controls across six experiments: a structured-placebo ablation (E1), a clamped partial factorial over state x gate x monologue x strategy (E2), hidden-stance inference against an honest single-shot LLM baseline (E3), and closed-loop dynamics driven against an OpenAI-compatible chat endpoint with decay-free and state-decoupled null controls (E4–E6). In small synthetic experiments on one or a few model endpoints, mode-labelled prompt additions reliably changed output length and question use, most strongly when the posterior was confident and with the private monologue as the leading but not isolated factor. A classifier-plus-accumulator inferred synthetic modes well above chance without an accuracy advantage over an honest single-shot baseline. Under a cyclic numerical drive the closed loop produced substantial finite-horizon path dependence, most of it accumulator arithmetic, with the size and mechanism of any additional feedback contribution left uncertain; a 100-turn basin test found no evidence of bistability, with relaxation still in progress. Every tested claim reproduces in a "true-but-softer" form once the missing controls are added. Scope is strictly the loop's measurable behaviour; no claim is made about consciousness or any Psi-threshold. Code and raw result data are released (MIT). Version 1.1 revises v1.0 after an external validity review: corrections of fact (E2 is 30 configurations / 180 replies and a partial factorial; the "numbers-only" control is renamed strategy-stripped; the gate channel is substantial, not minor; "loop area" is a mean vertical separation), an E1 reproducibility caveat, and claims softened to what the designs support. See the paper's Revision history. No data were re-run.

Grant Wilson · 0 citations