Large language models (LLMs) are capable of completing a variety of tasks, but remain unpredictable and intractable. Representation Control (RepControl) seeks to resolve this problem through targeted interventions that modify high-level representations of concepts such as honesty, harmfulness or power-seeking. We forma...
GPS-Bench is introduced, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence.
Linh Le, Melanie Bui, My Chiffon Nguyen et al.· 0 citations
A causal model of how sandbagging is carried in the residual stream is proposed, which predicts a window of layers, after the last sandbagging write and before the answer commit, in which a single-layer reference graft of the sandbagging axis to its honest value restores the full capability.
Hong-Fu Tan, Linh Le, David Williams-King· 0 citations
Sandbagging, in which a model deliberately underperforms on an evaluation despite retaining the underlying capability, threatens the safety evaluations that frontier-model governance depends on. The Elicitation Game found that fine-tuning elicits hidden capability from sandbagging model organisms whereas additive activ...
Linh Le, Hong-Fu Tan, David Williams-King· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.