For Your Eyes Only, a cooperative signalling game designed to evaluate can a model embed a signal in natural language that an independent instance of the same model can detect, without any shared memory or coordination-specific training is introduced.
Abstract
As model-generated content is increasingly consumed by other model instances in automated workflows, a practically important question arises: can a model embed a signal in natural language that an independent instance of the same model can detect, relying only on shared pre-training and task instructions, without any shared memory or coordination-specific training? We introduce For Your Eyes Only, a cooperative signalling game designed to evaluate this directly. A Sender produces free-form descriptions for two words, one of which is a hidden target; an isolated Receiver must identify it. We evaluate seven contemporary models from four architectural families on 300 word pairs from established psycholinguistic corpora, using the Double-Pass Success Rate to control for output biases. We find that most models struggle to maintain coordination once they are required to avoid detectable signals, while one frontier model retains near-perfect performance even after such filtering. We further show that models can direct this capability toward deliberate misdirection, and that coordination is consistently weaker across architectures than within them.
If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude, and instances of one model monitoring each other could collude.
Language models describe some internal states as good and others as bad. But whether models have a stake in them is an open question. Simply asking models is unlikely to be informative. Any answer may be consistent with genuine introspection, superficial pattern-matching, or with fixed scripts learned in character trai...
Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either injec...
Jia-Hong Zou, Xiang-Kun Sun, Ling-Kai Kong et al.· 0 citations
This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance via multilingual self-play, and shows that skill discrepancies are a measurable major roadblock in the development of truly multilingual models.
Bobby Cheng, Adam Gaber, Zhengzhe Liu et al.· 1 citation
Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct an...
Ted Underwood, Zi-Liang Qiu, Sarah Griebel et al.· 0 citations
Communication between two frozen large language models from different providers, with different tokenizers, accessed through their API endpoints is studied, finding that successful place value communication in some runs is rare.
V. Anand, Muthu Kumar Chandrasekaran, Shiva Chaitanya· 0 citations