Skip to content

Communication between Frozen Large Language Models via Prompt Optimization in a Referential Game

Sep 2026 · 0 citations · 35 references
Computer Science

TL;DR

Communication between two frozen large language models from different providers, with different tokenizers, accessed through their API endpoints is studied, finding that successful place value communication in some runs is rare.

Abstract

We study communication between two frozen large language models from different providers, with different tokenizers, accessed through their API endpoints. The two play a referential game: one sees an object and describes it in a short fixed-length message over a small alphabet; the other must pick that object out of a candidate set. Neither model's weights are updated. Each agent's prompt is rewritten by an isolated prompt optimizer whose reflection model reads that agent's scored interactions. In the positional setting, optimized prompts carry a shared code that generalizes to held-out objects above a measured no-codebook baseline, including when the memory window is removed. In a second setting, independent per-letter blocks no longer fit within the message, although a whole-object place value code does. The base system fails to establish reliable communication: the sender struggles to retain an injective rule, and the receiver has too few confirmed examples in view. A sender collision penalty, retention of successful interactions, and sequential optimization enable successful place value communication in some runs. Outcomes vary across runs and reflection models. In successful runs, the protocol is written into the optimized prompts, where it can be read and audited directly.

View source

Similar papers

#natural language process... Preprint Sep 2026

Lost with a Map: Conversational State and Behavioral Reliability in Language Models

Task-oriented dialogue requires maintaining and updating information across turns, yet language models expose no explicit belief-state object. We study how conversational state is represented, updated, and used inside eight instruction-tuned language models from four families on MultiWOZ and SGD. Structure and values s...

Atahan Dokme, Larry Heck · 0 citations
#artificial intelligence Preprint Aug 2026

Exploring Collaboration between a language and a non-language agent

To solve LLM collaboration with non-language agents, latent state internalization is introduced, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state.

Harini S.I., Somesh Singh, Yaman Kumar Singla et al. · 1 citation
#natural language process... Preprint Sep 2026

Where a Model Sends Its Own Repeated Token

Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We take that input and ask where the model goes when it does not. For each token t, read argmax...

Nicolás Vera Zúñiga · 0 citations
#artificial intelligence Preprint Sep 2026

The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models

When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? We propose a round-trip protocol that answers this question empirically for tree-structured expressio...

Xavier Suau, A. F. de las Morenas, L. Zappella et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random

It is found that a literal-similarity baseline with no pragmatics outperforms most tested language models, that adding a pragmatic layer over two baseline similarity sources moves choosers toward random rather than toward the Bayesian reference, and that a standard labelled multiple-choice format carries no measurable...

Cris Huynh · 1 citation

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.