Skip to content
Preprint

GEB-Bench: Abstract Structures Told in Many Voices

Aug 2026 · 0 citations · 33 references
Computer Science

TL;DR

Evaluating twelve open and proprietary models, it is found that abstraction failure is lawful, and the central finding is a gap between recognition and cross-voice mapping: models identify a structure within one voice far better than they carry it across voices; every model pays this tax, and mapping strong enough to narrow it appears only at the frontier tier.

Abstract

Can a model look at a river delta and a lightning bolt and see that they share a structure? We introduce GEB-Bench, a benchmark whose unit is an abstract structural motif--self-reference, a strange loop, a Mobius twist--in the spirit of Godel, Escher, Bach. Each motif is told in several voices: a natural scene whose composition is the structure, a folk story whose telling enacts it through a mechanically checkable form device, a mathematical theorem, and a programmatic skeleton; surface parameters are declared nuisance variables and never scored. Motifs, voices, and the structural changes between them form a small cross-modal category, and GEB-Bench's tasks are its questions. Evaluating twelve open and proprietary models, we find that abstraction failure is lawful. The central finding is a gap between recognition and cross-voice mapping: models identify a structure within one voice far better than they carry it across voices; every model pays this tax, and mapping strong enough to narrow it appears only at the frontier tier. Two patterns support it. Errors align more strongly with the designed formal geometry than with measured perceptual geometries, and frontier models from different vendors converge on the same wrong answers; and surface complexity taxes every model that reads structure, with capacity buying headroom rather than immunity. GEB-Bench is fully generative and released with its pipeline.

View source

Similar papers

Preprint Aug 2026

Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace

Reading a base Qwen3-Omni with a logit lens at the audio-token positions, it is found that the answer to a spoken question becomes legible - in words - in the model's middle layers, before it emits any token.

Jia-Jun Fan, Jing-Yuan Li, Prashanth Gurunath Shivakumar et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EnigmaForge: The Question Is Hidden in the Story

Most benchmarks hand the model a question. EnigmaForge hands it a stack of old documents and no question at all. Buried in the letters, receipts, and logbook margins is a small logic puzzle whose solution is unique - proved by a SAT solver at generation time, with an ablation certificate showing every clue is load-bear...

Daniel C. Eisner · 0 citations
#natural language process... Preprint Sep 2026

ScorePrompts: Natural-Language Exploration of Symbolic Music Scores through Analysis

We present ScorePrompts, an interactive system in which users upload a score, receive natural-language descriptions of its musical structure, ask questions about specific passages, and inspect the corresponding analysis results in staff notation. Specialist MIR components first estimate harmony, tonality, cadences, for...

Emmanouil Karystinaios, Gerhard Widmer · 0 citations
Preprint Aug 2026

Probing Character-level Transformers for the Spanish L-shaped Morphome

When a transformer learns an irregular morphological pattern, what has it learned? Our test case is the Spanish \emph{L-shaped morphome}, a complex irregular pattern in which the verb's stem alternates in exactly the first-person singular indicative and all subjunctive forms, and whose membership no phonological, seman...

Akhilesh Kakolu Ramarao, Kevin Tang, Wiebke Petersen et al. · 0 citations
Open access Oct 2026

2×6

2×6 consists of short “stanzories”—stanzas that are also stories, each one relating an encounter between two people. Appearing in English, French, Spanish, Russian, Japanese, and Polish, the stanzories are generated by a similar underlying process, even as they do not neatly correspond to one another the way a trans...

Nick Montfort, Serge Bouchardon, C. León et al. · 0 citations
#natural language process... Preprint Sep 2026

MUSE: A Theory-Harnessed Story Engine for Vibe Narrativizing

LLMs have been able to generate fluent prose, but high-quality stories also require coordinated decisions about plot, character, and language across planning, drafting, and revision. We formulate Vibe Narrativizing as turning natural-language writing requirements into a finished story. MUSE, a Theory-Harnessed Story En...

Jian-Xiang Ma, Xiaocui Yang, Da-Ling Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.