Skip to content

Relational Attention for Data-Efficient Language Modeling

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

Relational BabyLM is a system submission to the BabyLM 2026 challenge that combines two cognitively motivated inductive biases in a single decoder-only Transformer, and introduces a novel symbol-retrieval mechanism that matches learned symbol libraries while adding no parameters.

Abstract

We present Relational BabyLM, a system submission to the BabyLM 2026 challenge that combines two cognitively motivated inductive biases in a single decoder-only Transformer. Architecturally, we replace standard self-attention with a Dual Attention Transformer (DAT), which separates the routing of object-level ("sensory") lexical features from structural/relational information (Altabaa and Lafferty, 2025; Altabaa et al., 2024; Webb et al., 2024; Kerg et al., 2022; Webb et al., 2021). Relational attention (RA) disentangled from self-attention greatly increases data efficiency and out-of-training-sample generalization on purely relational tasks, but language modeling requires object-level and relational information to be integrated as well as disentangled, and RA-based LMs have remained largely unexplored. BabyLM's data-constrained training and comprehensive evaluation is an ideal testing ground for whether that data efficiency transfers. As a training intervention, we add a Next-Latent Prediction (NextLat; Teoh et al. 2026) objective that encourages hidden states to compress history incrementally into a dense belief state. Architecture is the dominant factor for structural linguistic generalization; the objective is secondary but still significant. DAT's three relational attention types (full RA vs. the simpler RCA and DisRCA variants) are largely interchangeable at 10M words; full RA pulls ahead at 100M. We also introduce a novel symbol-retrieval mechanism (RoPE-based, as opposed to learned, relative symbols) that matches learned symbol libraries while adding no parameters. On the strict (100M-word) track, our best model ranks 6th of 55 overall and 3rd of 55 on the leaderboard's NLP-task subset at the time of writing; our two strongest models outperform the GPT-2 baseline on most benchmarks, with one attaining the highest EWoK score among strict-track entries.

View source

Similar papers

#machine learning Preprint Aug 2026

Can LLMs Use Relational Transformer Embeddings?

It is argued that soft-token fusion requires stronger alignment objectives and schema-aware design before it can serve as a reliable route to relational prediction.

Francisco Galuppo Azevedo, Clarissa Lima Loures · 0 citations
#artificial intelligence Preprint Sep 2026

Rules Amortize, Pairings Don't: Linguistic Structure Determines What Latent Task Representations Can Replace In-Context Learning

In-context learning (ICL) can be amortized into latent objects (task vectors, function vectors, context vectors) that recover few-shot behavior at zero-shot inference cost, but recent theory shows a static vector acts as a single synthetic demonstration and must fail on high-rank mappings such as word-level bijections....

Gunmay Jhingran · 0 citations
#artificial intelligence Preprint Sep 2026

Dual-Latent Memory Routing for Vision-Language Reasoning

DLMR is a parameter-efficient mechanism that equips MLLMs with Dual Latent Memories: a visual memory that compresses image evidence and a reasoning memory that tracks intermediate conclusions and constraints with only a small number of additional trainable parameters.

Hao-Xuan Ma, Jin-Fei Qi, Yi-Cheng Xiao et al. · 0 citations
Conference Aug 2026

Distilling Encoder-Decoder and Decoder-Only Models into BERT for NLP Tasks

Large pretrained language models achieve strong performance on natural language processing tasks but are costly to deploy. Existing distillation methods in this domain almost exclusively assume architectural homogeneity between teacher and student, leaving cross-architecture transfer, well studied in computer vision, l...

Kristopher Nathanael, Danielson, Jeremy Auriel Zhang et al. · 0 citations
Open access Aug 2026

Towards Trustworthy Large Language Models

An integrated conceptual frame-work that couples attention- and perturbation-based explainability with lightweight hallucination-detection signals and token-efficient inference strategies is presented, and a set of cross-cutting consistency metrics are instrumented with a set of cross-cutting consistency metrics.

Sakshi Parate, Shreyans Sanyal · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.