Skip to content

Dual-Latent Memory Routing for Vision-Language Reasoning

Sep 2026 · 0 citations · 53 references
Computer Science

TL;DR

DLMR is a parameter-efficient mechanism that equips MLLMs with Dual Latent Memories: a visual memory that compresses image evidence and a reasoning memory that tracks intermediate conclusions and constraints with only a small number of additional trainable parameters.

Abstract

Multimodal large language models (MLLMs) have recently made strong progress in vision-language reasoning, yet their performance often degrades as generations grow longer. A key factor is that they frequently lose track of earlier visual evidence and intermediate constraints under a monolithic growing context. Inspired by how humans separately recall what they see and what they infer when solving complex tasks, we propose DLMR, a parameter-efficient mechanism that equips MLLMs with Dual Latent Memories: a visual memory that compresses image evidence and a reasoning memory that tracks intermediate conclusions and constraints. A Router then dynamically decides which memory and how much to reuse during inference, preserving visual grounding while maintaining coherent long-horizon reasoning. DLMR is trained in three stages, from latent memory construction to selective router learning, while keeping the base MLLM frozen, yielding substantial gains on both general and reasoning benchmarks with only a small number of additional trainable parameters. Analyses further show interpretable, state-dependent routing with specialized memory roles and reduced decoding tokens over long generations. Code is available at https://github.com/Hunter-Wrynn/DLMR.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

ConvMem: Convolutional Memory for Long-Context Reasoning

While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective context by reading text in segments and iteratively updating a fixed-size memory. However,...

Hongming Zhang, Zhao-Zhen Gu, Fengshuo Bai et al. · 0 citations
Preprint Aug 2026

TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory

TRAM (TRajectory-derived Auxiliary Memory), a training-free method that augments standard decoding with an auxiliary memory pathway derived from the model's own reasoning trajectory, shows that TRAM improves performance over vanilla decoding on mathematical, scientific, and general visual reasoning tasks without additi...

Kang Liu, Zi-Jing Wang, Yongkang Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Structured Sparse Memory for Recurrent Reasoning

Recurrent models trained from scratch have recently become competitive on ARC-style reasoning tasks, but the usual framing around small recurrent backbones overlooks two important parts of the system: task-conditioned memory and synthetic augmentation data. We study this regime through CHARM, a compact hybrid ARC model...

Zi-Xuan Zhao, Samuel Wheeler, Neil Getty et al. · 0 citations
#machine learning Preprint Sep 2026

Latent-Aligned Reasoning for Multimodal Recommendation

LARK (Latent-Aligned Reasoning frameworK), a two-stage latent reasoning framework with complementary alignment mechanisms within a single VLM, achieves state-of-the-art performance across multiple recommendation architectures, with controlled ablations confirming the distinct contribution of each component.

Jia-Rui Jin, An-Ya Ji · 0 citations
Preprint Aug 2026

ChronoVision: Temporal Reasoning via Latent State Reconstruction

This work proposes ChronoVision, a multimodal framework designed to align visual logic with latent imagery, and introduces Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task.

Yi-Fan Shen, Jian Xu, Boyi Li et al. · 1 citation
#artificial intelligence Preprint Sep 2026

ARM: Attention with Routed-Memory for Learnable Sparse Control

Despite advances in long-context inference, large language models (LLMs) remain fundamentally limited by the key-value (KV) caching mechanisms that are necessary for stable computation. Techniques such as selective token eviction and pruning have vastly mitigated these issues, but often discard core information to mana...

Qiu-Hao Zeng, Jerry M. Huang, Peng Lu et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.