Skip to content
Preprint

Localize, Then Reason: Visual Latent Structural Reasoning for Molecular Properties and Edits

Aug 2026 · 0 citations · 30 references
Computer Science Biology

TL;DR

This work proposes Visual Latent Structural Reasoning (VLSR), an end-to-end framework that jointly learns localization and reasoning from molecular images, and central to the approach is a localize-then-reason strategy.

Abstract

Local chemical perception and property reasoning are both essential for understanding how molecular structure determines properties. Current LLM-based chemical reasoning methods either receive SMILES/molecular images together with descriptions of local motifs, or reason directly from molecular images. Neither approach enables the model to focus on chemically meaningful regions before reasoning. To address this gap, we propose Visual Latent Structural Reasoning (VLSR), an end-to-end framework that jointly learns localization and reasoning from molecular images. Central to our approach is a localize-then-reason strategy. VLSR first learns to locate chemically meaningful regions in a molecular image. It then reasons about their property effects in a compact latent workspace before producing the final answer. Under the same inference setup, this design achieves 9.6X higher throughput than a comparable textual-reasoning baseline.

View source

Similar papers

Open access Jul 2026

RelAgent: a multi-agent solution for molecular relationship grounding

RelAgent decomposes the task into three interpretable stages: entity extraction, substructure localization, and ontology-guided relationship reasoning, and then uses verifier agents to rank structurally plausible candidates to support fine-grained reasoning over molecular substructure.

Rubing Chen, Jiaxin Wu, C. Zhang et al. · 0 citations
Book Open access Aug 2026

Don't Just Encode But See: A Data-Centric Paradigm for Visual Molecular Understanding in Large Language Models

MolGlass is proposed, a data-centric paradigm for visual molecular understanding in vision-language models (VLMs) that injects chemical priors directly into the visual input through chemical-aware visual augmentations, without modifying model architectures or training molecule-specific encoders.

Runqing Xu, Xiaotang Wang, Chunfeng Gao et al. · 0 citations
Preprint Jul 2026

SAGE-Net: Semantics-Augmented Geometric Encoder for Material Property Prediction

Reliable structure-property modeling is crucial for accelerating materials discovery, where crystal graphs and structure-derived crystallographic descriptions provide complementary geometric and semantic information. Existing multimodal materials models primarily incorporate textual information through post-encoding fusion, latent-space alignment, or attention-based representation interaction mechanisms. However, in most cases, crystallographic semantics are introduced after structural encoding and therefore cannot directly guide the formation of atom-level crystal-graph representations. Here, we present Semantics-Augmented Geometric Encoder Network (SAGE-Net), a flexible multimodal framework that injects description-derived chemical and crystallographic semantics into geometric message passing. SAGE-Net introduces Semantic-Guided Message Passing (SGMP), which gates atom-level updates and enables crystallographic semantics to directly modulate local geometric interactions across multiple graph neural network (GNN) backbones. Across benchmarks covering bandgap, mechanical, transport-related properties, and synthesizability assessment, the SAGE-Net instantiated with different GNN backbones achieves the lowest MAE on eight out of ten JARVIS-DFT regression targets and delivers strong or highly competitive performance against both structure-based and multimodal baselines. For synthesizability assessment, the SAGE-Net demonstrate outstanding classification performance and high recall rates. Interpretability analysis unravels that SAGE-Net effectively captures physically interpretable crystallographic features, viz. space group, dimensionality, polyhedral environments, among others. Together, these results demonstrate SGMP-based SAGE-Net as a general and transferable framework for deeply integrated multimodal materials learning.

Guanghui Zhang, Yuxuan Yao, Kieran B. Spooner et al. · 0 citations
Preprint Jul 2026

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics, is proposed and revealed, revealing complementary strengths of pixel-space expression and text-based reasoning.

Xu Wang, Kaixiang Yao, Miao Pan et al. · 1 citation
Aug 2026

PINS: A Physics-Informed Generative Framework for De Novo Structure Elucidation from 1D NMR Spectra.

PINS (Physics-Informed NMR Structure elucidation model), a generative framework that explicitly bridges the gap between spectral data and molecular topology by enforcing multiphysical priors, provides a trustworthy, automated strategy for decoding novel chemical structures in data-scarce regimes.

Pengfei Liu, Cuimei Liu, Laiqun Xia et al. · 0 citations