Skip to content

PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models

Aug 2026 · 0 citations · 24 references
Computer Science

TL;DR

PolyComp, a procedurally generated and verified benchmark that stresses visual recognition and compositional spatial reasoning, is introduced, and the observed accuracy spread across geometry families is larger than across presentation formats.

Abstract

We introduce PolyComp, a procedurally generated and verified benchmark that stresses visual recognition and compositional spatial reasoning. In each problem, a model must identify which of four options shows a pair of polycube components that can be combined to form a target solid. The benchmark contains 120 problems across four geometry families, and each problem has three different presentation formats using either a single image or multiple images. The random guessing baseline is 25%. Across the three presentations (360 presented problems per model), GPT-5.6 Sol with max effort attains 50.0% accuracy (95% problem-cluster CI 43.3-56.7%) at a mean cost of \$0.951 per presented problem, Claude Fable 5 with max effort attains 39.4% (33.1-46.1%) at \$0.701, and Gemini 3.1 Pro Preview with thinking level high attains 27.5% (22.8-32.5%), near the 25% random guessing baseline, at \$0.350. The observed accuracy spread across geometry families is larger than across presentation formats. We present a problem development and evaluation protocol, cost and token accounting, and release the 120 problems.

View source

Similar papers

Preprint Jul 2026

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit{\ours{}}, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth. Across 95K instances at four grid densities (4$\times$4 to 16$\times$16), we find that \textbf{zero-shot VLMs largely lack geometric reasoning}: only one of five frontier models (GPT-5.5) exceeds random baseline on 4$\times$4 puzzles, while all others perform at chance level. While supervised fine-tuning achieves $>$97\% on 4$\times$4, \textbf{all models collapse on larger grids}: GPT-5.5 drops from 70\% to near-random on 8$\times$8, and even fine-tuned models fall below 5\% on 12$\times$12. This ``scaling cliff''suggests current architectures cannot maintain consistent constraint satisfaction as the number of pieces increases. \ours{} establishes scalable geometric reasoning as an open challenge for vision-language models.

Shawn Li, Wei Yang, Jike Zhong et al. · 0 citations
Preprint Aug 2026

Disentangling 3D Modeling from Spatial Reasoning

This work proposes the Disentangled Spatial Reasoner (DiSR), a simple yet effective framework that reconstructs the physical world into structured 3D evidence using off-the-shelf expert perception models and fine-tunes an LLM with LoRA to perform reasoning solely over this explicit geometric evidence.

Haoze Sun, Jiequan Cui, Qingshan Xu et al. · 0 citations
#natural language process... Preprint Aug 2026

Reactivating Test-Time Scaling for Plane Geometry Problem Solving

Plane geometry problem (PGP) solving has become a critical benchmark for multimodal reasoning because it requires accurate visual perception and precise multi-step symbolic deduction. Although test-time scaling (TTS) has demonstrated remarkable success in general mathematical reasoning, it fails to scale effectively under the symbolic-program paradigm for plane geometry. We identify two key obstacles: limited reasoning diversity induced by rigid symbolic programs and insufficient explicit visual grounding before symbolic deduction. To address these issues, we propose Multi-Trace Synthesis (MTS), which converts each symbolic program into heterogeneous reasoning traces, including executable Python scripts and CoT-augmented variants. We further propose Perception-Augmented (PA) training, which parses diagrams into structured semantic clauses before deduction, and Consensus-Guided Multi-Trace Ensemble (CG-MTE) for efficient self-adaptive inference. Experiments on three geometry benchmarks show that our method consistently improves PGP-solving across model scales and achieves strong performance against both general-purpose MLLMs and specialized geometry solvers. Under test-time scaling, CG-MTE achieves comparable accuracy to high-budget self-consistency while reducing sampling cost by up to 8x. Code and data are publicly available at https://github.com/Jason8Kang/ReTTS-PGPS.

Xiao-Qiang Kang, Shen-Gen Wu, Maizhen Ning et al. · 0 citations
Preprint Aug 2026

ReRef-3D: A Benchmark for Spatial Referring Expression-Guided 3D Scene Rearrangement

We introduce ReRef-3D, a benchmark for language-guided placement in 3D scenes. It contains 33,826 instructions across 998 CLEVR-derived scenes, spanning 16 placement families and direct, one-hop, and two-hop references. Each instruction must be resolved into a valid new placement position. Given that an instruction defines a region of acceptable placements rather than one coordinate, our evaluation inserts a prediction into the scene, recomputes relations, and tests relation satisfaction and physical validity. Each instruction also includes a verified naturalized rewrite. After fine-tuning, LLaVA-3D, 3D-LLM, and PlaceIt3D produce valid placements for 68.3%, 31.6%, and 22.4% of instructions, respectively. Across models, relation satisfaction surpasses physical validity, relations such as nearest and between are the most difficult, and phrasing has minimal effect on performance.

Mary Lynn Martin, Yifei Zhang, M. Palmer et al. · 0 citations
Preprint Aug 2026

OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D objects and rarely address the fine-grained geometry and millimeter-level tolerances required in industrial mechanical design. We introduce OmniMech, the first million-scale benchmark for evaluating VLMs on executable CAD generation from industrial manufacturing data. OmniMech contains more than 251,000 fully dimensioned and toleranced 2D orthographic drawings, paired with native CAD models, multi-view renderings, mesh, STEP and B-rep representations, and rich semantic annotations. The benchmark includes four tasks: (1) parametric CAD program synthesis from engineering drawings; (2) diagram-to-3D reasoning for geometrically and structurally consistent reconstruction; (3) annotation-grounded reasoning over dimensions, symbols, feature callouts, and manufacturing constraints; and (4) tool-augmented agentic reasoning using visualization, measurement, CAD execution, and verification tools. Experiments show that current VLMs and CAD-specialized models still struggle with executable program synthesis, fine-grained 3D reconstruction, and reliable enforcement of dimensions and tolerances. We will release the benchmark data, evaluation code, and tool interfaces to support future research.

Taiting Lu, Runze Liu, Ziwei Dong et al. · 0 citations
Open access 2026

Stimulus-Design Confounds in Progressive Rendered 3-D VLM Evaluation

Vision-language model (VLM) evaluation on rendered 3D stimuli is a computer graphics stimulus-design problem: camera, visible faces, silhouettes, shading, and answer format together determine what evidence the model receives. We introduce a progressive surface-disclosure protocol that reveals mesh faces at matched surface-area budgets, renders them as texture-free views, and queries the VLM under category-choice, object-choice, or free-response tasks. We evaluate it on Core-100, a controlled set of 100 everyday 3D mesh models, with five local VLMs. Disclosure ordering produces large threshold differences: in the multi-view object-choice setting, random face disclosure reaches a censored mean threshold of 28.5%, while principal-axis sweep and spatially contiguous growth require 45.7% and 46.4%. The direction repeats across all five models and survives a projected-coverage adjustment and a random-growth ablation that removes the large-face seed prior. Because the forced-choice first-hit metric accumulates chance, we pair the censored thresholds, which size the effect, with a label-shuffle chance-corrected comparison, which confirms it is above chance. At the 10% budget random disclosure beats its shuffle baseline by 28.7 points while connected barely does, so only spatially distributed orderings clear chance at low budgets. Task format matters too: object-choice succeeds on 85.1% of sequences, while strict free-response naming succeeds on 56.8% (68.2% with a fixed alias table). Silhouette-only and coverage-matched experiments rule out interior shading and coverage magnitude, narrowing the driver to the distributed projected shape; a contrastive CLIP baseline shows a much weaker same-direction effect, suggesting the sensitivity is amplified in generative VLMs. The results position progressive rendered 3D VLM evaluation as a measurement protocol whose reports should include disclosure distribution and connectivity, view and task policy, scoring rules, and censoring choices.

Heeyoung Park, Jinyoung Choi, Youngjin Park · 0 citations

Related blog posts