Skip to content
Open access

Stimulus-Design Confounds in Progressive Rendered 3-D VLM Evaluation

2026 · IEEE Access · Vol 14, pp. 122330-122345 · 0 citations · 35 references

Abstract

Vision-language model (VLM) evaluation on rendered 3D stimuli is a computer graphics stimulus-design problem: camera, visible faces, silhouettes, shading, and answer format together determine what evidence the model receives. We introduce a progressive surface-disclosure protocol that reveals mesh faces at matched surface-area budgets, renders them as texture-free views, and queries the VLM under category-choice, object-choice, or free-response tasks. We evaluate it on Core-100, a controlled set of 100 everyday 3D mesh models, with five local VLMs. Disclosure ordering produces large threshold differences: in the multi-view object-choice setting, random face disclosure reaches a censored mean threshold of 28.5%, while principal-axis sweep and spatially contiguous growth require 45.7% and 46.4%. The direction repeats across all five models and survives a projected-coverage adjustment and a random-growth ablation that removes the large-face seed prior. Because the forced-choice first-hit metric accumulates chance, we pair the censored thresholds, which size the effect, with a label-shuffle chance-corrected comparison, which confirms it is above chance. At the 10% budget random disclosure beats its shuffle baseline by 28.7 points while connected barely does, so only spatially distributed orderings clear chance at low budgets. Task format matters too: object-choice succeeds on 85.1% of sequences, while strict free-response naming succeeds on 56.8% (68.2% with a fixed alias table). Silhouette-only and coverage-matched experiments rule out interior shading and coverage magnitude, narrowing the driver to the distributed projected shape; a contrastive CLIP baseline shows a much weaker same-direction effect, suggesting the sensitivity is amplified in generative VLMs. The results position progressive rendered 3D VLM evaluation as a measurement protocol whose reports should include disclosure distribution and connectivity, view and task policy, scoring rules, and censoring choices.

Read PDF