Evaluating multimodal foundation models for engineering drawing assessment and spatial reasoning
Abstract
This work investigates two different aspects of contemporary multimodal vision language models (VLMs) regarding their analysis and understanding of engineering technical drawings: (i) their capacity to grade drawings based on their text, image and graphic analysis, and (ii) their engineering spatial reasoning capabilities. Using a collection of one hundred authentic engineering drawing assessments and a novel drawing-to-isometric correspondence task, we examined how two widespread commercially available VLM models, namely Gemini (Google DeepMind) and Copilot (Microsoft), perform across multiple repeats of the same task. Both experiments were benchmarked to human experts who fulfilled the same task as a baseline comparison. Results from the drawing assessment and marking experiment showed a notable capacity of both VLMs to grade submissions, both within 10% error relative to human marking results. The intra-marking consistency and the ranking ability varied across models, with Copilot displaying higher reliability levels closer to human level. The isometric-matching task, however, demonstrated that current VLMs struggle with highly visual abstraction tasks, as their model-matching capacity was well below human marker’s performance. We conclude that, although VLMs exhibit high performance in analytical tasks involving optical recognition, text and data mining, they appear to still lack sufficient abstraction capacity consistent with human mental processes such as mental rotation, graphic memory retention and shape matching. A cross-analysis further revealed that spatial reasoning performance and grading accuracy are moderately associated for Gemini, but independent for Copilot, suggesting fundamentally different underlying processing strategies between the two models.