Evaluating multimodal foundation models for engineering drawing assessment and spatial reasoning
This work investigates two different aspects of contemporary multimodal vision language models (VLMs) regarding their analysis and understanding of engineering technical drawings: (i) their capacity to grade drawings based on their text, image and graphic analysis, and (ii) their engineering spatial reasoning capabilit...