Automating industrial CAD design and manufacturing places distinctive demands on multimodal foundation models: the model must see engineering drawings and 3D geometry screenshots, write correct parametric-modelling scripts and Windows COM API code, and cover the full range from single parts to assemblies. General-purpose multimodal models fall short on these tasks, while single-task fine-tuning is too narrow to support the diverse calls that upper-layer agents issue. We build IndustryForge-27B on top of Qwen3.5-VL-27B by curating and integrating six industrial-CAD sub-corpora totalling $\sim$52k multimodal samples---CAD Visual QA (CAD-VQA), parametric CAD code (text2cadquery), assembly-level CAD code (text2cadquery-assembly), and three COM sub-corpora for Inventor / SolidWorks (com_2d / com_3d / com_assembly)---and training with a unified multi-task SFT recipe. Across four CAD-domain benchmarks IndustryForge-27B lifts the base model by $+33.65$~pp on average and outperforms the strong closed-source model GPT-5.4 on all four; across eleven general-capability benchmarks it retains, and slightly improves upon, the base model ($+1.56$~pp mean, no catastrophic forgetting). IndustryForge-27B will serve as the common substrate for downstream industrial-agent projects, providing a unified starting point for a full-stack industrial agent that spans from CAD design to industrial-software operation, from parts to assemblies, and from single-shot generation to closed-loop self-improvement.
Nianchen Deng, Jiaxin Ai, Tao Hu et al.· 0 citations
AssemCAD is presented, an axiom-grounded framework for production-ready CAD assembly generation from natural language that extends Text-to-CAD from isolated part generation toward production-ready mechanical assembly design.
MUSE is a framework that enables iterative self-improvement through a hierarchical Memory Module that organizes cross-domain insights to facilitate the orchestration of long-horizon workflows and demonstrates that MUSE’s performance scales with the accumulation of insights and exhibits strong cross-task transferability.
Cheng Yang, Xuemeng Yang, Licheng Wen et al.· Annual Meeting of the Associ...· 2 citations· ⚡1