Skip to content
Preprint

Towards Physics-Faithful Generation of Scientific Diagrams

Aug 2026 · 0 citations
Computer Science

TL;DR

Princigram shows that explicit physics-structured supervision improves the physical faithfulness of generated scientific diagrams, and curate and structurally annotate 4.3 million physics images, carry expert-level annotation, and adapt a unified multimodal backbone.

Abstract

Text-to-image generation has reached photorealistic quality, yet state-of-the-art systems remain unreliable at producing scientific diagrams, whose value depends not on appearance but on physical faithfulness: correct force directions, valid coordinate systems, consistent thermodynamic states, and equations matching the depicted scenario. Trained on web imagery with physically shallow captions, generic models produce diagrams that look plausible but are physically wrong, harmful in education and scientific communication. We present Princigram, a physics-faithful scientific-diagram generator, and its data pipeline. Our central advance is Structured Physical Chain-of-Thought (SP-CoT): a per-subdiscipline schema that decomposes a physics diagram into an explicit multi-step reasoning chain across six subdisciplines, from scene identification through force or process analysis to governing laws and synthesis. Unlike free-form chain-of-thought, SP-CoT follows a fixed schema with strict fidelity rules that separate visually grounded facts from physically inferred reasoning and type all mathematics symbolically; it serves both as dense training supervision and, at inference, as a structured"thinking"prompt. With it we curate and structurally annotate 4.3 million physics images, of which 115,037 carry expert-level annotation, and adapt a unified multimodal backbone. We further introduce VeriphyT2IBench, whose questions are derived from each held-out diagram's own structured annotation: each diagram becomes an item-specific bank of binary questions about its objects, forces, and states, so a judge model's score decomposes into named physical facts rather than one holistic number. On the physics subset of GenExam and on VeriphyT2IBench, Princigram shows that explicit physics-structured supervision improves the physical faithfulness of generated scientific diagrams.

View source

Similar papers

Preprint Sep 2026

Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation

Recent advancements in unified generative models (UGMs) and world simulators have achieved unprecedented results in visual perception and synthesis. However, these models primarily rely on surface-level event alignment, leaving the capacity for high-level visual reasoning underexplored. True visual generative intellige...

Yu-Tong Liu, Nan Huang, Xu Cao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PhysMAS: Physics-Grounded Multi-Agent Synthesis of Compositional 4D Gaussians

Efficient, fully automatic, and physically plausible 4D Gaussian synthesis is an important goal for dynamic scene generation. Recent physics-based methods couple 3D Gaussians with the Material Point Method (MPM) to generate physically driven motion, but extending this paradigm to heterogeneous multi-part objects and in...

Jiang Qin, Chun-Ji Lv, Yang-Guang Wei et al. · 0 citations
Preprint Aug 2026

NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation

The advancement of Embodied AI necessitates high-quality simulation assets that faithfully mirror the real world. However, transforming raw visual observations into simulation-ready scenes remains challenging due to the lack of physical grounding and scene-level interactivity in current image-to-URDF methods. We propos...

Yumeng He, Yi-Chen Song, Xiao-Tian Yang et al. · 1 citation
Preprint Sep 2026

Principia: Relational Physics Tests for Video Models

Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law,...

Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan et al. · 0 citations
Preprint Sep 2026

Physics as the label for measuring and correcting materials reasoning in multimodal models

Vision-language and language models increasingly interpret materials data, yet benchmarks report that they hallucinate invalid properties and violate physical law. Evaluation matches final answers to scarce human labels, while discovery agents verify final proposals or density functional theory (DFT) execution. Neither...

Hasan Kurban, Rasul Khanbayov, Mustafa Kurban · 0 citations
Preprint Sep 2026

SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation

In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have enabled scientific image generation, yet verifying the correctn...

Jia-Li Chen, Zheng-Teng Lin, Zu-Qi Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.