Skip to content

PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides

Aug 2026 · 0 citations · 91 references
Computer Science

TL;DR

PPTBench establishes a measurable testbed for studying visual coding and advancing agents toward more reliable visual creation, and shows that agents can generally produce valid slide files, but struggle to produce high-quality reconstructions that faithfully recover the semantics and visual structure of the target.

Abstract

Coding agents are increasingly moving beyond text-based software tasks to reconstruct visual targets through code. This capability, visual coding, requires agents to translate their understanding of visual targets into executable code. Slides provide a natural testbed for this capability, combining rich visual structure with objects that can be programmatically created and edited. To measure this capability, we introduce PPTBench, a benchmark for reconstructing scientific flow diagrams as editable PowerPoint slides. PPTBench covers 500 scientific flow-diagram tasks across 10 presentation domains, drawn from real research papers, and is evaluated with a four-stage agentic judge covering artifact validity, process and connector fidelity, rendering quality, and visual fidelity. Evaluation of ten models across 36 model--harness--effort configurations shows a substantial gap in reliable visual coding: the best configuration achieves 77.34, while the median across configurations is 24.38. Fine-grained analysis shows that agents can generally produce valid slide files, but struggle to produce high-quality reconstructions that faithfully recover the semantics and visual structure of the target. Further analysis shows that increasing reasoning effort primarily improves hard-gate passage rather than mean detail quality on each configuration's passing tasks, while configurations with more inspection tend to achieve higher overall scores. PPTBench establishes a measurable testbed for studying visual coding and advancing agents toward more reliable visual creation.

View source

Similar papers

Preprint Aug 2026

Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations

Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understanding. We study whether this representation can serve as operational context for agentic coding, where an agent must navigate repositories, edit source files, and verify executable...

Weijie Liang, Yuanfeng Song, Xing Chen et al. · 0 citations
#computer vision Review Sep 2026

ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts

Multimodal coding agents are expected to turn visual inputs into usable artifacts, and they act through a harness, the layer of tools, context management, and execution environment around the model. Existing evaluations often isolate short tool calls, API traces, or screenshot resemblance, and a low score under these p...

Li-Yang Fan, Chi Wei, Yi-Tai Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes

Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing render can hide incorrect spatial relations, intersecting obj...

Xiao-Kang Ye, Siddhant Hitesh Mantri, Zi-Meng Chen et al. · 0 citations
Preprint Aug 2026

VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise existing visualizations to repair flawed charts or adapt them to desired styles. Existing benc...

Mizanur Rahman, Arshia Azimlu, Shadikur Rahman et al. · 2 citations
#machine learning Preprint Aug 2026

ViCo: Visual-oriented Coding with Self-Reflection for Chart Replication

This paper addresses the challenge of generating high-quality academic charts that match the visual standards of human-authored papers. While existing AI agents can produce well-structured text and code, their generated visualizations often lack the stylistic and semantic fidelity of human designs. Advanced coding agen...

Jiaxin Duan, Dian-Jiao-Shuai Zhao, Jia-Bing Leng et al. · 0 citations
#computer vision Preprint Sep 2026

Editable Visual Design

While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control...

Jun-Yan Ye, Wei Liu, Dongzhi Jiang et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.