Skip to content
Preprint

Learning visual representations for compositional analysis of artworks and photographs

Aug 2026 · 0 citations · 51 references
Computer Science

TL;DR

This work compares two parallel paradigms for composition analysis: a human-inspired method grounded in perceptual grouping, and fine-tuned foundation models enabled by recent large-scale compositional datasets.

Abstract

Composition, the deliberate arrangement of visual elements, is central to how meaning, emotion, and aesthetic quality are conveyed in artwork, yet it remains among the least formalized dimensions of visual understanding. Prior work highlights a persistent gap in learning meaningful compositional representations, attributing it to semantic bias and suggesting that human-inspired approaches may be key. We compare two parallel paradigms for composition analysis: a human-inspired method grounded in perceptual grouping, and fine-tuned foundation models enabled by recent large-scale compositional datasets. The human-inspired approach uses object-centric models for region-level decomposition and a graph attention network to capture spatial relationships between elements. Both paradigms are evaluated on composition score/category prediction, compositional image retrieval, and visual saliency detection. With frozen encoders, the human-inspired method achieves competitive performance while remaining interpretable. When sufficient data enables fine-tuning, large self-supervised models outperform significantly, but at the cost of interpretability and cross-domain generalization. Code and pre-trained models are available on GitHub.

View source

Similar papers

Preprint Aug 2026

Abstract4D: A Large-Scale Dataset and Framework for Understanding the Visual Language of Abstract Art

This work analyzes the semantic structure of abstract art through large-scale embedding visualization, uncovering how perceptual relationships organize artistic meaning, and establishes benchmark tasks for classification, cross-modal retrieval, and text-to-image generation to evaluate how AI models perceive and reprodu...

Hao-Wei Zhang, Yuanpei Zhao, Ji-Zhe Zhou et al. · 0 citations
Preprint Aug 2026

MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding

MMArt is introduced, a large-scale dataset of 74,234 WikiArt paintings, each annotated with four independently annotated perspectives plus a harmonized unified caption, produced by specialized vision-language models or human annotation and validated through complementary quality evaluations.

Shuai Wang, Wang-Yuan Ding, Yixian Shen et al. · 0 citations
Review Open access 2026

A Review of Fine-Grained Visual Categorization with Deep Learning

: Fine-grained visual categorization (FGVC) presents a class of recognition problems in which the discriminative signal is spatially concentrated, visually subtle, and easily destroyed by the preprocessing and augmentation strategies that serve coarse recognition well. Where standard image classification requires a mod...

Richard Adusei, G. Abdul-Salaam · 0 citations
#artificial intelligence Preprint Aug 2026

Conducting Stylistic Analysis of Paintings through an Art-History Agent

This approach converts detailed visual features into descriptive terms, addressing a key challenge in art history, and connects the use of images as data with the semantic concerns of humanists, establishing vision-based computational art history as an area for future growth.

M. Walton, Astrid Harth · 0 citations
Open access Aug 2026

Structured Creative Evaluation for Text-to-Image Generative AI Models

In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...

W.-C. Ma, Q. Zhang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.