Skip to content
Preprint

ProFocus: Interpreting Affective Experience in Artistic Images with Progressive Visual Focusing

Aug 2026 · 0 citations · 55 references
Computer Science

TL;DR

ProFocus is a novel framework that models affective experience in artistic images via progressive visual focusing inspired by a hierarchical cognitive theory of human aesthetic appreciation and consistently outperforms state-of-the-art methods in both emotion recognition and affective explanation.

Abstract

Interpreting the emotional responses triggered by images is central to achieving emotional intelligence. Compared with natural images, visual art is intentionally created to elicit emotional responses from its viewers through abstract concepts and visual metaphors, making affective interpretation particularly challenging. However, most existing methods rely on general-purpose visual embeddings (e.g., CLIP), failing to capture the nuanced cues underlying artistic emotion. To address this gap, we propose \textbf{ProFocus}, a novel framework that models affective experience in artistic images via progressive visual focusing. The key idea is to model visual representation learning inspired by a hierarchical cognitive theory of human aesthetic appreciation. Technically, ProFocus contains two core components: a Hierarchical Art Critic (HAC) and a Progressive Hint Fusion (PHF) module. HAC leverages multimodal large language models to generate structured linguistic priors at three cognitive levels--atmospheric style, narrative subjects, and concrete details--thereby translating artistic perception into coherent semantic guidance. Building upon these priors, PHF departs from conventional cross-modal fusion by sequentially injecting the hierarchical hints into visual features, enabling a progressive focusing process that mirrors human perception. This design allows the model to capture subtle affective cues and produce more faithful explanations. Extensive experiments on the ArtEmis v1.0 and v2.0 datasets demonstrate that ProFocus consistently outperforms state-of-the-art methods in both emotion recognition and affective explanation. Project page: https://github.com/Zhang-Zhiyan/ProFocus.

View source

Similar papers

Preprint Sep 2026

Beyond Emotion Prompts: Fine-Grained Text-to-Image Generation Driven by Valence-Arousal-Dominance

Although text-to-image models can accurately depict subjects and scenes, creators still struggle to specify the fine-grained emotions an image should convey without rewriting its content description. Natural language can suggest emotions, but it offers no control scale with stable meanings and ordered intensities. We p...

Ming-Lan Li, Yue-Yue Fang, Xie-Ping Gao · 0 citations
Review Sep 2026

Organization of Valence and Arousal in Vision-Language Representations of Built Environments: Insights from the EMOIS Dataset

Visual perception of built environments contributes to the affective impressions that people form in everyday life. However, how these impressions are represented within vision foundation models remains largely unexplored. To support the systematic investigation of this subject, we introduce the Emotional Impression of...

Madoka Yonekura, Katsunori Kohda, Nobuhiko Muramoto et al. · 0 citations
#artificial intelligence Review Open access Nov 2026

Generation and Perception: A Computational Evaluation Method for Visual Quality and Emotional Impact in AI Artworks

Key contributions include proposing a multi-task learning framework for jointly optimizing visual quality and emotion, establishing the inaugural VAWE-Art dataset comprising 5,000 AI-generated images with 20-dimensional emotional annotations, and providing computational foundations for emotion-controllable generative a...

Hengju Gang · 0 citations
Aug 2026

Visual Mapping of Multimodal Emotion Features for Intelligent Interaction Design.

By enabling machines to comprehend, interpret, and react to human affective states, visual mapping of multimodal emotion traits is crucial to intelligent interface design. Nevertheless, current multimodal emotion detection algorithms primarily focus on predictive performance, providing little insight into interpretabil...

Wan-Bao Ge, Zhen-Hua Yang, Zhen-Hu Liu et al. · 0 citations
Open access Sep 2026

Dual-pathway processing of AI-generated Chinese ink painting: evidence from eye-tracking, EEG, and artistic expertise

AI-driven generative art is changing the scope of artistic creation, but the principles behind its aesthetics and the perception of its outcomes remain relatively unexplored, especially concerning culturally specific art forms. The present study investigates how viewers cognitively, affectively, and evaluatively proc...

Qiang Qin, Jing-Chin Chai, Wan-Ni Zhong · 0 citations
Preprint Aug 2026

Learning visual representations for compositional analysis of artworks and photographs

This work compares two parallel paradigms for composition analysis: a human-inspired method grounded in perceptual grouping, and fine-tuned foundation models enabled by recent large-scale compositional datasets.

F. Behrad, T. Tuytelaars, Johan Wagemans · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.