ProFocus is a novel framework that models affective experience in artistic images via progressive visual focusing inspired by a hierarchical cognitive theory of human aesthetic appreciation and consistently outperforms state-of-the-art methods in both emotion recognition and affective explanation.
Abstract
Interpreting the emotional responses triggered by images is central to achieving emotional intelligence. Compared with natural images, visual art is intentionally created to elicit emotional responses from its viewers through abstract concepts and visual metaphors, making affective interpretation particularly challenging. However, most existing methods rely on general-purpose visual embeddings (e.g., CLIP), failing to capture the nuanced cues underlying artistic emotion. To address this gap, we propose \textbf{ProFocus}, a novel framework that models affective experience in artistic images via progressive visual focusing. The key idea is to model visual representation learning inspired by a hierarchical cognitive theory of human aesthetic appreciation. Technically, ProFocus contains two core components: a Hierarchical Art Critic (HAC) and a Progressive Hint Fusion (PHF) module. HAC leverages multimodal large language models to generate structured linguistic priors at three cognitive levels--atmospheric style, narrative subjects, and concrete details--thereby translating artistic perception into coherent semantic guidance. Building upon these priors, PHF departs from conventional cross-modal fusion by sequentially injecting the hierarchical hints into visual features, enabling a progressive focusing process that mirrors human perception. This design allows the model to capture subtle affective cues and produce more faithful explanations. Extensive experiments on the ArtEmis v1.0 and v2.0 datasets demonstrate that ProFocus consistently outperforms state-of-the-art methods in both emotion recognition and affective explanation. Project page: https://github.com/Zhang-Zhiyan/ProFocus.
Although text-to-image models can accurately depict subjects and scenes, creators still struggle to specify the fine-grained emotions an image should convey without rewriting its content description. Natural language can suggest emotions, but it offers no control scale with stable meanings and ordered intensities. We p...
Visual perception of built environments contributes to the affective impressions that people form in everyday life. However, how these impressions are represented within vision foundation models remains largely unexplored. To support the systematic investigation of this subject, we introduce the Emotional Impression of...
Madoka Yonekura, Katsunori Kohda, Nobuhiko Muramoto et al.· 0 citations
Key contributions include proposing a multi-task learning framework for jointly optimizing visual quality and emotion, establishing the inaugural VAWE-Art dataset comprising 5,000 AI-generated images with 20-dimensional emotional annotations, and providing computational foundations for emotion-controllable generative a...
Hengju Gang· Journal of Engineering, Proj...· 0 citations
By enabling machines to comprehend, interpret, and react to human affective states, visual mapping of multimodal emotion traits is crucial to intelligent interface design. Nevertheless, current multimodal emotion detection algorithms primarily focus on predictive performance, providing little insight into interpretabil...
Wan-Bao Ge, Zhen-Hua Yang, Zhen-Hu Liu et al.· Journal of Visualized Experi...· 0 citations
AI-driven generative art is changing the scope of artistic creation, but the principles behind its aesthetics and the perception of its outcomes remain relatively unexplored, especially concerning culturally specific art forms. The present study investigates how viewers cognitively, affectively, and evaluatively proc...
This work compares two parallel paradigms for composition analysis: a human-inspired method grounded in perceptual grouping, and fine-tuned foundation models enabled by recent large-scale compositional datasets.
F. Behrad, T. Tuytelaars, Johan Wagemans· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.