Aug 2026· Journal of Business, Social and Technology· 0 citations· 25 references
TL;DR
It is demonstrated that integrating spatial and frequency-domain representations through a dual-branch Vision Transformer architecture enhances photo aesthetic assessment performance.
Abstract
Background: The rapid growth of digital media has increased the demand for automated image aesthetic assessment (IAA) in social media, creative industries, and e-commerce platforms. Although Vision Transformer (ViT)-based models have demonstrated promising performance, most existing approaches rely primarily on spatial representations while overlooking frequency-domain information, which captures complementary characteristics such as sharpness, texture, noise patterns, and bokeh effects.
Objective: This study proposes a Dual-Branch Late Fusion FFT-ViT architecture with concatenation-based fusion as an approach for integrating dual-domain representations to improve photo aesthetic assessment (PAA).
Methods: The proposed architecture consists of two parallel branches. The RGB branch employs a Vision Transformer (ViT-Small) to extract spatial and compositional features, while the FFT branch utilizes ViT-Tiny to capture frequency-domain characteristics associated with texture, sharpness, and image details. The extracted features from both branches are fused using a concatenation strategy before being passed to the regression layer.
Results: The proposed Dual-Branch FFT-ViT with concatenation fusion achieved the best performance, obtaining a PLCC of 0.7336, SRCC of 0.7347, MSE of 0.0193, MAE of 0.1117, and RMSE of 0.1388. Compared with the RGB-only ViT baseline, the proposed model improved the PLCC score by 0.0124, demonstrating the effectiveness of integrating spatial and frequency-domain features for aesthetic score prediction.
Conclusion: This study demonstrates that integrating spatial and frequency-domain representations through a dual-branch Vision Transformer architecture enhances photo aesthetic assessment performance.
A Swin Transformer-based NR-IQA method comprising three core modules, which improves Spearman Rank-Order Correlation Coefficient and Pearson Linear Correlation Coefficient over the best-performing comparison method, validating its effectiveness for no-reference image quality prediction.
Xiao-Meng Xia, Jia Yong, Yi-Biao Long et al.· PeerJ Computer Science· 0 citations
The rapid development of AI image generation technology has created an urgent need for systematic aesthetic evaluation of generated visual content. Existing computer vision assessment methods are often limited to technical image parameters and insufficiently consider multidimensional artistic judgment, texture details,...
Lingbo Yang, N. Yang· Advanced Electromagnetics· 0 citations
A hybrid multimodal framework combining intermediate cross-attention fusion with subsequent late fusion is proposed, which provides the strongest overall performance.
A cross-domain AI-generated IQA via content-distortion awareness (CDAQA) is proposed, designed to update the existing IQA model for AGIs and achieves higher accuracy and stability in cross-domain AGIs tasks.
Shun Zhu, Xi-Chen Yang, De-Chun Zhao et al.· Proceedings of the Thirty-Fi...· 0 citations
Most existing image aesthetic assessment methods rely on global representations or simple aggregation of local features. This makes it difficult to model semantic region information and provide region-level explanations. To address this issue, we propose a Semantic Segmentation-Guided Cross-Attention Fusion (SSG-CAF) m...