Skip to content
Preprint

A Multi-View and Confusion-Guided Ensemble Framework for Robust Synthetic Image Attribution

Sep 2026 · 0 citations · 9 references
Computer Science

TL;DR

This report presents a multi-view and confusion-guided ensemble framework for the Synthetic Image Attribution Challenge of the DLMMDD Workshop at ICANN 2026, and introduces a dedicated binary expert classifier that is selectively activated under low-confidence conditions.

Abstract

Synthetic image attribution (SIA) has become increasingly important with the rapid advancement of text-to-image generation models. However, accurately identifying the source model of a generated image remains challenging due to the growing similarity among modern diffusion-based generators and the presence of diverse post-processing operations. In this report, we present a multi-view and confusion-guided ensemble framework for the Synthetic Image Attribution Challenge of the DLMMDD Workshop at ICANN 2026. Our approach integrates multiple complementary architectures, including FFT-ConvNeXt, DINOv2, CLIP, and Xception, to capture diverse attribution cues from frequency, semantic, and forensic perspectives. To improve robustness against unknown degradations and image manipulations, extensive data augmentation strategies are employed during training, simulating realistic post-processing operations such as compression, resizing, grayscale conversion, and blur. Furthermore, we analyze the confusion patterns of the ensemble model and observe severe ambiguity between Stable Diffusion 3 and Stable Diffusion 3.5. To address this issue, we introduce a dedicated binary expert classifier that is selectively activated under low-confidence conditions. We additionally apply class-adaptive confidence calibration to improve the discrimination of challenging classes such as Tencent Hunyuan. The proposed framework achieved 99.53% on the public leaderboard and 99.20% on the private leaderboard. The source code and implementation details are publicly available at https://github.com/ZOMIN28/SIA.

View source

Similar papers

2026

MF2DA: Multi-Level Feature Fusion for Robust Detection and Attribution of Universal AI-Generated Images

The proliferation of hyper-realistic AI-generated images poses significant threats to digital information integrity and forensic accountability. Existing detection methodologies, however, face three critical bottlenecks: vulnerability to real-world distortions such as social media compression, inadequate specialization...

Wen-Peng Mu, Qiang Xu, Yi-Ning Zhang et al. · 1 citation
Open access Aug 2026

Structured Creative Evaluation for Text-to-Image Generative AI Models

In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...

W.-C. Ma, Q. Zhang · 0 citations
Aug 2026

STAFuse: Scene-Text Aggregation-Guided Composite Degradation-Robust Infrared and Visible Image Fusion

Infrared and visible image fusion aims to integrate complementary information from source images to generate high-quality fusion images that serve downstream tasks. However, the differentiated representation of image scene content, the unpredictability of degradation modes in source images, and the complexity of compos...

Ting Lv, Hong Jiang, Yu Liu · 0 citations
Preprint Aug 2026

Scalable Black-Box Model Attribution for Images

The rapid proliferation of generative models raises the model attribution problem: given only an image, can we determine which model produced it? We propose a lightweight CNN to solve this problem in a strict black box setting. The CNN operates on multiple image patches to handle varying image size and improve accuracy...

Asaf Livne, Amir Jevnisek, S. Avidan · 0 citations
Preprint Aug 2026

Source-Agnostic Image Translation Based on Latent Aware Adaptive Masking

In this work, we propose a source-agnostic framework that dynamically refines a binary mask throughout the reverse diffusion process by computing the discrepancies of a pretrained diffusion model's prediction for each latent time step. Rather than relying on a fixed threshold, our method introduces a time-dependent sta...

Tomislav Dobricki, Byung-woo Hong · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.