A data-generation pipeline that captions real photographs with a vision–language model and regenerates them with modern text-to-image systems, producing semantically aligned real/synthetic pairs that isolate generative artifacts from image content is described.
Abstract
We study the detection of AI-generated images and contribute toward detectors that are accurate, trustworthy, and honestly evaluated. We first describe a data-generation pipeline that captions real photographs with a vision–language model and regenerates them with modern text-to-image systems, producing semantically aligned real/synthetic pairs that isolate generative artifacts from image content. We then build a lightweight, CPU-deployable spectral detector that fuses an RGB backbone with a radial residual-frequency branch, and show through a controlled ablation that the frequency cue mainly contributes calibration and false-positive control rather than raw ranking. To keep a single saturated score from overstating readiness, we package the detector with a multi-axis evaluation suite and a worst-group summary metric. Finally, we extend the same recipe to a dedicated face-deepfake detector that adds a neighboring-pixel-relationship branch and generalizes well to generators unseen in training. We release the models, the data protocols, and all per-sample scores as a reproducible reference point for synthetic-image forensics.
GurAI is proposed, a transparent logistic late-fusion method that combines Rich384 and DeMamba logits that suggests that transparent late fusion can exploit complementary detector strengths more effectively than architectural redesign alone when facing generator diversity.
GenSyn10 is introduced, a CIFAR-10-aligned synthetic image dataset of 60,000 images generated using three architecturally diverse state-of-the-art models, enabling controlled and systematic evaluation of out-of-distribution (OOD) generalization to novel generators.
Md Faraz Kabir Khan, Saeed Anwar, G. Hassan· 0 citations
Across cross-generator, post-processing, and in-the-wild benchmarks, PE-SPC surpasses the previous DINOv3 baseline and achieves new state-of-the-art results.
Wei-Han Cai, Hao Tan, Zichang Tan et al.· 0 citations
VendorBench-100, a cross-paradigm benchmark that evaluates 36 representative models using a single adversarial 100-image corpus, a unified output schema, and a common evaluation framework, is introduced.
S. N. Deshmukh, Md. Rashidunnabi, Nelton Tiago Gemo et al.· 0 citations
This work introduces a dual-branch ensemble framework fusing Semantic Deep Learning with Mathematical Forensic Feature Extraction, highlighting the practicality and scalability of mathematical forensics for real-world deployment.
The rapid advancement of generative AI has enabled the creation of highly realistic deepfake media, posing significant threats, including misinformation, digital identity theft, fraud, and manipulation of public opinion. AI-generated image (AIGI) detection is reliably challenging due to the diversity of generative methods and the subtle artifacts they leave behind. In this work, we propose GenRes, a novel framework for generative residual learning via a neural tensor network, which models fine-grained relational features between original and transformed samples to enhance generalization. To address scenarios involving multiple generative transformations, we introduce GenRes++, which employs a learnable attention mechanism to aggregate relational features across multiple transformed samples and enables the model to focus on the most informative cues. Both models leverage PE-Core as a feature extractor, providing generalized and semantically rich embeddings that improve cross-domain performance and enable the detection of AIGI generated by unseen methods. Comprehensive experiments on multiple benchmark datasets demonstrate that the proposed GenRes++ approach outperforms existing methods.
Kutub Uddin, Nusrat Tasnim, Awais Khan et al.· 2 citations