Skip to content
Preprint

A Convolutional Layer Activation Dimensionality Reduction for Out-of-Distribution and Adversarial Attack Detection Methods

Aug 2026 · 0 citations · 58 references
Computer Science

TL;DR

This paper carefully analyzes two state-of-the-art detection methods and their dimensionality reductions for convolutional layers and develops a novel reduction method with a controllable high-compression level.

Abstract

Despite the success of convolutional neural networks in image classification tasks and their general application in multi-modal models, their susceptibility to out-of-distribution and adversarial attack samples raises concerns regarding trustworthiness and safety. Among the approaches to tackle such issues, detection methods that analyze the model's intermediate activations to estimate a confidence score are a promising family that evaluates the decision process, relying on a dimensionality reduction step to enable efficient downstream processing of the high-dimensional activations. However, when considering convolutional layers, the dimensionality reduction methods in the literature either lack a mechanism to control the compression/information-loss trade-off or yield large representations. In this paper, we carefully analyze two state-of-the-art detection methods and their dimensionality reductions for convolutional layers and develop a novel reduction method with a controllable high-compression level. We extend these two state-of-the-art detection methods, enabling the usage of any dimensionality reduction, and evaluate their performance on out-of-distribution and adversarial attack detection. Results show that the detection methods with the proposed dimensionality reduction consistently perform better than, or comparable to, the strongest alternative. Furthermore, the proposed method is shown to reduce computation and memory footprints, given that it has the highest compression among the compared methods.

View source

Similar papers

Conference Open access 2026

Passive Detection of GAN-Generated Images: A Structured Investigation of Spatial and Frequency-Domain Approaches

The rapid advancements in generative adversarial networks (GANs) have led to the production of highly realistic synthetic images, posing severe threats to the credibility and authenticity of digital media across social platforms, news outlets, and official documents. Passive detection methods tackle this problem by ide...

Lin-Jie Lu · 0 citations
Sep 2026

Seeing Through Threats: Adversarial Detection Through Explainability (ADEx)

Deep Neural Networks (DNNs) remain vulnerable to adversarial perturbations, raising significant concerns in image processing applications, particularly in high-stakes domains such as medical imaging and security-critical systems. Most existing defense strategies are limited by domain specificity, architectural dependen...

Syamantak Sarkar, Nirmal Joseph, Sudhish N. George et al. · 0 citations
Preprint Sep 2026

Detection of Adversarial Attacks on Super-Resolvers Using Spectral Features

The integration of deep learning models into image preprocessing pipelines such as super-resolution introduces a largely unexplored attack vector for adversaries targeting downstream tasks. To ensure trustworthiness of critical imaging pipelines, we must be able to detect adversarial behavior within preprocessing model...

Emma J. Reid, Haley Duba-Sullivan, Tony G. Allen · 0 citations
Conference Open access 2026

Image Deepfake Detection Technologies: A Comprehensive Investigation of Methodologies, Challenges, and Future Trends

The rapid advancement of deep generative models, especially Generative Adversarial Networks (GANs) and Diffusion Models, has escalated the creation of highly realistic synthetic media, posing significant threats to information security through misinformation and fraud. The core of current detection methodologies encomp...

Tian-Hua Tang · 0 citations
Conference Sep 2026

FSED: A Feature-Space Ensemble Defense for Detecting Adversarial Examples

This paper introduces a Feature-Space Ensemble Defense (FSED) framework, which involves adversarial training, joint confidence calibration, and class-conditional Mahalanobis feature-space anomaly scoring to facilitate powerful adversarial detection. The proposed method considers the final-layer uncertainty. It also mod...

Aliza Saadi, Vanya Shafiq, Aaleen Zainab et al. · 0 citations
Review Open access Sep 2026

Sensing Deepfake Detection: A Survey of Detection Architectures, Adversarial Challenges, and Critical Applications in Political, Educational, and Military Domains

Deepfake technology has advanced swiftly, enabling the rapid production of hyper-realistic synthetic media that pose considerable threats to digital security, privacy, military operations, and information integrity. This paper extensively examines visual intelligence and computer vision methodologies for deepfake detec...

Alexandros Gazis, Stylianos Pappas, T. Vavouras et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.