Skip to content
Preprint

The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report

Aug 2026 · 0 citations · 54 references
Computer Science

TL;DR

Based on a three-year effort with Phonexia, a commercial speaker-recognition vendor, barriers encountered while building and deploying a detector are connected to concrete research and coordination proposals: shared standards for commercially usable datasets, realistic deployment benchmarks, and scores that non- experts can act on.

Abstract

Synthetic speech detection benchmarks now report sub-1% error rates on some in-domain evaluations, yet performance degrades under unseen attacks, channel mismatch, and distribution shift. Based on a three-year effort with Phonexia, a commercial speaker-recognition vendor, we report barriers encountered while building and deploying a detector. Many public benchmarks are not licensed for commercial model development. Real inputs are not four-second clean clips but long, codec-degraded, sometimes partially synthetic recordings. And when a calibrated system returns a log-likelihood ratio of 2.5, no one can tell the customer what it means for their decision. Rather than proposing a new model, we connect these barriers to concrete research and coordination proposals: shared standards for commercially usable datasets, realistic deployment benchmarks, and scores that non-experts can act on. These observations come from one project and should be tested in other settings.

View source

Similar papers

Preprint Aug 2026

AT-ADD: A Benchmark and Challenge for Robust and All-Type Audio Deepfake Detection

Recent audio generation models can synthesize high-fidelity speech, environmental sound, singing voice, and music, creating new risks for multimedia trust. Existing audio deepfake detection (ADD) benchmarks remain predominantly speech-centric and often underrepresent realistic channel variation and diverse audio types....

Yuan-Kun Xie, Hao-Nan Cheng, Jiayi Zhou et al. · 1 citation
Preprint Aug 2026

AT-ADD: All-Type Audio Deepfake Detection Challenge Summary

This paper summarizes the ACM Multimedia 2026 AT-ADD Grand Challenge on all-type audio deepfake detection. AT-ADD contains two tracks: robust speech deepfake detection under realistic acoustic and channel variations, and type-agnostic detection over speech, environmental sound, singing voice, and music. We describe the...

Yuan-Kun Xie, Hao-Nan Cheng, Jiayi Zhou et al. · 0 citations
Preprint Sep 2026

GenTraceBench: A Benchmark for Tracing Audio Deepfakes Across Pre- and Post-training Stages

Modern text-to-speech (TTS) systems are rarely deployed as unchanged pre-trained models. They are often adapted through supervised fine-tuning (SFT) or preference optimization such as DPO and GRPO. This raises a practical question for audio deepfake forensics: do fingerprints learned from a foundation generator remain...

Li Wang, Kun-Yu Feng, Wan Lin et al. · 0 citations
Preprint Sep 2026

Tracing and Relearning Detection Evidence in Text-to-Speech Systems

Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-BigVGAN pipeline. Since vocoder reconstruction of a real mel ca...

Eunji Shin, Kyudan Jung, Jihwan Kim et al. · 0 citations
Preprint Sep 2026

Domain-Incremental Learning for Multi-Channel Replay Speech Detection

This work frames this as Domain-Incremental Learning over acoustic environments and presents the first continual learning benchmark for multi-channel replay speech detection, evaluating a state-of-the-art beamformer-based detector over all 24 environment orderings of the ReMASC corpus with five seeds.

Michael Neri · 0 citations
Preprint Aug 2026

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

MADBench is introduced, the first benchmark that treats speech and environmental audio as distinct acoustic components, enabling component-aware evaluation of audio deepfake detection across independently manipulated forgery sources and establishes a rigorous foundation for future research into robust, component-aware...

Yan-Qiu Li, Yang Xiao, Jisheng Bai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.