Skip to content

Before the Warning Comes Too Late: Incremental Phone-Scam Detection from Speech

Jul 2026 · 0 citations · 32 references
Computer Science

TL;DR

StreamFraudNet is introduced, which processes incoming audio through overlapping bounded-context windows using a frozen self-supervised speech encoder, recurrent temporal modeling, and learned aggregation of latent window scores to demonstrate that fraud risk can be scored incrementally from raw speech without transcripts or temporal annotations.

Abstract

We study weakly supervised incremental telecom fraud detection from raw telephone audio, where training provides only conversation-level labels and predictions must be updated before a call ends. We introduce StreamFraudNet, which processes incoming audio through overlapping bounded-context windows using a frozen self-supervised speech encoder, recurrent temporal modeling, and learned aggregation of latent window scores. On a controlled English benchmark, StreamFraudNet achieves a ROC--AUC of \(0.9953\), significantly outperforming acoustic and mean-pooling baselines while remaining competitive with strong global temporal models. The model produces its first prediction after 10 seconds of audio, updates every 2 seconds, and operates faster than real time on the evaluated server hardware. Ablations identify recurrent temporal context as the principal contributor to performance. These results demonstrate that fraud risk can be scored incrementally from raw speech without transcripts or temporal annotations, while highlighting the need for latency-aware training to improve early prediction.

View source

Similar papers

Preprint Sep 2026

Domain-Incremental Learning for Multi-Channel Replay Speech Detection

Replay attacks are the most accessible threat to voice-controlled systems, and the acoustic cues that expose them are strongly modulated by the environment in which the attack is mounted. A detector deployed in the field therefore has to absorb new acoustic conditions over time, ideally without revisiting past recordin...

Michael Neri · 0 citations
#machine learning Preprint Sep 2026

Zero-Shot Temporal Localisation of Audio Deepfakes in Multi-Speaker Conversations

Voice-cloning fraud increasingly relies on surgical injection: a genuine conversation in which only one or two sentences are replaced by synthetic speech. Utterance-level deepfake detectors emit a single real/fake label per clip and cannot report where the synthetic speech lies. We formalise this as Temporal Deepfake L...

Soumyadeep Roy · 0 citations
#artificial intelligence Preprint Sep 2026

Look Less, Hear Better: Jointly Rewarded GRPO for Streaming ASR

AWED, a word-level emission-delay metric defined relative to the acoustic end of each word, is introduced and post-train a DSM recognizer with GRPO under a reward that scores transcription accuracy and measured delay jointly, advancing the accuracy--latency Pareto frontier of streaming ASR without architectural change.

Xiu-Wen Zheng · 0 citations
#natural language process... Preprint Sep 2026

FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech

Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and addition...

Puneet Mathur, Dinesh Manocha · 0 citations
#natural language process... Preprint Sep 2026

I'll Keep an Ear Out: Teaching AudioLLMs Proactive Audio Assistance

Interrupt and Silent Modeling (ISM), a model-agnostic paradigm that embeds proactive decisions into LLM decoding via two special tokens, is proposed, capturing four states: onset detection, sustained-relevance triggering, irrelevance suppression, and de-duplication.

Amit Kumar Singh Yadav, Ritvik Shrivastava, Xuan Zhang et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.