StreamFraudNet is introduced, which processes incoming audio through overlapping bounded-context windows using a frozen self-supervised speech encoder, recurrent temporal modeling, and learned aggregation of latent window scores to demonstrate that fraud risk can be scored incrementally from raw speech without transcri...
K. Vo, A. T. D. Dinh, Tien-Ta Tai et al.· 0 citations
Results clarify when causal cross-query memory improves repair and when broader memory representations remain preferable, and show that negative memory contributes modestly, the value of type conditioning and lexical-dense ranking is dataset dependent, and schema-local experience provides the most consistent benefit.
K. Vo, Tam Minh Chu, A. T. D. Dinh et al.· 0 citations
Text-compatible JEPA objectives must preserve multiple plausible completions rather than compress them into a single latent point, showing that text-compatible JEPA objectives must preserve multiple plausible completions rather than compress them into a single latent point.
A. T. D. Dinh, K. Vo· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.