Skip to content
Conference Open access

DeepL Voice: Real-Time Speech-to-Speech Translation

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · 0 citations · 7 references

Abstract

DeepL Voice is a real-time speech-to-speech translation system for global business communication, following a pragmatic incremental approach: developing a production-grade cascaded speech-to-speech-translation (S2ST) system, while exploring end-to-end solutions in parallel. The production system (launched November 2024) achieves competitive transcription quality through proprietary real-time ASR models and eliminates translation "flickering" via stable text streaming while maintaining low latency. Supporting 18 input languages and 30+ target languages, it offers DeepL Voice for Meetings (Microsoft Teams/Zoom integration), DeepL Voice for Conversations (mobile apps), as well as the DeepL API for Voice. Key features include customizable formality and glossary support for business-appropriate communication, with voice cloning TTS under development.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.