Qwen3.8-LiveTranslate Debuts With Interleave Architecture
QwenTeam has introduced Qwen3.8-LiveTranslate, a real-time simultaneous interpretation system built on an Interleave architecture. Released on 2026/09/18, the model weaves audio and text together into a single stream, dropping average lagging (LAAL) from 2.8 seconds in the previous generation down to 2.3 seconds.
The system incorporates a Hybrid-MoE-based Thinker–Talker two-module design. The Thinker processes video, audio, source text, and translation sequentially, while the Talker synthesizes the translation into speech that preserves the timbre of the original speaker.
Qwen3.8-LiveTranslate adds three primary features to support real-world use cases:
- Real-time speaker separation with sentence attribution and voice cloning support
- Synchronized source-and-translation output on a single bilingual screen
- Long-context disambiguation for proper nouns and terminology across conversation turns
The model supports input audio and output text across 60 languages, alongside output audio capabilities for 29 languages. Evaluations across 70 language directions on the FLEURS audio test set and 14 language directions on the Omnilingua-MSpeaker benchmark show improvements in translation quality, faithfulness, fluency, conciseness, and speech recognition and synthesis.
Related AI News
Enjoyed this? Get more in your inbox.
Weekly AI breakthroughs, tool reviews, and practical guides.