 Qwen3-TTS is a powerful open-source text-to-speech model supporting voice cloning, voice design, and multilingual generation across 10 languages
3-second voice cloning: Clone any voice with just 3 seconds of audio input using Qwen3-TTS base models
State-of-the-art performance: Outperforms competitors like MiniMax, ElevenLabs, and SeedTTS in voice quality and speaker similarity
Dual-track streaming architecture: Achieve ultra-low latency of 97ms for real-time applications with Qwen3-TTS
Apache 2.0 license: Fully open-source models ranging from 0.6B to 1.7B parameters, available on HuggingFace and GitHub