Qwen3-TTS Open Source: Voice Cloning & Design in 10 Languages, 97ms Latency
Qwen3-TTS open-source speech synthesis models support 10 languages, voice cloning, and design. 0.6B-1.7B scales with 97ms latency streaming, outperforming commercial TTS services.
Published 249 days ago. Content may be outdated.
In January 2025, Alibaba Cloud’s Qwen team officially open-sourced the complete Qwen3-TTS series, including VoiceDesign, CustomVoice, and Base variants across 0.6B and 1.7B parameter scales—5 models in total—covering deployment scenarios from edge devices to data centers.
Key Highlights
1. Fully Open Source, No Licensing Fees
The entire Qwen3-TTS series is completely open-source, supporting on-premises deployment and customization, eliminating per-minute API costs and providing enterprise-grade text-to-speech technology to developers.
2. 10-Language Support
Covers Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian, with multiple dialectal voice profiles to meet global application needs.
3. Three Core Capabilities
- Voice Design: Generate new voices from natural language descriptions without predefined constraints
- Custom Voice: 9 premium timbres with instruction-based control over timbre, emotion, and prosody
- Voice Clone: Rapid 3-second voice cloning from reference audio
4. Ultra-Low Latency Streaming
Innovative Dual-Track hybrid streaming architecture achieves extreme bidirectional streaming speed:
- First audio packet wait time equivalent to just one character
- End-to-end synthesis latency as low as 97ms
- Meets rigorous real-time interactive scenario demands
5. Powerful Speech Representation
- Self-developed Qwen3-TTS-Tokenizer-12Hz for efficient acoustic compression and high-dimensional semantic modeling
- Fully preserves paralinguistic information and acoustic environmental features
- High-speed, high-fidelity speech reconstruction through lightweight non-DiT architecture
Model Architecture & Performance
Model List
| Model | Parameters | Features | Streaming | Instruction Control |
|---|---|---|---|---|
| Qwen3-TTS-12Hz-1.7B-VoiceDesign | 1.7B | Description-based voice generation | ✅ | ✅ |
| Qwen3-TTS-12Hz-1.7B-CustomVoice | 1.7B | 9 premium timbres + instruction control | ✅ | ✅ |
| Qwen3-TTS-12Hz-1.7B-Base | 1.7B | 3-second rapid clone + fine-tuning | ✅ | - |
| Qwen3-TTS-12Hz-0.6B-CustomVoice | 0.6B | 9 premium timbres | ✅ | - |
| Qwen3-TTS-12Hz-0.6B-Base | 0.6B | 3-second rapid clone + fine-tuning | ✅ | - |
Performance Comparison
Zero-shot speech generation on Seed-TTS test set (lower WER is better):
| Model | Chinese WER | English WER |
|---|---|---|
| CosyVoice 3 | 0.71 | 1.45 |
| MiniMax-Speech | 0.83 | 1.65 |
| Qwen3-TTS-12Hz-1.7B-Base | 0.77 | 1.24 |
| Qwen3-TTS-12Hz-0.6B-Base | 0.92 | 1.32 |
| FireRedTTS 2 | 1.14 | 1.95 |
Qwen3-TTS demonstrates excellent content consistency, with performance matching or exceeding commercial services.
Quick Start
Environment Setup
# Create Python 3.12 environment
conda create -n qwen3-tts python=3.12 -y
conda activate qwen3-tts
# Install qwen-tts package
pip install -U qwen-tts
# Recommended: Install FlashAttention 2 to reduce GPU memory usage
pip install -U flash-attn --no-build-isolation
1. Custom Voice Generation
import torch
import soundfile as sf
from qwen_tts import Qwen3TTSModel
model = Qwen3TTSModel.from_pretrained(
"Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
device_map="cuda:0",
dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
)
# Single inference
wavs, sr = model.generate_custom_voice(
text="Actually, I've really noticed that I'm particularly good at observing others' emotions.",
language="English",
speaker="Ryan",
instruct="Speak in a very angry tone",
)
sf.write("output_custom_voice.wav", wavs[0], sr)
# Batch inference
wavs, sr = model.generate_custom_voice(
text=[
"Actually, I've really noticed that I'm particularly good at observing others' emotions.",
"She said she would be here by noon."
],
language=["English", "English"],
speaker=["Ryan", "Aiden"],
instruct=["", "Very happy."]
)
Supported 9 Premium Speakers:
| Speaker | Description | Native Language |
|---|---|---|
| Vivian | Bright, slightly edgy young female voice | Chinese |
| Serena | Warm, gentle young female voice | Chinese |
| Uncle_Fu | Seasoned male voice with low, mellow timbre | Chinese |
| Dylan | Youthful Beijing male voice with clear, natural timbre | Chinese (Beijing Dialect) |
| Eric | Lively Chengdu male voice with slightly husky brightness | Chinese (Sichuan Dialect) |
| Ryan | Dynamic male voice with strong rhythmic drive | English |
| Aiden | Sunny American male voice with clear midrange | English |
| Ono_Anna | Playful Japanese female voice with light, nimble timbre | Japanese |
| Sohee | Warm Korean female voice with rich emotion | Korean |
2. Voice Design
import torch
import soundfile as sf
from qwen_tts import Qwen3TTSModel
model = Qwen3TTSModel.from_pretrained(
"Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign",
device_map="cuda:0",
dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
)
wavs, sr = model.generate_voice_design(
text="It's in the top drawer... wait, it's empty? No way, that's impossible! I'm sure I put it there!",
language="English",
instruct="Speak in an incredulous tone, but with a hint of panic beginning to creep into your voice.",
)
sf.write("output_voice_design.wav", wavs[0], sr)
3. Voice Clone
import torch
import soundfile as sf
from qwen_tts import Qwen3TTSModel
model = Qwen3TTSModel.from_pretrained(
"Qwen/Qwen3-TTS-12Hz-1.7B-Base",
device_map="cuda:0",
dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
)
ref_audio = "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3-TTS-Repo/clone.wav"
ref_text = "Okay. Yeah. I resent you. I love you. I respect you. But you know what? You blew it! And thanks to you."
wavs, sr = model.generate_voice_clone(
text="I am solving the equation: x = [-b ± √(b²-4ac)] / 2a? Nobody can — it's a disaster (◍•͈⌔•͈◍), very sad!",
language="English",
ref_audio=ref_audio,
ref_text=ref_text,
)
sf.write("output_voice_clone.wav", wavs[0], sr)
4. Launch Local Web UI
# CustomVoice model
qwen-tts-demo Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice --ip 0.0.0.0 --port 8000
# VoiceDesign model
qwen-tts-demo Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign --ip 0.0.0.0 --port 8000
# Base model
qwen-tts-demo Qwen/Qwen3-TTS-12Hz-1.7B-Base --ip 0.0.0.0 --port 8000
Then visit http://<your-ip>:8000 to experience it.
vLLM Support
vLLM officially provides day-0 support! Use vLLM-Omni for deployment and inference:
# Clone vLLM-Omni repository
git clone https://github.com/vllm-project/vllm-omni.git
cd vllm-omni/examples/offline_inference/qwen3_tts
# Single sample with CustomVoice task
python end2end.py --query-type CustomVoice
# Batch sample with CustomVoice task
python end2end.py --query-type CustomVoice --use-batch-sample
# VoiceDesign task inference
python end2end.py --query-type VoiceDesign
# Base model ICL mode inference
python end2end.py --query-type Base --mode-tag icl
DashScope API Service
In addition to open-source models, Alibaba Cloud provides DashScope API services for faster and more efficient experiences:
- Real-time API (CustomVoice): China Docs | International Docs
- Real-time API (VoiceClone): China Docs | International Docs
- Real-time API (VoiceDesign): China Docs | International Docs
Application Scenarios
- Education & Cultural Preservation: Support dialect digitization, language learning, and cultural protection
- Content Creation: High-quality voiceovers for short videos, games, and audiobooks
- Intelligent Interaction: Optimize customer service, voice assistants, and accessibility services
- Cross-language Applications: Multilingual support for global products
- Personalized Services: Create unique brand voices through voice cloning and design
Technical Advantages
Universal End-to-End Architecture
Utilizes discrete multi-codebook LM architecture for full-information end-to-end speech modeling, completely bypassing information bottlenecks and cascading errors inherent in traditional LM+DiT schemes, significantly enhancing versatility, generation efficiency, and performance ceiling.
Intelligent Text Understanding & Voice Control
Supports natural language instruction-driven speech generation with flexible control over multi-dimensional acoustic attributes like timbre, emotion, and prosody. Deep integration of text semantic understanding enables adaptive adjustment of tone, rhythm, and emotional expression for lifelike “what you imagine is what you hear” output.
Strong Robustness
Demonstrates markedly improved robustness to noisy input text, intelligently handling complex text with automatic adaptation of tone and fluency.
Summary
The full open-sourcing of Qwen3-TTS marks the democratization of enterprise-grade text-to-speech technology, providing developers with unprecedented flexibility by offering quality matching or exceeding commercial competitors while eliminating per-minute API costs. Dual model sizes (0.6B and 1.7B) enable flexible deployment from edge devices to data centers, while 10-language support and streaming architecture make it ideal for multilingual applications.
Related Links:
- GitHub Repository: https://github.com/QwenLM/Qwen3-TTS
- Hugging Face Models: https://huggingface.co/Qwen
- ModelScope Models: https://modelscope.cn/organization/Qwen
More Articles