StableLearn Logo

Search Content

News 5 min read

Qwen3-TTS Open Source: Voice Cloning & Design in 10 Languages, 97ms Latency

Qwen3-TTS open-source speech synthesis models support 10 languages, voice cloning, and design. 0.6B-1.7B scales with 97ms latency streaming, outperforming commercial TTS services.

Cover image for Qwen3-TTS Open Source: Voice Cloning & Design in 10 Languages, 97ms Latency

Published 249 days ago. Content may be outdated.

In January 2025, Alibaba Cloud’s Qwen team officially open-sourced the complete Qwen3-TTS series, including VoiceDesign, CustomVoice, and Base variants across 0.6B and 1.7B parameter scales—5 models in total—covering deployment scenarios from edge devices to data centers.

Key Highlights

1. Fully Open Source, No Licensing Fees

The entire Qwen3-TTS series is completely open-source, supporting on-premises deployment and customization, eliminating per-minute API costs and providing enterprise-grade text-to-speech technology to developers.

2. 10-Language Support

Covers Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian, with multiple dialectal voice profiles to meet global application needs.

3. Three Core Capabilities

  • Voice Design: Generate new voices from natural language descriptions without predefined constraints
  • Custom Voice: 9 premium timbres with instruction-based control over timbre, emotion, and prosody
  • Voice Clone: Rapid 3-second voice cloning from reference audio

4. Ultra-Low Latency Streaming

Innovative Dual-Track hybrid streaming architecture achieves extreme bidirectional streaming speed:

  • First audio packet wait time equivalent to just one character
  • End-to-end synthesis latency as low as 97ms
  • Meets rigorous real-time interactive scenario demands

5. Powerful Speech Representation

  • Self-developed Qwen3-TTS-Tokenizer-12Hz for efficient acoustic compression and high-dimensional semantic modeling
  • Fully preserves paralinguistic information and acoustic environmental features
  • High-speed, high-fidelity speech reconstruction through lightweight non-DiT architecture

Model Architecture & Performance

Model List

ModelParametersFeaturesStreamingInstruction Control
Qwen3-TTS-12Hz-1.7B-VoiceDesign1.7BDescription-based voice generation✅✅
Qwen3-TTS-12Hz-1.7B-CustomVoice1.7B9 premium timbres + instruction control✅✅
Qwen3-TTS-12Hz-1.7B-Base1.7B3-second rapid clone + fine-tuning✅-
Qwen3-TTS-12Hz-0.6B-CustomVoice0.6B9 premium timbres✅-
Qwen3-TTS-12Hz-0.6B-Base0.6B3-second rapid clone + fine-tuning✅-

Performance Comparison

Zero-shot speech generation on Seed-TTS test set (lower WER is better):

ModelChinese WEREnglish WER
CosyVoice 30.711.45
MiniMax-Speech0.831.65
Qwen3-TTS-12Hz-1.7B-Base0.771.24
Qwen3-TTS-12Hz-0.6B-Base0.921.32
FireRedTTS 21.141.95

Qwen3-TTS demonstrates excellent content consistency, with performance matching or exceeding commercial services.

Quick Start

Environment Setup

   # Create Python 3.12 environment
conda create -n qwen3-tts python=3.12 -y
conda activate qwen3-tts

# Install qwen-tts package
pip install -U qwen-tts

# Recommended: Install FlashAttention 2 to reduce GPU memory usage
pip install -U flash-attn --no-build-isolation

1. Custom Voice Generation

   import torch
import soundfile as sf
from qwen_tts import Qwen3TTSModel

model = Qwen3TTSModel.from_pretrained(
    "Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
    device_map="cuda:0",
    dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",
)

# Single inference
wavs, sr = model.generate_custom_voice(
    text="Actually, I've really noticed that I'm particularly good at observing others' emotions.",
    language="English",
    speaker="Ryan",
    instruct="Speak in a very angry tone",
)
sf.write("output_custom_voice.wav", wavs[0], sr)

# Batch inference
wavs, sr = model.generate_custom_voice(
    text=[
        "Actually, I've really noticed that I'm particularly good at observing others' emotions.", 
        "She said she would be here by noon."
    ],
    language=["English", "English"],
    speaker=["Ryan", "Aiden"],
    instruct=["", "Very happy."]
)

Supported 9 Premium Speakers:

SpeakerDescriptionNative Language
VivianBright, slightly edgy young female voiceChinese
SerenaWarm, gentle young female voiceChinese
Uncle_FuSeasoned male voice with low, mellow timbreChinese
DylanYouthful Beijing male voice with clear, natural timbreChinese (Beijing Dialect)
EricLively Chengdu male voice with slightly husky brightnessChinese (Sichuan Dialect)
RyanDynamic male voice with strong rhythmic driveEnglish
AidenSunny American male voice with clear midrangeEnglish
Ono_AnnaPlayful Japanese female voice with light, nimble timbreJapanese
SoheeWarm Korean female voice with rich emotionKorean

2. Voice Design

   import torch
import soundfile as sf
from qwen_tts import Qwen3TTSModel

model = Qwen3TTSModel.from_pretrained(
    "Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign",
    device_map="cuda:0",
    dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",
)

wavs, sr = model.generate_voice_design(
    text="It's in the top drawer... wait, it's empty? No way, that's impossible! I'm sure I put it there!",
    language="English",
    instruct="Speak in an incredulous tone, but with a hint of panic beginning to creep into your voice.",
)
sf.write("output_voice_design.wav", wavs[0], sr)

3. Voice Clone

   import torch
import soundfile as sf
from qwen_tts import Qwen3TTSModel

model = Qwen3TTSModel.from_pretrained(
    "Qwen/Qwen3-TTS-12Hz-1.7B-Base",
    device_map="cuda:0",
    dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",
)

ref_audio = "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3-TTS-Repo/clone.wav"
ref_text = "Okay. Yeah. I resent you. I love you. I respect you. But you know what? You blew it! And thanks to you."

wavs, sr = model.generate_voice_clone(
    text="I am solving the equation: x = [-b ± √(b²-4ac)] / 2a? Nobody can — it's a disaster (◍•͈⌔•͈◍), very sad!",
    language="English",
    ref_audio=ref_audio,
    ref_text=ref_text,
)
sf.write("output_voice_clone.wav", wavs[0], sr)

4. Launch Local Web UI

   # CustomVoice model
qwen-tts-demo Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice --ip 0.0.0.0 --port 8000

# VoiceDesign model
qwen-tts-demo Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign --ip 0.0.0.0 --port 8000

# Base model
qwen-tts-demo Qwen/Qwen3-TTS-12Hz-1.7B-Base --ip 0.0.0.0 --port 8000

Then visit http://<your-ip>:8000 to experience it.

vLLM Support

vLLM officially provides day-0 support! Use vLLM-Omni for deployment and inference:

   # Clone vLLM-Omni repository
git clone https://github.com/vllm-project/vllm-omni.git
cd vllm-omni/examples/offline_inference/qwen3_tts

# Single sample with CustomVoice task
python end2end.py --query-type CustomVoice

# Batch sample with CustomVoice task
python end2end.py --query-type CustomVoice --use-batch-sample

# VoiceDesign task inference
python end2end.py --query-type VoiceDesign

# Base model ICL mode inference
python end2end.py --query-type Base --mode-tag icl

DashScope API Service

In addition to open-source models, Alibaba Cloud provides DashScope API services for faster and more efficient experiences:

Application Scenarios

  1. Education & Cultural Preservation: Support dialect digitization, language learning, and cultural protection
  2. Content Creation: High-quality voiceovers for short videos, games, and audiobooks
  3. Intelligent Interaction: Optimize customer service, voice assistants, and accessibility services
  4. Cross-language Applications: Multilingual support for global products
  5. Personalized Services: Create unique brand voices through voice cloning and design

Technical Advantages

Universal End-to-End Architecture

Utilizes discrete multi-codebook LM architecture for full-information end-to-end speech modeling, completely bypassing information bottlenecks and cascading errors inherent in traditional LM+DiT schemes, significantly enhancing versatility, generation efficiency, and performance ceiling.

Intelligent Text Understanding & Voice Control

Supports natural language instruction-driven speech generation with flexible control over multi-dimensional acoustic attributes like timbre, emotion, and prosody. Deep integration of text semantic understanding enables adaptive adjustment of tone, rhythm, and emotional expression for lifelike “what you imagine is what you hear” output.

Strong Robustness

Demonstrates markedly improved robustness to noisy input text, intelligently handling complex text with automatic adaptation of tone and fluency.

Summary

The full open-sourcing of Qwen3-TTS marks the democratization of enterprise-grade text-to-speech technology, providing developers with unprecedented flexibility by offering quality matching or exceeding commercial competitors while eliminating per-minute API costs. Dual model sizes (0.6B and 1.7B) enable flexible deployment from edge devices to data centers, while 10-language support and streaming architecture make it ideal for multilingual applications.

Related Links:

Share Article

More Articles