Z-Image: Alibaba's 6B Open-Source Text-to-Image Model | 8-Step Generation Rivals Closed-Source SOTA
Z-Image is Alibaba Tongyi's 6B parameter open-source text-to-image model, rivaling closed-source SOTA! 8-step inference with sub-second latency, runs on 16GB VRAM. Supports Chinese & English text rendering, leads in AI Arena human preference evaluation.
Published 295 days ago. Content may be outdated.
Z-Image is the latest high-efficiency image generation foundation model released by Alibaba’s Tongyi Lab. This model is truly impressive — with 6 billion parameters, it can generate high-quality images in just 8 inference steps, achieving sub-second latency on enterprise H800 GPUs, and even runs smoothly on consumer GPUs with 16GB VRAM!
Imagine simply entering a text description, and Z-Image quickly generates photorealistic images while accurately rendering Chinese and English text. This combination of efficiency and quality represents another significant breakthrough in image generation.
For example, you can ask Z-Image to generate an image of “a Chinese girl in red Hanfu, holding a folding fan, with Xi’an’s Giant Wild Goose Pagoda in the background” — it not only accurately understands these complex descriptions but also generates stunning high-quality images!
Key Highlights
- Ultra Efficient — Only 8 inference steps, sub-second latency on H800
- Low VRAM Requirements — Runs on consumer GPUs with 16GB VRAM
- Photorealistic Quality — Excellent realistic image generation capability
- Bilingual Text Rendering — Accurately renders complex Chinese and English text
- Strong Instruction Following — Precisely understands and executes text descriptions
- Open Source — Model weights available on Hugging Face and ModelScope
Model Architecture
S3-DiT Architecture
Z-Image adopts an innovative Scalable Single-Stream DiT (S3-DiT) architecture. Unlike traditional dual-stream approaches, S3-DiT concatenates text, visual semantic tokens, and image VAE tokens at the sequence level as a unified input stream, maximizing parameter efficiency.
This design brings several key advantages:
- Higher Parameter Efficiency — Single-stream architecture utilizes model parameters more fully than dual-stream methods
- Faster Inference — Unified input stream reduces computational overhead
- Better Scalability — Architecture design supports flexible model scaling
Model Variants
Z-Image currently has three variants:
| Model | Description | Status |
|---|---|---|
| Z-Image-Turbo | Distilled version, 8-step inference, sub-second latency | ✅ Released |
| Z-Image-Base | Non-distilled base model, supports community fine-tuning | 🔜 Coming Soon |
| Z-Image-Edit | Image editing version, supports natural language editing instructions | 🔜 Coming Soon |
Core Technologies
Decoupled-DMD Distillation Algorithm
Z-Image-Turbo’s efficient inference benefits from the Decoupled-DMD distillation algorithm. The core insight of this algorithm is that the success of existing DMD (Distribution Matching Distillation) methods comes from two independently cooperating mechanisms:
- CFG Augmentation (CA) — The main engine driving the distillation process
- Distribution Matching (DM) — Acts as a regularizer, ensuring output stability and quality
By decoupling these two mechanisms, the team was able to study and optimize them independently, ultimately developing an improved distillation pipeline that significantly enhances few-step generation performance.
DMDR Reinforcement Learning Integration
Building on Decoupled-DMD, the team further proposed DMDR, synergistically integrating reinforcement learning (RL) with DMD:
- RL Unlocks DMD’s Performance Potential — Enhances the model’s generation capability
- DMD Effectively Regularizes RL — Maintains output stability
This integration further improves semantic alignment, aesthetic quality, and structural coherence, generating images with richer high-frequency details.
Model Download
Hugging Face
- Model Weights: Tongyi-MAI/Z-Image-Turbo
- Online Demo: Z-Image-Turbo Space
ModelScope
- Model Weights: Tongyi-MAI/Z-Image-Turbo
- Online Demo: Z-Image-Turbo Online Experience
Environment Setup
System Requirements
| Item | Requirement |
|---|---|
| Operating System | Linux / Windows / MacOS |
| Python | 3.8+ |
| PyTorch | 2.0+ |
| CUDA | 11.8+ (Recommended) |
| GPU | NVIDIA GPU (16GB+ VRAM recommended) |
| Memory | 16GB+ RAM |
Step 1: Install diffusers
Z-Image has been integrated into the official 🤗 diffusers library. You need to install the latest version from source:
# Install latest diffusers from source
pip install git+https://github.com/huggingface/diffusers
Why install from source?
Z-Image support has been merged into the official diffusers repository through two PRs (#12703 and #12715). To get the latest Z-Image support, you need to install from source.
Step 2: Install Other Dependencies
# Install PyTorch (choose based on your CUDA version)
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu118
# Optional: Install Flash Attention for better performance
pip install flash-attn --no-build-isolation
Step 3: Verify Installation
# Test if installation is successful
import torch
from diffusers import ZImagePipeline
print(f"PyTorch version: {torch.__version__}")
print(f"CUDA available: {torch.cuda.is_available()}")
print("Z-Image installation successful!")
Quick Start
Basic Image Generation
import torch
from diffusers import ZImagePipeline
# 1. Load model
# Use bfloat16 for best performance
print("Loading Z-Image-Turbo model...")
pipe = ZImagePipeline.from_pretrained(
"Tongyi-MAI/Z-Image-Turbo",
torch_dtype=torch.bfloat16,
low_cpu_mem_usage=False,
)
pipe.to("cuda")
print("Model loaded!")
# 2. Set prompt
prompt = "Young Chinese woman in red Hanfu, intricate embroidery. Impeccable makeup, red floral forehead pattern. Elaborate high bun, golden phoenix headdress, red flowers, beads. Holds round folding fan with lady, trees, bird. Neon lightning-bolt lamp (⚡️), bright yellow glow, above extended left palm. Soft-lit outdoor night background, silhouetted tiered pagoda (Xi'an Giant Wild Goose Pagoda), blurred colorful distant lights."
# 3. Generate image
print("Generating image...")
image = pipe(
prompt=prompt,
height=1024,
width=1024,
num_inference_steps=9, # Actually 8 DiT forward passes
guidance_scale=0.0, # Turbo model guidance should be 0
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
# 4. Save result
image.save("z_image_example.png")
print("Image saved to: z_image_example.png")
Chinese Prompt Generation
Z-Image has excellent support for Chinese prompts:
import torch
from diffusers import ZImagePipeline
# Load model
pipe = ZImagePipeline.from_pretrained(
"Tongyi-MAI/Z-Image-Turbo",
torch_dtype=torch.bfloat16,
)
pipe.to("cuda")
# Chinese prompt
prompt = "一位穿着精美汉服的中国古典美女,站在江南水乡的石桥上,手持油纸伞,背景是烟雨蒙蒙的白墙黑瓦建筑,柳树依依,水面倒影,画面充满诗意"
# Generate image
image = pipe(
prompt=prompt,
height=1024,
width=1024,
num_inference_steps=9,
guidance_scale=0.0,
generator=torch.Generator("cuda").manual_seed(123),
).images[0]
image.save("chinese_beauty.png")
print("Chinese prompt image generation complete!")
Text Rendering Example
Z-Image excels at Chinese and English text rendering:
import torch
from diffusers import ZImagePipeline
pipe = ZImagePipeline.from_pretrained(
"Tongyi-MAI/Z-Image-Turbo",
torch_dtype=torch.bfloat16,
)
pipe.to("cuda")
# Prompt with text
prompt = 'A professional business card design with the text "创意设计工作室" and "Creative Design Studio" elegantly displayed, minimalist style, white background, gold accents'
image = pipe(
prompt=prompt,
height=1024,
width=1024,
num_inference_steps=9,
guidance_scale=0.0,
generator=torch.Generator("cuda").manual_seed(456),
).images[0]
image.save("text_rendering.png")
print("Text rendering image generation complete!")
Advanced Configuration
Enable Flash Attention
If your GPU supports Flash Attention, enable it for better efficiency:
import torch
from diffusers import ZImagePipeline
pipe = ZImagePipeline.from_pretrained(
"Tongyi-MAI/Z-Image-Turbo",
torch_dtype=torch.bfloat16,
)
pipe.to("cuda")
# Enable Flash Attention 2
pipe.transformer.set_attention_backend("flash")
# Or enable Flash Attention 3 (if supported)
# pipe.transformer.set_attention_backend("_flash_3")
# Generate image
image = pipe(
prompt="A beautiful sunset over the ocean",
height=1024,
width=1024,
num_inference_steps=9,
guidance_scale=0.0,
).images[0]
Model Compilation Acceleration
Use PyTorch compilation for further inference acceleration:
import torch
from diffusers import ZImagePipeline
pipe = ZImagePipeline.from_pretrained(
"Tongyi-MAI/Z-Image-Turbo",
torch_dtype=torch.bfloat16,
)
pipe.to("cuda")
# Compile DiT model (first run will be slower, subsequent runs faster)
pipe.transformer.compile()
# Generate image
image = pipe(
prompt="A futuristic city skyline at night",
height=1024,
width=1024,
num_inference_steps=9,
guidance_scale=0.0,
).images[0]
CPU Offload (Low VRAM Devices)
If your VRAM is limited, enable CPU offload:
import torch
from diffusers import ZImagePipeline
pipe = ZImagePipeline.from_pretrained(
"Tongyi-MAI/Z-Image-Turbo",
torch_dtype=torch.bfloat16,
)
# Enable CPU offload to save VRAM
pipe.enable_model_cpu_offload()
# Generate image
image = pipe(
prompt="A cozy coffee shop interior",
height=1024,
width=1024,
num_inference_steps=9,
guidance_scale=0.0,
).images[0]
Practical Use Cases
Batch Image Generation
import torch
from diffusers import ZImagePipeline
def batch_generate_images(prompts, output_dir="outputs"):
"""
Batch generate images
"""
import os
os.makedirs(output_dir, exist_ok=True)
# Load model
pipe = ZImagePipeline.from_pretrained(
"Tongyi-MAI/Z-Image-Turbo",
torch_dtype=torch.bfloat16,
)
pipe.to("cuda")
results = []
for i, prompt in enumerate(prompts):
print(f"Generating image {i+1}/{len(prompts)}...")
image = pipe(
prompt=prompt,
height=1024,
width=1024,
num_inference_steps=9,
guidance_scale=0.0,
generator=torch.Generator("cuda").manual_seed(i * 42),
).images[0]
output_path = os.path.join(output_dir, f"image_{i:03d}.png")
image.save(output_path)
results.append({
"prompt": prompt,
"output": output_path
})
print(f"✓ Saved: {output_path}")
print(f"Batch generation complete! Generated {len(results)} images")
return results
# Usage example
prompts = [
"A serene Japanese garden with cherry blossoms",
"A cyberpunk street scene with neon lights",
"A cozy library with warm lighting",
"A majestic mountain landscape at sunrise",
]
results = batch_generate_images(prompts)
Different Size Image Generation
import torch
from diffusers import ZImagePipeline
def generate_multiple_sizes(prompt, sizes):
"""
Generate images at different sizes
"""
pipe = ZImagePipeline.from_pretrained(
"Tongyi-MAI/Z-Image-Turbo",
torch_dtype=torch.bfloat16,
)
pipe.to("cuda")
results = []
for width, height in sizes:
print(f"Generating {width}x{height} image...")
image = pipe(
prompt=prompt,
height=height,
width=width,
num_inference_steps=9,
guidance_scale=0.0,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
output_path = f"output_{width}x{height}.png"
image.save(output_path)
results.append({
"size": (width, height),
"output": output_path
})
print(f"✓ Saved: {output_path}")
return results
# Usage example
prompt = "A beautiful landscape with mountains and lake"
sizes = [
(1024, 1024), # Square
(1024, 768), # Landscape
(768, 1024), # Portrait
]
results = generate_multiple_sizes(prompt, sizes)
Random Seed Exploration
import torch
from diffusers import ZImagePipeline
def explore_seeds(prompt, num_seeds=5):
"""
Explore effects of different random seeds
"""
pipe = ZImagePipeline.from_pretrained(
"Tongyi-MAI/Z-Image-Turbo",
torch_dtype=torch.bfloat16,
)
pipe.to("cuda")
results = []
for seed in range(num_seeds):
print(f"Generating with seed {seed}...")
image = pipe(
prompt=prompt,
height=1024,
width=1024,
num_inference_steps=9,
guidance_scale=0.0,
generator=torch.Generator("cuda").manual_seed(seed),
).images[0]
output_path = f"seed_{seed:03d}.png"
image.save(output_path)
results.append({
"seed": seed,
"output": output_path
})
print(f"✓ Seed {seed} complete")
print(f"Exploration complete! Generated {len(results)} images with different seeds")
return results
# Usage example
prompt = "A magical forest with glowing mushrooms"
results = explore_seeds(prompt, num_seeds=5)
Performance
According to Alibaba AI Arena Elo human preference evaluation, Z-Image-Turbo demonstrates high competitiveness compared to other leading models and achieves state-of-the-art performance among open-source models.
Performance Comparison
| Feature | Z-Image-Turbo | Other Open-Source Models |
|---|---|---|
| Inference Steps | 8 steps | Usually 20-50 steps |
| Inference Latency | Sub-second | Several seconds |
| VRAM Requirements | 16GB | Usually 24GB+ |
| Text Rendering | Chinese & English | Usually English only |
| Instruction Following | Excellent | Average |
Gallery
Photorealistic Image Generation
Z-Image-Turbo excels at photorealistic image generation, producing high-quality photo-grade images while maintaining excellent aesthetic quality.
Bilingual Text Rendering
Z-Image-Turbo can accurately render complex Chinese and English text, a rare capability among image generation models.
Prompt Enhancement and Reasoning
Through the prompt enhancer, the model has reasoning capabilities, able to go beyond surface descriptions and tap into underlying world knowledge.
Resources
| Resource Type | Link |
|---|---|
| Official Website | Z-Image Blog |
| GitHub Repository | Tongyi-MAI/Z-Image |
| Technical Report | Z-Image Report |
| Art Gallery | Web Art Gallery |
Online Experience
Want to try Z-Image first? Access the online demos directly:
| Platform | Link |
|---|---|
| Hugging Face | Z-Image-Turbo Demo |
| ModelScope | Z-Image-Turbo Online Experience |
FAQ
Why should guidance_scale be set to 0?
Z-Image-Turbo is a distilled model that has internalized the effects of CFG (Classifier-Free Guidance), so no additional guidance is needed.
How to get better text rendering results?
Explicitly specify the text content to render in your prompt, wrap text in quotes, and describe the text style and position.
What resolutions does the model support?
1024x1024 resolution is recommended. Other common aspect ratios like 1024x768, 768x1024 are also supported.
How to reduce VRAM usage?
Use pipe.enable_model_cpu_offload() to enable CPU offload, or reduce the generation resolution.
More Articles