StableLearn Logo

Search Content

CV 8 min read

Z-Image: Alibaba's 6B Open-Source Text-to-Image Model | 8-Step Generation Rivals Closed-Source SOTA

Z-Image is Alibaba Tongyi's 6B parameter open-source text-to-image model, rivaling closed-source SOTA! 8-step inference with sub-second latency, runs on 16GB VRAM. Supports Chinese & English text rendering, leads in AI Arena human preference evaluation.

Cover image for Z-Image: Alibaba's 6B Open-Source Text-to-Image Model | 8-Step Generation Rivals Closed-Source SOTA

Published 295 days ago. Content may be outdated.

Z-Image is the latest high-efficiency image generation foundation model released by Alibaba’s Tongyi Lab. This model is truly impressive — with 6 billion parameters, it can generate high-quality images in just 8 inference steps, achieving sub-second latency on enterprise H800 GPUs, and even runs smoothly on consumer GPUs with 16GB VRAM!

Imagine simply entering a text description, and Z-Image quickly generates photorealistic images while accurately rendering Chinese and English text. This combination of efficiency and quality represents another significant breakthrough in image generation.

For example, you can ask Z-Image to generate an image of “a Chinese girl in red Hanfu, holding a folding fan, with Xi’an’s Giant Wild Goose Pagoda in the background” — it not only accurately understands these complex descriptions but also generates stunning high-quality images!

Key Highlights

  • Ultra Efficient — Only 8 inference steps, sub-second latency on H800
  • Low VRAM Requirements — Runs on consumer GPUs with 16GB VRAM
  • Photorealistic Quality — Excellent realistic image generation capability
  • Bilingual Text Rendering — Accurately renders complex Chinese and English text
  • Strong Instruction Following — Precisely understands and executes text descriptions
  • Open Source — Model weights available on Hugging Face and ModelScope

Model Architecture

S3-DiT Architecture

Z-Image adopts an innovative Scalable Single-Stream DiT (S3-DiT) architecture. Unlike traditional dual-stream approaches, S3-DiT concatenates text, visual semantic tokens, and image VAE tokens at the sequence level as a unified input stream, maximizing parameter efficiency.

This design brings several key advantages:

  • Higher Parameter Efficiency — Single-stream architecture utilizes model parameters more fully than dual-stream methods
  • Faster Inference — Unified input stream reduces computational overhead
  • Better Scalability — Architecture design supports flexible model scaling

Model Variants

Z-Image currently has three variants:

ModelDescriptionStatus
Z-Image-TurboDistilled version, 8-step inference, sub-second latency✅ Released
Z-Image-BaseNon-distilled base model, supports community fine-tuning🔜 Coming Soon
Z-Image-EditImage editing version, supports natural language editing instructions🔜 Coming Soon

Core Technologies

Decoupled-DMD Distillation Algorithm

Z-Image-Turbo’s efficient inference benefits from the Decoupled-DMD distillation algorithm. The core insight of this algorithm is that the success of existing DMD (Distribution Matching Distillation) methods comes from two independently cooperating mechanisms:

  • CFG Augmentation (CA) — The main engine driving the distillation process
  • Distribution Matching (DM) — Acts as a regularizer, ensuring output stability and quality

By decoupling these two mechanisms, the team was able to study and optimize them independently, ultimately developing an improved distillation pipeline that significantly enhances few-step generation performance.

DMDR Reinforcement Learning Integration

Building on Decoupled-DMD, the team further proposed DMDR, synergistically integrating reinforcement learning (RL) with DMD:

  • RL Unlocks DMD’s Performance Potential — Enhances the model’s generation capability
  • DMD Effectively Regularizes RL — Maintains output stability

This integration further improves semantic alignment, aesthetic quality, and structural coherence, generating images with richer high-frequency details.


Model Download

Hugging Face

ModelScope


Environment Setup

System Requirements

ItemRequirement
Operating SystemLinux / Windows / MacOS
Python3.8+
PyTorch2.0+
CUDA11.8+ (Recommended)
GPUNVIDIA GPU (16GB+ VRAM recommended)
Memory16GB+ RAM

Step 1: Install diffusers

Z-Image has been integrated into the official 🤗 diffusers library. You need to install the latest version from source:

   # Install latest diffusers from source
pip install git+https://github.com/huggingface/diffusers

Why install from source?

Z-Image support has been merged into the official diffusers repository through two PRs (#12703 and #12715). To get the latest Z-Image support, you need to install from source.

Step 2: Install Other Dependencies

   # Install PyTorch (choose based on your CUDA version)
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu118

# Optional: Install Flash Attention for better performance
pip install flash-attn --no-build-isolation

Step 3: Verify Installation

   # Test if installation is successful
import torch
from diffusers import ZImagePipeline

print(f"PyTorch version: {torch.__version__}")
print(f"CUDA available: {torch.cuda.is_available()}")
print("Z-Image installation successful!")

Quick Start

Basic Image Generation

   import torch
from diffusers import ZImagePipeline

# 1. Load model
# Use bfloat16 for best performance
print("Loading Z-Image-Turbo model...")
pipe = ZImagePipeline.from_pretrained(
    "Tongyi-MAI/Z-Image-Turbo",
    torch_dtype=torch.bfloat16,
    low_cpu_mem_usage=False,
)
pipe.to("cuda")
print("Model loaded!")

# 2. Set prompt
prompt = "Young Chinese woman in red Hanfu, intricate embroidery. Impeccable makeup, red floral forehead pattern. Elaborate high bun, golden phoenix headdress, red flowers, beads. Holds round folding fan with lady, trees, bird. Neon lightning-bolt lamp (⚡️), bright yellow glow, above extended left palm. Soft-lit outdoor night background, silhouetted tiered pagoda (Xi'an Giant Wild Goose Pagoda), blurred colorful distant lights."

# 3. Generate image
print("Generating image...")
image = pipe(
    prompt=prompt,
    height=1024,
    width=1024,
    num_inference_steps=9,  # Actually 8 DiT forward passes
    guidance_scale=0.0,     # Turbo model guidance should be 0
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]

# 4. Save result
image.save("z_image_example.png")
print("Image saved to: z_image_example.png")

Chinese Prompt Generation

Z-Image has excellent support for Chinese prompts:

   import torch
from diffusers import ZImagePipeline

# Load model
pipe = ZImagePipeline.from_pretrained(
    "Tongyi-MAI/Z-Image-Turbo",
    torch_dtype=torch.bfloat16,
)
pipe.to("cuda")

# Chinese prompt
prompt = "一位穿着精美汉服的中国古典美女,站在江南水乡的石桥上,手持油纸伞,背景是烟雨蒙蒙的白墙黑瓦建筑,柳树依依,水面倒影,画面充满诗意"

# Generate image
image = pipe(
    prompt=prompt,
    height=1024,
    width=1024,
    num_inference_steps=9,
    guidance_scale=0.0,
    generator=torch.Generator("cuda").manual_seed(123),
).images[0]

image.save("chinese_beauty.png")
print("Chinese prompt image generation complete!")

Text Rendering Example

Z-Image excels at Chinese and English text rendering:

   import torch
from diffusers import ZImagePipeline

pipe = ZImagePipeline.from_pretrained(
    "Tongyi-MAI/Z-Image-Turbo",
    torch_dtype=torch.bfloat16,
)
pipe.to("cuda")

# Prompt with text
prompt = 'A professional business card design with the text "创意设计工作室" and "Creative Design Studio" elegantly displayed, minimalist style, white background, gold accents'

image = pipe(
    prompt=prompt,
    height=1024,
    width=1024,
    num_inference_steps=9,
    guidance_scale=0.0,
    generator=torch.Generator("cuda").manual_seed(456),
).images[0]

image.save("text_rendering.png")
print("Text rendering image generation complete!")

Advanced Configuration

Enable Flash Attention

If your GPU supports Flash Attention, enable it for better efficiency:

   import torch
from diffusers import ZImagePipeline

pipe = ZImagePipeline.from_pretrained(
    "Tongyi-MAI/Z-Image-Turbo",
    torch_dtype=torch.bfloat16,
)
pipe.to("cuda")

# Enable Flash Attention 2
pipe.transformer.set_attention_backend("flash")

# Or enable Flash Attention 3 (if supported)
# pipe.transformer.set_attention_backend("_flash_3")

# Generate image
image = pipe(
    prompt="A beautiful sunset over the ocean",
    height=1024,
    width=1024,
    num_inference_steps=9,
    guidance_scale=0.0,
).images[0]

Model Compilation Acceleration

Use PyTorch compilation for further inference acceleration:

   import torch
from diffusers import ZImagePipeline

pipe = ZImagePipeline.from_pretrained(
    "Tongyi-MAI/Z-Image-Turbo",
    torch_dtype=torch.bfloat16,
)
pipe.to("cuda")

# Compile DiT model (first run will be slower, subsequent runs faster)
pipe.transformer.compile()

# Generate image
image = pipe(
    prompt="A futuristic city skyline at night",
    height=1024,
    width=1024,
    num_inference_steps=9,
    guidance_scale=0.0,
).images[0]

CPU Offload (Low VRAM Devices)

If your VRAM is limited, enable CPU offload:

   import torch
from diffusers import ZImagePipeline

pipe = ZImagePipeline.from_pretrained(
    "Tongyi-MAI/Z-Image-Turbo",
    torch_dtype=torch.bfloat16,
)

# Enable CPU offload to save VRAM
pipe.enable_model_cpu_offload()

# Generate image
image = pipe(
    prompt="A cozy coffee shop interior",
    height=1024,
    width=1024,
    num_inference_steps=9,
    guidance_scale=0.0,
).images[0]

Practical Use Cases

Batch Image Generation

   import torch
from diffusers import ZImagePipeline

def batch_generate_images(prompts, output_dir="outputs"):
    """
    Batch generate images
    """
    import os
    os.makedirs(output_dir, exist_ok=True)
    
    # Load model
    pipe = ZImagePipeline.from_pretrained(
        "Tongyi-MAI/Z-Image-Turbo",
        torch_dtype=torch.bfloat16,
    )
    pipe.to("cuda")
    
    results = []
    
    for i, prompt in enumerate(prompts):
        print(f"Generating image {i+1}/{len(prompts)}...")
        
        image = pipe(
            prompt=prompt,
            height=1024,
            width=1024,
            num_inference_steps=9,
            guidance_scale=0.0,
            generator=torch.Generator("cuda").manual_seed(i * 42),
        ).images[0]
        
        output_path = os.path.join(output_dir, f"image_{i:03d}.png")
        image.save(output_path)
        
        results.append({
            "prompt": prompt,
            "output": output_path
        })
        
        print(f"✓ Saved: {output_path}")
    
    print(f"Batch generation complete! Generated {len(results)} images")
    return results

# Usage example
prompts = [
    "A serene Japanese garden with cherry blossoms",
    "A cyberpunk street scene with neon lights",
    "A cozy library with warm lighting",
    "A majestic mountain landscape at sunrise",
]

results = batch_generate_images(prompts)

Different Size Image Generation

   import torch
from diffusers import ZImagePipeline

def generate_multiple_sizes(prompt, sizes):
    """
    Generate images at different sizes
    """
    pipe = ZImagePipeline.from_pretrained(
        "Tongyi-MAI/Z-Image-Turbo",
        torch_dtype=torch.bfloat16,
    )
    pipe.to("cuda")
    
    results = []
    
    for width, height in sizes:
        print(f"Generating {width}x{height} image...")
        
        image = pipe(
            prompt=prompt,
            height=height,
            width=width,
            num_inference_steps=9,
            guidance_scale=0.0,
            generator=torch.Generator("cuda").manual_seed(42),
        ).images[0]
        
        output_path = f"output_{width}x{height}.png"
        image.save(output_path)
        
        results.append({
            "size": (width, height),
            "output": output_path
        })
        
        print(f"✓ Saved: {output_path}")
    
    return results

# Usage example
prompt = "A beautiful landscape with mountains and lake"
sizes = [
    (1024, 1024),  # Square
    (1024, 768),   # Landscape
    (768, 1024),   # Portrait
]

results = generate_multiple_sizes(prompt, sizes)

Random Seed Exploration

   import torch
from diffusers import ZImagePipeline

def explore_seeds(prompt, num_seeds=5):
    """
    Explore effects of different random seeds
    """
    pipe = ZImagePipeline.from_pretrained(
        "Tongyi-MAI/Z-Image-Turbo",
        torch_dtype=torch.bfloat16,
    )
    pipe.to("cuda")
    
    results = []
    
    for seed in range(num_seeds):
        print(f"Generating with seed {seed}...")
        
        image = pipe(
            prompt=prompt,
            height=1024,
            width=1024,
            num_inference_steps=9,
            guidance_scale=0.0,
            generator=torch.Generator("cuda").manual_seed(seed),
        ).images[0]
        
        output_path = f"seed_{seed:03d}.png"
        image.save(output_path)
        
        results.append({
            "seed": seed,
            "output": output_path
        })
        
        print(f"✓ Seed {seed} complete")
    
    print(f"Exploration complete! Generated {len(results)} images with different seeds")
    return results

# Usage example
prompt = "A magical forest with glowing mushrooms"
results = explore_seeds(prompt, num_seeds=5)

Performance

According to Alibaba AI Arena Elo human preference evaluation, Z-Image-Turbo demonstrates high competitiveness compared to other leading models and achieves state-of-the-art performance among open-source models.

Performance Comparison

FeatureZ-Image-TurboOther Open-Source Models
Inference Steps8 stepsUsually 20-50 steps
Inference LatencySub-secondSeveral seconds
VRAM Requirements16GBUsually 24GB+
Text RenderingChinese & EnglishUsually English only
Instruction FollowingExcellentAverage

Photorealistic Image Generation

Z-Image-Turbo excels at photorealistic image generation, producing high-quality photo-grade images while maintaining excellent aesthetic quality.

Bilingual Text Rendering

Z-Image-Turbo can accurately render complex Chinese and English text, a rare capability among image generation models.

Prompt Enhancement and Reasoning

Through the prompt enhancer, the model has reasoning capabilities, able to go beyond surface descriptions and tap into underlying world knowledge.


Resources

Resource TypeLink
Official WebsiteZ-Image Blog
GitHub RepositoryTongyi-MAI/Z-Image
Technical ReportZ-Image Report
Art GalleryWeb Art Gallery

Online Experience

Want to try Z-Image first? Access the online demos directly:

PlatformLink
Hugging FaceZ-Image-Turbo Demo
ModelScopeZ-Image-Turbo Online Experience

FAQ

Why should guidance_scale be set to 0?

Z-Image-Turbo is a distilled model that has internalized the effects of CFG (Classifier-Free Guidance), so no additional guidance is needed.

How to get better text rendering results?

Explicitly specify the text content to render in your prompt, wrap text in quotes, and describe the text style and position.

What resolutions does the model support?

1024x1024 resolution is recommended. Other common aspect ratios like 1024x768, 768x1024 are also supported.

How to reduce VRAM usage?

Use pipe.enable_model_cpu_offload() to enable CPU offload, or reduce the generation resolution.

Share Article

More Articles