Baidu ERNIE-Image: 8B Open-Source Text-to-Image AI Beats Larger Models
Baidu ERNIE-Image: 8B parameter open-source text-to-image AI model achieves SOTA performance. GenEval 0.8856, LongTextBench 0.96+. Superior text rendering, runs on 24GB VRAM, 8-step Turbo mode.
Published 167 days ago. Content may be outdated.
Baidu has quietly achieved something remarkable.
Today, Baidu’s ERNIE-Image team officially open-sourced the text-to-image model ERNIE-Image, which has only 8B parameters but achieves performance in multiple authoritative benchmarks that matches or even exceeds larger open-source models.
Even more impressive, it excels at text rendering—a notoriously difficult challenge—whether it’s posters, infographics, UI interfaces, or long-form text layouts, it generates them with precision.
Here are the key numbers:
- GenEval overall score 0.8856, best performance in single object generation and attribute binding
- LongTextBench scores above 0.96 for both English and Chinese, crushing competitors in text rendering
- 24GB VRAM to run, deployable on consumer-grade GPUs
- 8-step inference (Turbo version) generates high-quality images
- Open-sourced on Hugging Face, supports Diffusers, SGLang, and other mainstream frameworks
What does this mean? A compact 8B model can now compete with models having tens of billions of parameters in critical capabilities like text rendering, instruction following, and structured generation.
Why Can 8B Parameters Beat Larger Models?
Core Architecture: Single-Stream DiT + Prompt Enhancer
ERNIE-Image’s core is a single-stream Diffusion Transformer (DiT) architecture, paired with a lightweight Prompt Enhancer.
| Component | Function | Features |
|---|---|---|
| Single-Stream DiT | Image generation backbone | Only 8B parameters, efficient and compact |
| Prompt Enhancer | Prompt expansion | Expands brief inputs into structured descriptions |
The brilliance of this design: using a lightweight Prompt Enhancer to compensate for the small model’s limitations in understanding complex prompts, allowing the 8B DiT to focus on image generation itself.
The result: Small model + Smart enhancement = Large model-level performance.
Dual-Version Strategy: Quality vs Speed
Baidu released two versions to meet different scenario needs:
| Version | Optimization Focus | Inference Steps | CFG | Use Cases |
|---|---|---|---|---|
| ERNIE-Image | General capability + Instruction fidelity | 50 steps | 4.0 | High-quality generation, detail-demanding |
| ERNIE-Image-Turbo | Speed + Aesthetics | 8 steps | 1.0 | Fast generation, real-time applications |
The Turbo version, optimized through DMD (Diffusion Model Distillation) and reinforcement learning, compresses inference steps from 50 to 8, achieving 6x+ speed improvement while maintaining high-quality output.
Benchmark Results: Comprehensive Dominance
GenEval: Strongest Instruction Following
GenEval is the authoritative benchmark for evaluating text-to-image models’ instruction following capabilities, covering single objects, multiple objects, counting, colors, positions, attribute binding, and more.
| Model | Single Object | Two Objects | Counting | Colors | Position | Attribute Binding | Overall |
|---|---|---|---|---|---|---|---|
| ERNIE-Image (w/o PE) | 1.0000 | 0.9596 | 0.7781 | 0.9282 | 0.8550 | 0.7925 | 0.8856 |
| ERNIE-Image (w/ PE) | 0.9906 | 0.9596 | 0.8187 | 0.8830 | 0.8625 | 0.7225 | 0.8728 |
| Qwen-Image | 0.9900 | 0.9200 | 0.8900 | 0.8800 | 0.7600 | 0.7700 | 0.8683 |
| FLUX.2-klein-9B | 0.9313 | 0.9571 | 0.8281 | 0.9149 | 0.7175 | 0.7400 | 0.8481 |
| Z-Image | 1.0000 | 0.9400 | 0.7800 | 0.9300 | 0.6200 | 0.7700 | 0.8400 |
Key Findings:
- Perfect single object generation, demonstrating solid foundational capabilities
- Attribute binding 0.7925, strongest instruction understanding in complex scenarios
- Overall score 0.8856, comprehensively leading similar open-source models
LongTextBench: Text Rendering Champion
LongTextBench specifically tests models’ rendering capabilities in long-text, dense-text scenarios—a recognized challenge for text-to-image models.
| Model | English | Chinese | Average |
|---|---|---|---|
| Seedream 4.5 | 0.9890 | 0.9873 | 0.9882 |
| ERNIE-Image (w/ PE) | 0.9804 | 0.9661 | 0.9733 |
| GLM-Image | 0.9524 | 0.9788 | 0.9656 |
| ERNIE-Image-Turbo (w/ PE) | 0.9675 | 0.9636 | 0.9655 |
| Qwen-Image | 0.9430 | 0.9460 | 0.9445 |
| Z-Image | 0.9350 | 0.9360 | 0.9355 |
| FLUX.2-klein-9B | 0.8642 | 0.2183 | 0.5413 |
Key Findings:
- Both English and Chinese above 0.96, text rendering capability second only to closed-source Seedream 4.5
- FLUX.2 collapses in Chinese scenarios (0.2183), while ERNIE-Image remains stable
- Turbo version maintains high score of 0.9655, achieving both speed and quality
OneIG: Comprehensive Capability Assessment
OneIG is a comprehensive benchmark evaluating alignment, text, reasoning, style, diversity, and other dimensions.
OneIG-EN (English Scenarios)
| Model | Alignment | Text | Reasoning | Style | Diversity | Overall |
|---|---|---|---|---|---|---|
| Nano Banana 2.0 | 0.8880 | 0.9440 | 0.3340 | 0.4810 | 0.2450 | 0.5780 |
| Seedream 4.5 | 0.8910 | 0.9980 | 0.3500 | 0.4340 | 0.2070 | 0.5760 |
| ERNIE-Image (w/ PE) | 0.8678 | 0.9788 | 0.3566 | 0.4309 | 0.2411 | 0.5750 |
| ERNIE-Image-Turbo (w/ PE) | 0.8676 | 0.9666 | 0.3537 | 0.4191 | 0.2212 | 0.5656 |
OneIG-ZH (Chinese Scenarios)
| Model | Alignment | Text | Reasoning | Style | Diversity | Overall |
|---|---|---|---|---|---|---|
| Nano Banana 2.0 | 0.8430 | 0.9830 | 0.3110 | 0.4610 | 0.2360 | 0.5670 |
| ERNIE-Image (w/ PE) | 0.8299 | 0.9539 | 0.3056 | 0.4342 | 0.2478 | 0.5543 |
| Qwen-Image | 0.8250 | 0.9630 | 0.2670 | 0.4050 | 0.2790 | 0.5480 |
| ERNIE-Image-Turbo (w/ PE) | 0.8258 | 0.9386 | 0.3043 | 0.4208 | 0.2281 | 0.5435 |
Key Findings:
- Strongest reasoning capability (0.3566), can understand complex logical relationships
- Balanced English and Chinese performance, no obvious weaknesses
- Turbo version minimal performance loss, highly practical
Five Core Advantages
1. Text Rendering: Best Choice for Posters, Infographics, UI Interfaces
ERNIE-Image excels in dense, long-form, layout-sensitive text generation tasks, particularly suitable for:
- Poster design: Multi-line text, complex layouts
- Infographics: Data visualization + text descriptions
- UI interfaces: Buttons, labels, menus, and other text elements
- Comic panels: Dialog boxes, narration, sound effect text
2. Instruction Following: Understanding Complex Prompts
Can accurately understand prompts containing multiple objects, detailed relationships, knowledge-intensive descriptions, such as:
- “A girl in a red dress standing next to a blue car, holding a yellow balloon"
- "A screenshot of a desktop browser webpage interface, with a standard browser frame at the top…“
3. Structured Generation: Multi-Panel, Storyboards, Grid Layouts
Excels in structured visual tasks:
- Posters: Title + body + decorative elements
- Comics: Multi-panel storyboards + dialog boxes
- Sticker packs: Grid layout + text annotations
- Infographics: Charts + data + explanatory text
4. Style Coverage: Realistic, Design-Oriented, Stylized
Supports multiple visual styles:
- Realistic photography: Natural lighting, authentic textures
- Design-oriented: Flat illustrations, vector styles
- Stylized aesthetics: Anime, hand-drawn, retro, etc.
5. Low Deployment Barrier: 24GB VRAM Sufficient
| Configuration | Description |
|---|---|
| VRAM | 24GB (RTX 3090 / 4090 level) |
| Inference Speed | Standard version 50 steps, Turbo version 8 steps |
| Deployment Methods | Diffusers, SGLang, ComfyUI |
Runs on consumer-grade GPUs, no need for expensive enterprise hardware.
Prompt Enhancer: The Small Model’s “Power-Up”
How It Works
Prompt Enhancer is a lightweight language model that expands brief user inputs into richer structured descriptions.
| Input | Output |
|---|---|
| ”A black and white Chinese rural dog" | "A black and white Chinese rural dog with distinct coloring, black ears and back, white chest and legs, alert eyes, standing on grass, sunlight shining on it…” |
Two Deployment Methods
Method 1: Built-in PE (Simple and Fast)
pipe = ErnieImagePipeline.from_pretrained(
"baidu/ERNIE-Image",
torch_dtype=torch.bfloat16,
).to("cuda")
image = pipe(
prompt="A black and white Chinese rural dog",
use_pe=True # Enable Prompt Enhancer
).images[0]
Method 2: Separate PE Deployment (Performance Optimization)
Deploy PE and DiT separately, PE using vLLM acceleration, DiT using SGLang inference, suitable for high-concurrency scenarios.
# Deploy PE (vLLM)
vllm serve ./ernie_image_pe --port 8888
# Deploy DiT (SGLang)
sglang serve --model-path baidu/ERNIE-Image-Turbo
Performance Comparison
| Configuration | GenEval | OneIG-EN | LongTextBench |
|---|---|---|---|
| w/ PE | 0.8728 | 0.5750 | 0.9733 |
| w/o PE | 0.8856 | 0.5537 | 0.9636 |
Key Findings:
- Stronger reasoning with PE (OneIG-EN: 0.5750 vs 0.5537)
- Stronger basic generation without PE (GenEval: 0.8856 vs 0.8728)
- Users can choose whether to enable PE based on scenarios
Open Source Ecosystem: Ready to Use
Supported Frameworks and Tools
| Tool | Description | Link |
|---|---|---|
| Diffusers | Hugging Face official inference library | Docs |
| SGLang | High-performance inference service | Docs |
| ComfyUI | Visual workflow | Template |
| Unsloth | GGUF quantization support | Docs |
| AI-Toolkit | Fine-tuning tools | Docs |
Quick Start
Diffusers (Recommended for Beginners)
import torch
from diffusers import ErnieImagePipeline
pipe = ErnieImagePipeline.from_pretrained(
"baidu/ERNIE-Image-Turbo", # or "baidu/ERNIE-Image"
torch_dtype=torch.bfloat16,
).to("cuda")
image = pipe(
prompt="A black and white Chinese rural dog",
height=1024,
width=1024,
num_inference_steps=8, # Turbo version uses 8 steps
guidance_scale=1.0, # Turbo version uses 1.0
use_pe=True
).images[0]
image.save("output.png")
SGLang (Recommended for Production)
# Start service
sglang serve --model-path baidu/ERNIE-Image-Turbo
# Send request
curl -X POST http://localhost:30000/generate \
-H "Content-Type: application/json" \
-d '{
"prompt": "A black and white Chinese rural dog",
"height": 1024,
"width": 1024,
"num_inference_steps": 8,
"guidance_scale": 1.0
}' \
--output output.png
Use Cases: Who Can Use It? How?
| Scenario | Target Users | Core Value |
|---|---|---|
| Content Creation | Media creators, designers | Quickly generate posters, illustrations, stickers |
| E-commerce Design | E-commerce operators, graphic designers | Product posters, promotional images, detail page graphics |
| UI Design | Product managers, UI designers | Rapid prototyping, interface mockups |
| Education & Publishing | Teachers, publishers | Textbook illustrations, infographics |
| Game Development | Indie game developers | Concept art, scene design, character concepts |
| Video Production | Video creators | Cover images, subtitle cards, transition animations |
Competitor Comparison: 8B Beats 9B+
| Model | Parameters | GenEval | LongTextBench | VRAM | Open Source |
|---|---|---|---|---|---|
| ERNIE-Image | 8B | 0.8856 | 0.9733 | 24GB | ✅ |
| FLUX.2-klein | 9B | 0.8481 | 0.5413 | 24GB+ | ✅ |
| Qwen-Image | Unknown | 0.8683 | 0.9445 | Unknown | ✅ |
| Z-Image | Unknown | 0.8400 | 0.9355 | Unknown | ✅ |
| Seedream 4.5 | Unknown | 0.5760 | 0.9882 | Unknown | ❌ |
Key Findings:
- 8B parameters achieve 9B+ model performance
- Text rendering capability second only to closed-source Seedream 4.5
- Comprehensive performance comprehensively leads similar open-source models
Conclusion: Victory of Small and Beautiful
The release of ERNIE-Image proves an important fact: bigger models aren’t always better—architecture design and training strategies matter equally.
8B parameters + smart Prompt Enhancer can match or exceed larger models’ performance, which is an important revelation for the entire industry.
More importantly, the 24GB VRAM deployment threshold allows more developers and creators to use high-quality text-to-image models, rather than being shut out by expensive hardware costs.
Baidu has quietly achieved something remarkable—not only open-sourcing the model but also providing complete toolchain support, from Diffusers to SGLang, from ComfyUI to fine-tuning tools. Developers can get started immediately, no waiting required.
The era of “small and beautiful” AI text-to-image generation has arrived.
Related Links:
- GitHub Repository: https://github.com/baidu/ERNIE-Image
- Hugging Face Model: https://huggingface.co/Baidu/ERNIE-Image
- Hugging Face Turbo Model: https://huggingface.co/Baidu/ERNIE-Image-Turbo
- Online Demo: https://huggingface.co/spaces/baidu/ERNIE-Image
- Official Blog: https://yiyan.baidu.com/blog/posts/ernie-image
More Articles