StableLearn Logo

Search Content

CV 6 min read

Baidu ERNIE-Image: 8B Open-Source Text-to-Image AI Beats Larger Models

Baidu ERNIE-Image: 8B parameter open-source text-to-image AI model achieves SOTA performance. GenEval 0.8856, LongTextBench 0.96+. Superior text rendering, runs on 24GB VRAM, 8-step Turbo mode.

Cover image for Baidu ERNIE-Image: 8B Open-Source Text-to-Image AI Beats Larger Models

Published 167 days ago. Content may be outdated.

Baidu has quietly achieved something remarkable.

Today, Baidu’s ERNIE-Image team officially open-sourced the text-to-image model ERNIE-Image, which has only 8B parameters but achieves performance in multiple authoritative benchmarks that matches or even exceeds larger open-source models.

Even more impressive, it excels at text rendering—a notoriously difficult challenge—whether it’s posters, infographics, UI interfaces, or long-form text layouts, it generates them with precision.

Here are the key numbers:

  • GenEval overall score 0.8856, best performance in single object generation and attribute binding
  • LongTextBench scores above 0.96 for both English and Chinese, crushing competitors in text rendering
  • 24GB VRAM to run, deployable on consumer-grade GPUs
  • 8-step inference (Turbo version) generates high-quality images
  • Open-sourced on Hugging Face, supports Diffusers, SGLang, and other mainstream frameworks

What does this mean? A compact 8B model can now compete with models having tens of billions of parameters in critical capabilities like text rendering, instruction following, and structured generation.

Why Can 8B Parameters Beat Larger Models?

Core Architecture: Single-Stream DiT + Prompt Enhancer

ERNIE-Image’s core is a single-stream Diffusion Transformer (DiT) architecture, paired with a lightweight Prompt Enhancer.

ComponentFunctionFeatures
Single-Stream DiTImage generation backboneOnly 8B parameters, efficient and compact
Prompt EnhancerPrompt expansionExpands brief inputs into structured descriptions

The brilliance of this design: using a lightweight Prompt Enhancer to compensate for the small model’s limitations in understanding complex prompts, allowing the 8B DiT to focus on image generation itself.

The result: Small model + Smart enhancement = Large model-level performance.

Dual-Version Strategy: Quality vs Speed

Baidu released two versions to meet different scenario needs:

VersionOptimization FocusInference StepsCFGUse Cases
ERNIE-ImageGeneral capability + Instruction fidelity50 steps4.0High-quality generation, detail-demanding
ERNIE-Image-TurboSpeed + Aesthetics8 steps1.0Fast generation, real-time applications

The Turbo version, optimized through DMD (Diffusion Model Distillation) and reinforcement learning, compresses inference steps from 50 to 8, achieving 6x+ speed improvement while maintaining high-quality output.

Benchmark Results: Comprehensive Dominance

GenEval: Strongest Instruction Following

GenEval is the authoritative benchmark for evaluating text-to-image models’ instruction following capabilities, covering single objects, multiple objects, counting, colors, positions, attribute binding, and more.

ModelSingle ObjectTwo ObjectsCountingColorsPositionAttribute BindingOverall
ERNIE-Image (w/o PE)1.00000.95960.77810.92820.85500.79250.8856
ERNIE-Image (w/ PE)0.99060.95960.81870.88300.86250.72250.8728
Qwen-Image0.99000.92000.89000.88000.76000.77000.8683
FLUX.2-klein-9B0.93130.95710.82810.91490.71750.74000.8481
Z-Image1.00000.94000.78000.93000.62000.77000.8400

Key Findings:

  • Perfect single object generation, demonstrating solid foundational capabilities
  • Attribute binding 0.7925, strongest instruction understanding in complex scenarios
  • Overall score 0.8856, comprehensively leading similar open-source models

LongTextBench: Text Rendering Champion

LongTextBench specifically tests models’ rendering capabilities in long-text, dense-text scenarios—a recognized challenge for text-to-image models.

ModelEnglishChineseAverage
Seedream 4.50.98900.98730.9882
ERNIE-Image (w/ PE)0.98040.96610.9733
GLM-Image0.95240.97880.9656
ERNIE-Image-Turbo (w/ PE)0.96750.96360.9655
Qwen-Image0.94300.94600.9445
Z-Image0.93500.93600.9355
FLUX.2-klein-9B0.86420.21830.5413

Key Findings:

  • Both English and Chinese above 0.96, text rendering capability second only to closed-source Seedream 4.5
  • FLUX.2 collapses in Chinese scenarios (0.2183), while ERNIE-Image remains stable
  • Turbo version maintains high score of 0.9655, achieving both speed and quality

OneIG: Comprehensive Capability Assessment

OneIG is a comprehensive benchmark evaluating alignment, text, reasoning, style, diversity, and other dimensions.

OneIG-EN (English Scenarios)

ModelAlignmentTextReasoningStyleDiversityOverall
Nano Banana 2.00.88800.94400.33400.48100.24500.5780
Seedream 4.50.89100.99800.35000.43400.20700.5760
ERNIE-Image (w/ PE)0.86780.97880.35660.43090.24110.5750
ERNIE-Image-Turbo (w/ PE)0.86760.96660.35370.41910.22120.5656

OneIG-ZH (Chinese Scenarios)

ModelAlignmentTextReasoningStyleDiversityOverall
Nano Banana 2.00.84300.98300.31100.46100.23600.5670
ERNIE-Image (w/ PE)0.82990.95390.30560.43420.24780.5543
Qwen-Image0.82500.96300.26700.40500.27900.5480
ERNIE-Image-Turbo (w/ PE)0.82580.93860.30430.42080.22810.5435

Key Findings:

  • Strongest reasoning capability (0.3566), can understand complex logical relationships
  • Balanced English and Chinese performance, no obvious weaknesses
  • Turbo version minimal performance loss, highly practical

Five Core Advantages

1. Text Rendering: Best Choice for Posters, Infographics, UI Interfaces

ERNIE-Image excels in dense, long-form, layout-sensitive text generation tasks, particularly suitable for:

  • Poster design: Multi-line text, complex layouts
  • Infographics: Data visualization + text descriptions
  • UI interfaces: Buttons, labels, menus, and other text elements
  • Comic panels: Dialog boxes, narration, sound effect text

2. Instruction Following: Understanding Complex Prompts

Can accurately understand prompts containing multiple objects, detailed relationships, knowledge-intensive descriptions, such as:

  • “A girl in a red dress standing next to a blue car, holding a yellow balloon"
  • "A screenshot of a desktop browser webpage interface, with a standard browser frame at the top…“

3. Structured Generation: Multi-Panel, Storyboards, Grid Layouts

Excels in structured visual tasks:

  • Posters: Title + body + decorative elements
  • Comics: Multi-panel storyboards + dialog boxes
  • Sticker packs: Grid layout + text annotations
  • Infographics: Charts + data + explanatory text

4. Style Coverage: Realistic, Design-Oriented, Stylized

Supports multiple visual styles:

  • Realistic photography: Natural lighting, authentic textures
  • Design-oriented: Flat illustrations, vector styles
  • Stylized aesthetics: Anime, hand-drawn, retro, etc.

5. Low Deployment Barrier: 24GB VRAM Sufficient

ConfigurationDescription
VRAM24GB (RTX 3090 / 4090 level)
Inference SpeedStandard version 50 steps, Turbo version 8 steps
Deployment MethodsDiffusers, SGLang, ComfyUI

Runs on consumer-grade GPUs, no need for expensive enterprise hardware.

Prompt Enhancer: The Small Model’s “Power-Up”

How It Works

Prompt Enhancer is a lightweight language model that expands brief user inputs into richer structured descriptions.

InputOutput
”A black and white Chinese rural dog""A black and white Chinese rural dog with distinct coloring, black ears and back, white chest and legs, alert eyes, standing on grass, sunlight shining on it…”

Two Deployment Methods

Method 1: Built-in PE (Simple and Fast)

   pipe = ErnieImagePipeline.from_pretrained(
    "baidu/ERNIE-Image",
    torch_dtype=torch.bfloat16,
).to("cuda")

image = pipe(
    prompt="A black and white Chinese rural dog",
    use_pe=True  # Enable Prompt Enhancer
).images[0]

Method 2: Separate PE Deployment (Performance Optimization)

Deploy PE and DiT separately, PE using vLLM acceleration, DiT using SGLang inference, suitable for high-concurrency scenarios.

   # Deploy PE (vLLM)
vllm serve ./ernie_image_pe --port 8888

# Deploy DiT (SGLang)
sglang serve --model-path baidu/ERNIE-Image-Turbo

Performance Comparison

ConfigurationGenEvalOneIG-ENLongTextBench
w/ PE0.87280.57500.9733
w/o PE0.88560.55370.9636

Key Findings:

  • Stronger reasoning with PE (OneIG-EN: 0.5750 vs 0.5537)
  • Stronger basic generation without PE (GenEval: 0.8856 vs 0.8728)
  • Users can choose whether to enable PE based on scenarios

Open Source Ecosystem: Ready to Use

Supported Frameworks and Tools

ToolDescriptionLink
DiffusersHugging Face official inference libraryDocs
SGLangHigh-performance inference serviceDocs
ComfyUIVisual workflowTemplate
UnslothGGUF quantization supportDocs
AI-ToolkitFine-tuning toolsDocs

Quick Start

   import torch
from diffusers import ErnieImagePipeline

pipe = ErnieImagePipeline.from_pretrained(
    "baidu/ERNIE-Image-Turbo",  # or "baidu/ERNIE-Image"
    torch_dtype=torch.bfloat16,
).to("cuda")

image = pipe(
    prompt="A black and white Chinese rural dog",
    height=1024,
    width=1024,
    num_inference_steps=8,  # Turbo version uses 8 steps
    guidance_scale=1.0,     # Turbo version uses 1.0
    use_pe=True
).images[0]

image.save("output.png")
   # Start service
sglang serve --model-path baidu/ERNIE-Image-Turbo

# Send request
curl -X POST http://localhost:30000/generate \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "A black and white Chinese rural dog",
    "height": 1024,
    "width": 1024,
    "num_inference_steps": 8,
    "guidance_scale": 1.0
  }' \
  --output output.png

Use Cases: Who Can Use It? How?

ScenarioTarget UsersCore Value
Content CreationMedia creators, designersQuickly generate posters, illustrations, stickers
E-commerce DesignE-commerce operators, graphic designersProduct posters, promotional images, detail page graphics
UI DesignProduct managers, UI designersRapid prototyping, interface mockups
Education & PublishingTeachers, publishersTextbook illustrations, infographics
Game DevelopmentIndie game developersConcept art, scene design, character concepts
Video ProductionVideo creatorsCover images, subtitle cards, transition animations

Competitor Comparison: 8B Beats 9B+

ModelParametersGenEvalLongTextBenchVRAMOpen Source
ERNIE-Image8B0.88560.973324GB✅
FLUX.2-klein9B0.84810.541324GB+✅
Qwen-ImageUnknown0.86830.9445Unknown✅
Z-ImageUnknown0.84000.9355Unknown✅
Seedream 4.5Unknown0.57600.9882Unknown❌

Key Findings:

  • 8B parameters achieve 9B+ model performance
  • Text rendering capability second only to closed-source Seedream 4.5
  • Comprehensive performance comprehensively leads similar open-source models

Conclusion: Victory of Small and Beautiful

The release of ERNIE-Image proves an important fact: bigger models aren’t always better—architecture design and training strategies matter equally.

8B parameters + smart Prompt Enhancer can match or exceed larger models’ performance, which is an important revelation for the entire industry.

More importantly, the 24GB VRAM deployment threshold allows more developers and creators to use high-quality text-to-image models, rather than being shut out by expensive hardware costs.

Baidu has quietly achieved something remarkable—not only open-sourcing the model but also providing complete toolchain support, from Diffusers to SGLang, from ComfyUI to fine-tuning tools. Developers can get started immediately, no waiting required.

The era of “small and beautiful” AI text-to-image generation has arrived.


Related Links:

Share Article

More Articles