Gemma 4 Release: Google's 4-Size Open Source Family with MoE Architecture
Google releases Gemma 4 multimodal models with dense and MoE architectures, 256K token context, native reasoning mode, and deployment from mobile to server. Four model sizes available under Apache 2.0.
Published 179 days ago. Content may be outdated.
Google DeepMind just dropped a bombshell! On April 2nd, the Gemma 4 series officially launched—a complete “family pack” ranging from 2.3B to 30.7B parameters, featuring both dense and MoE (Mixture of Experts) architectures, context windows up to 256K tokens, and native reasoning mode. Even better? Full Apache 2.0 open source, deployable from smartphones to servers.
Let’s look at the numbers: 4 model sizes, 256K token context, 140+ languages, native reasoning mode. What does this mean? It can read an entire technical document in one go and perform reasoning analysis on it.
Model Family: From Mobile to Server, Full Coverage
Gemma 4 isn’t flying solo—it launches with 4 models covering everything from mobile devices to data centers.
Dense Models: Three Brothers, Each with Strengths
| Model | Parameters | Context Length | Modalities | Target |
|---|---|---|---|---|
| E2B | 2.3B effective (5.1B with embeddings) | 128k tokens | Text, Image, Audio | Mobile/Edge |
| E4B | 4.5B effective (8.0B with embeddings) | 128k tokens | Text, Image, Audio | Laptop/Edge |
| 31B Dense | 30.7B | 256k tokens | Text, Image | Workstation/Server |
The “E” in E2B and E4B stands for “effective parameters.” These smaller models use Per-Layer Embeddings (PLE) technology—simply put, each decoder layer has its own small embedding table. The embedding tables are large but only used for fast lookups, so the effective parameter count is much smaller than the total. This design is optimized for on-device deployment, enabling efficient operation on phones and laptops.
MoE Model: Perfect Balance of Speed and Capability
| Model | Total Parameters | Active Parameters | Context Length | Experts | Modalities |
|---|---|---|---|---|---|
| 26B A4B | 25.2B | 3.8B | 256k tokens | 8 active / 128 total + 1 shared | Text, Image |
The “A” in 26B A4B stands for “active parameters.” Through the MoE architecture, only a 4B parameter subset is activated during inference, running almost as fast as a 4B model while delivering performance close to a 26B model. This is an excellent choice for fast inference.
The key is Hybrid Attention mechanism: interweaving local sliding window attention with global attention, ensuring the final layer is always global. This design provides the processing speed and low memory footprint of lightweight models without sacrificing the deep awareness needed for complex long-context tasks.
Performance Explosion: SOTA Across Multiple Benchmarks
The data doesn’t lie. Gemma 4 achieves SOTA across multiple benchmarks.
Core Capability Comparison
| Benchmark | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B |
|---|---|---|---|---|---|
| MMLU Pro | 85.2% | 82.6% | 69.4% | 60.0% | 67.6% |
| AIME 2026 No Tools | 89.2% | 88.3% | 42.5% | 37.5% | 20.8% |
| LiveCodeBench v6 | 80.0% | 77.1% | 52.0% | 44.0% | 29.1% |
| Codeforces ELO | 2150 | 1718 | 940 | 633 | 110 |
| GPQA Diamond | 84.3% | 82.3% | 58.6% | 43.4% | 42.4% |
Look at these numbers: 31B scores 89.2% on AIME 2026 (math competition), 80.0% on LiveCodeBench (coding ability), and reaches 2150 Codeforces ELO—what level is this? It already surpasses most human programmers.
Even more impressive, the smallest E2B achieves 60.0% on MMLU Pro, only 7.6 percentage points behind Gemma 3 27B’s 67.6%, with less than one-tenth the parameters.
Multimodal Capability: Comprehensive Vision Understanding
| Capability | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 E4B | Gemma 4 E2B |
|---|---|---|---|---|
| MMMU Pro | 76.9% | 73.8% | 52.6% | 44.2% |
| MATH-Vision | 85.6% | 82.4% | 59.5% | 52.4% |
| OmniDocBench 1.5 (lower is better) | 0.131 | 0.149 | 0.181 | 0.290 |
For vision understanding, 31B scores 85.6% on MATH-Vision and achieves an edit distance of only 0.131 on document parsing (OmniDocBench)—meaning it can accurately recognize and understand document content with maxed-out OCR capability.
Audio Capability: E2B/E4B Exclusive
| Capability | Gemma 4 E4B | Gemma 4 E2B |
|---|---|---|
| CoVoST (Speech Translation) | 35.54 | 33.47 |
| FLEURS (Speech Recognition, lower is better) | 0.08 | 0.09 |
E2B and E4B natively support audio input with excellent performance in speech recognition and translation tasks.
Long Context: Truly “Reading Entire Books”
| Context Length | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 E4B | Gemma 4 E2B |
|---|---|---|---|---|
| MRCR v2 8-needle 128k (avg) | 66.4% | 44.1% | 25.4% | 19.1% |
31B scores 66.4% on 128K token long-context tasks. What does this mean? It can read a medium-length technical book in one go and answer questions about it.
Core Features: Not Just Multimodal, But Reasoning Mode
Gemma 4 isn’t just a simple multimodal model—it comes with a bunch of “black tech.”
Feature Overview
| Feature | Supported Models | Description |
|---|---|---|
| Thinking Mode | All models | Built-in reasoning mode allowing step-by-step thinking before answering |
| Long Context | All models | E2B/E4B support 128K tokens, 26B A4B/31B support 256K tokens |
| Image Understanding | All models | Object detection, document/PDF parsing, OCR, handwriting recognition, variable aspect ratios and resolutions |
| Video Understanding | All models | Analyze video by processing frame sequences |
| Audio Processing | E2B/E4B | Automatic Speech Recognition (ASR) and speech translation (multiple languages) |
| Function Calling | All models | Native support for structured tool use, enabling agent workflows |
| Multilingual | All models | Out-of-box support for 35+ languages, pretrained on 140+ languages |
Black Tech: Thinking Mode
The most eye-catching feature is Thinking Mode.
Simply put, the model “thinks” before answering. By including the {"<|think|>"} token in the system prompt, the model outputs its internal reasoning process before giving the final answer.
This is Gemma 4’s unique “emergent ability,” enabling it to handle more complex reasoning tasks. Performance on AIME 2026 (math competition) and Codeforces (programming competition) is the best proof.
Variable Image Resolution: Balance Between Detail and Speed
Gemma 4 supports controlling image resolution through configurable vision token budgets. Supported token budgets: 70, 140, 280, 560, 1120.
- Low budget (70/140): Suitable for classification, captioning, or video understanding tasks—fast
- High budget (560/1120): Suitable for OCR, document parsing, or reading small text—detailed
This design allows the model to flexibly adjust based on task requirements, quickly processing large numbers of frames while accurately recognizing details.
How to Use? Transformers and llama.cpp Both Supported
Using Gemma 4 is very simple—both Transformers and llama.cpp are already supported.
Quick Start
import torch
from transformers import AutoProcessor, AutoModelForCausalLM
MODEL_ID = "google/gemma-4-E2B-it"
# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
dtype=torch.bfloat16,
device_map="auto"
)
# Construct dialogue
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Write a short joke about saving RAM."},
]
# Process input
text = processor.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False # Set to True to enable reasoning mode
)
inputs = processor(text=text, return_tensors="pt").to(model.device)
input_len = inputs["input_ids"].shape[-1]
# Generate output
outputs = model.generate(**inputs, max_new_tokens=1024)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
# Parse thinking output
processor.parse_response(response)
Best Practices
| Configuration | Recommended Value | Description |
|---|---|---|
| temperature | 1.0 | Standardized sampling configuration |
| top_p | 0.95 | Standardized sampling configuration |
| top_k | 64 | Standardized sampling configuration |
| Thinking Mode | Enable as needed | Include think token in system prompt |
| Image Token Budget | 70-1120 | Choose based on task requirements |
Technical Deep Dive: Hybrid Attention + MoE
Core Architecture
| Component | Function | Purpose |
|---|---|---|
| Hybrid Attention | Local sliding window + global attention | Balance speed and long-context understanding |
| Per-Layer Embeddings | Independent embedding table per layer | Improve on-device parameter efficiency |
| Mixture of Experts (MoE) | 8 active / 128 total experts | Only activate subset during inference for speed |
| Unified KV + p-RoPE | Global layer optimization | Reduce long-context memory usage |
The key is Hybrid Attention mechanism: interweaving local sliding window attention with global attention, ensuring the final layer is always global. This design provides the processing speed and low memory footprint of lightweight models without sacrificing the deep awareness needed for complex long-context tasks.
For MoE models, by activating only a 4B parameter subset during inference, 26B A4B runs almost as fast as a 4B model while delivering performance close to a 26B model.
Application Scenarios: From Content Creation to Agents
| Application | Core Capability | Real Value |
|---|---|---|
| Content Creation | Text generation, code generation | Poetry, scripts, marketing copy, code completion |
| Chatbots | Conversational AI | Customer service, virtual assistants, interactive apps |
| Document Processing | OCR, document parsing | Extract, interpret, and summarize visual data |
| Audio Processing | ASR, speech translation | Meeting transcription, multilingual translation |
| Agent Workflows | Function calling | Structured tool use, autonomous agents |
| Research & Education | NLP research, language learning | Algorithm development, grammar correction, writing practice |
Safety and Ethics: Google’s AI Principles
Gemma 4 is developed by Google DeepMind and, like the proprietary Gemini models, undergoes rigorous safety evaluation.
Evaluation Methods
- CSAM Filtering: Strict filtering applied at multiple stages of data preparation
- Sensitive Data Filtering: Automated techniques filter personal information and other sensitive data
- Content Safety Assessment: Prevent generation of harmful content (child sexual abuse, dangerous content, explicit pornography, hate speech, harassment)
Evaluation Results
Across all safety tests, Gemma 4 shows significant improvements in all content safety categories compared to previous Gemma models. Overall, Gemma 4 models significantly outperform Gemma 3 in safety while reducing unjustified refusals to low levels.
Final Thoughts
The release of Gemma 4 marks a new stage for open-source multimodal models.
No longer “good enough to run” open-source models, but achieving: full-scenario coverage from mobile to server, dense + MoE dual architecture, native reasoning mode, 256K token long context, 140+ language support—each capability is strong individually, but combined they’re a dimensional strike.
More importantly, Google not only released the models but also provided Apache 2.0 open-source license, Transformers and llama.cpp support, detailed technical documentation and best practices. Developers can get started immediately, no waiting.
The future of multimodal open-source models may arrive faster than we imagine.
Reference Resources
- Hugging Face: https://huggingface.co/google/gemma-4-E2B-it
- Official Documentation: https://ai.google.dev/gemma/docs/core/model_card_4
More Articles