StableLearn Logo

Search Content

News 7 min read

Gemma 4 Release: Google's 4-Size Open Source Family with MoE Architecture

Google releases Gemma 4 multimodal models with dense and MoE architectures, 256K token context, native reasoning mode, and deployment from mobile to server. Four model sizes available under Apache 2.0.

Cover image for Gemma 4 Release: Google's 4-Size Open Source Family with MoE Architecture

Published 179 days ago. Content may be outdated.

Google DeepMind just dropped a bombshell! On April 2nd, the Gemma 4 series officially launched—a complete “family pack” ranging from 2.3B to 30.7B parameters, featuring both dense and MoE (Mixture of Experts) architectures, context windows up to 256K tokens, and native reasoning mode. Even better? Full Apache 2.0 open source, deployable from smartphones to servers.

Let’s look at the numbers: 4 model sizes, 256K token context, 140+ languages, native reasoning mode. What does this mean? It can read an entire technical document in one go and perform reasoning analysis on it.

Model Family: From Mobile to Server, Full Coverage

Gemma 4 isn’t flying solo—it launches with 4 models covering everything from mobile devices to data centers.

Dense Models: Three Brothers, Each with Strengths

ModelParametersContext LengthModalitiesTarget
E2B2.3B effective (5.1B with embeddings)128k tokensText, Image, AudioMobile/Edge
E4B4.5B effective (8.0B with embeddings)128k tokensText, Image, AudioLaptop/Edge
31B Dense30.7B256k tokensText, ImageWorkstation/Server

The “E” in E2B and E4B stands for “effective parameters.” These smaller models use Per-Layer Embeddings (PLE) technology—simply put, each decoder layer has its own small embedding table. The embedding tables are large but only used for fast lookups, so the effective parameter count is much smaller than the total. This design is optimized for on-device deployment, enabling efficient operation on phones and laptops.

MoE Model: Perfect Balance of Speed and Capability

ModelTotal ParametersActive ParametersContext LengthExpertsModalities
26B A4B25.2B3.8B256k tokens8 active / 128 total + 1 sharedText, Image

The “A” in 26B A4B stands for “active parameters.” Through the MoE architecture, only a 4B parameter subset is activated during inference, running almost as fast as a 4B model while delivering performance close to a 26B model. This is an excellent choice for fast inference.

The key is Hybrid Attention mechanism: interweaving local sliding window attention with global attention, ensuring the final layer is always global. This design provides the processing speed and low memory footprint of lightweight models without sacrificing the deep awareness needed for complex long-context tasks.

Performance Explosion: SOTA Across Multiple Benchmarks

The data doesn’t lie. Gemma 4 achieves SOTA across multiple benchmarks.

Core Capability Comparison

BenchmarkGemma 4 31BGemma 4 26B A4BGemma 4 E4BGemma 4 E2BGemma 3 27B
MMLU Pro85.2%82.6%69.4%60.0%67.6%
AIME 2026 No Tools89.2%88.3%42.5%37.5%20.8%
LiveCodeBench v680.0%77.1%52.0%44.0%29.1%
Codeforces ELO21501718940633110
GPQA Diamond84.3%82.3%58.6%43.4%42.4%

Look at these numbers: 31B scores 89.2% on AIME 2026 (math competition), 80.0% on LiveCodeBench (coding ability), and reaches 2150 Codeforces ELO—what level is this? It already surpasses most human programmers.

Even more impressive, the smallest E2B achieves 60.0% on MMLU Pro, only 7.6 percentage points behind Gemma 3 27B’s 67.6%, with less than one-tenth the parameters.

Multimodal Capability: Comprehensive Vision Understanding

CapabilityGemma 4 31BGemma 4 26B A4BGemma 4 E4BGemma 4 E2B
MMMU Pro76.9%73.8%52.6%44.2%
MATH-Vision85.6%82.4%59.5%52.4%
OmniDocBench 1.5 (lower is better)0.1310.1490.1810.290

For vision understanding, 31B scores 85.6% on MATH-Vision and achieves an edit distance of only 0.131 on document parsing (OmniDocBench)—meaning it can accurately recognize and understand document content with maxed-out OCR capability.

Audio Capability: E2B/E4B Exclusive

CapabilityGemma 4 E4BGemma 4 E2B
CoVoST (Speech Translation)35.5433.47
FLEURS (Speech Recognition, lower is better)0.080.09

E2B and E4B natively support audio input with excellent performance in speech recognition and translation tasks.

Long Context: Truly “Reading Entire Books”

Context LengthGemma 4 31BGemma 4 26B A4BGemma 4 E4BGemma 4 E2B
MRCR v2 8-needle 128k (avg)66.4%44.1%25.4%19.1%

31B scores 66.4% on 128K token long-context tasks. What does this mean? It can read a medium-length technical book in one go and answer questions about it.

Core Features: Not Just Multimodal, But Reasoning Mode

Gemma 4 isn’t just a simple multimodal model—it comes with a bunch of “black tech.”

Feature Overview

FeatureSupported ModelsDescription
Thinking ModeAll modelsBuilt-in reasoning mode allowing step-by-step thinking before answering
Long ContextAll modelsE2B/E4B support 128K tokens, 26B A4B/31B support 256K tokens
Image UnderstandingAll modelsObject detection, document/PDF parsing, OCR, handwriting recognition, variable aspect ratios and resolutions
Video UnderstandingAll modelsAnalyze video by processing frame sequences
Audio ProcessingE2B/E4BAutomatic Speech Recognition (ASR) and speech translation (multiple languages)
Function CallingAll modelsNative support for structured tool use, enabling agent workflows
MultilingualAll modelsOut-of-box support for 35+ languages, pretrained on 140+ languages

Black Tech: Thinking Mode

The most eye-catching feature is Thinking Mode.

Simply put, the model “thinks” before answering. By including the {"<|think|>"} token in the system prompt, the model outputs its internal reasoning process before giving the final answer.

This is Gemma 4’s unique “emergent ability,” enabling it to handle more complex reasoning tasks. Performance on AIME 2026 (math competition) and Codeforces (programming competition) is the best proof.

Variable Image Resolution: Balance Between Detail and Speed

Gemma 4 supports controlling image resolution through configurable vision token budgets. Supported token budgets: 70, 140, 280, 560, 1120.

  • Low budget (70/140): Suitable for classification, captioning, or video understanding tasks—fast
  • High budget (560/1120): Suitable for OCR, document parsing, or reading small text—detailed

This design allows the model to flexibly adjust based on task requirements, quickly processing large numbers of frames while accurately recognizing details.

How to Use? Transformers and llama.cpp Both Supported

Using Gemma 4 is very simple—both Transformers and llama.cpp are already supported.

Quick Start

   import torch
from transformers import AutoProcessor, AutoModelForCausalLM

MODEL_ID = "google/gemma-4-E2B-it"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    dtype=torch.bfloat16,
    device_map="auto"
)

# Construct dialogue
messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Write a short joke about saving RAM."},
]

# Process input
text = processor.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False  # Set to True to enable reasoning mode
)
inputs = processor(text=text, return_tensors="pt").to(model.device)
input_len = inputs["input_ids"].shape[-1]

# Generate output
outputs = model.generate(**inputs, max_new_tokens=1024)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)

# Parse thinking output
processor.parse_response(response)

Best Practices

ConfigurationRecommended ValueDescription
temperature1.0Standardized sampling configuration
top_p0.95Standardized sampling configuration
top_k64Standardized sampling configuration
Thinking ModeEnable as neededInclude think token in system prompt
Image Token Budget70-1120Choose based on task requirements

Technical Deep Dive: Hybrid Attention + MoE

Core Architecture

ComponentFunctionPurpose
Hybrid AttentionLocal sliding window + global attentionBalance speed and long-context understanding
Per-Layer EmbeddingsIndependent embedding table per layerImprove on-device parameter efficiency
Mixture of Experts (MoE)8 active / 128 total expertsOnly activate subset during inference for speed
Unified KV + p-RoPEGlobal layer optimizationReduce long-context memory usage

The key is Hybrid Attention mechanism: interweaving local sliding window attention with global attention, ensuring the final layer is always global. This design provides the processing speed and low memory footprint of lightweight models without sacrificing the deep awareness needed for complex long-context tasks.

For MoE models, by activating only a 4B parameter subset during inference, 26B A4B runs almost as fast as a 4B model while delivering performance close to a 26B model.

Application Scenarios: From Content Creation to Agents

ApplicationCore CapabilityReal Value
Content CreationText generation, code generationPoetry, scripts, marketing copy, code completion
ChatbotsConversational AICustomer service, virtual assistants, interactive apps
Document ProcessingOCR, document parsingExtract, interpret, and summarize visual data
Audio ProcessingASR, speech translationMeeting transcription, multilingual translation
Agent WorkflowsFunction callingStructured tool use, autonomous agents
Research & EducationNLP research, language learningAlgorithm development, grammar correction, writing practice

Safety and Ethics: Google’s AI Principles

Gemma 4 is developed by Google DeepMind and, like the proprietary Gemini models, undergoes rigorous safety evaluation.

Evaluation Methods

  • CSAM Filtering: Strict filtering applied at multiple stages of data preparation
  • Sensitive Data Filtering: Automated techniques filter personal information and other sensitive data
  • Content Safety Assessment: Prevent generation of harmful content (child sexual abuse, dangerous content, explicit pornography, hate speech, harassment)

Evaluation Results

Across all safety tests, Gemma 4 shows significant improvements in all content safety categories compared to previous Gemma models. Overall, Gemma 4 models significantly outperform Gemma 3 in safety while reducing unjustified refusals to low levels.

Final Thoughts

The release of Gemma 4 marks a new stage for open-source multimodal models.

No longer “good enough to run” open-source models, but achieving: full-scenario coverage from mobile to server, dense + MoE dual architecture, native reasoning mode, 256K token long context, 140+ language support—each capability is strong individually, but combined they’re a dimensional strike.

More importantly, Google not only released the models but also provided Apache 2.0 open-source license, Transformers and llama.cpp support, detailed technical documentation and best practices. Developers can get started immediately, no waiting.

The future of multimodal open-source models may arrive faster than we imagine.


Reference Resources

Share Article

More Articles