Qwen3.5: 9B Beats 120B, 0.8B Runs Video on Phones—Full MoE Family
Qwen3.5 MoE series (0.8B-397B) with Gated DeltaNet hybrid attention. 9B outperforms 120B models, 0.8B processes video on smartphones. 1M context, native multimodal, 201 languages.
Published 210 days ago. Content may be outdated.
Alibaba’s Qwen team officially released the Qwen3.5 model family on February 16, 2026 — the latest flagship of the Qwen series. The headline model, Qwen3.5-397B-A17B, packs 397 billion total parameters while activating only 17 billion per forward pass, achieving new state-of-the-art results across reasoning, coding, and multimodal understanding as an open-source model.
On March 2, the Qwen3.5 Small Model Series went open-source — Qwen3.5-0.8B, 2B, 4B, and 9B, all under Apache 2.0. The 9B model scores 81.7 on GPQA Diamond, beating OpenAI’s GPT-OSS-120B (71.5) — a model 13.5x its size — while the 0.8B can process video on a smartphone. This is the Qwen3.5 architecture advantage playing out at every scale.
What makes Qwen3.5 remarkable isn’t incremental improvement — it’s fundamental architectural innovation. By introducing Gated DeltaNet hybrid attention and native multimodal training from scratch, Qwen3.5 matches or surpasses GPT-5.2 and Claude Opus 4.5 on 80% of evaluated dimensions while using a fraction of the active compute.
Architecture Innovation: Why Qwen3.5 Is Different
Gated DeltaNet: Breaking the Attention Bottleneck
Standard Transformer self-attention has a critical weakness: compute scales quadratically with sequence length. Processing 1 million tokens costs 100x more than processing 100K tokens — practically infeasible for real-world applications.
Qwen3.5’s solution is Gated DeltaNet — a linear attention variant that incorporates recurrent neural network concepts:
| Feature | Standard Attention | Gated DeltaNet |
|---|---|---|
| Complexity | O(n²) quadratic | Near O(n) linear |
| 1M token support | Prohibitively expensive | Feasible and efficient |
| Long-range deps | Full attention required | Gated state updates |
| Memory usage | Scales with sequence length | Near-constant |
The mechanism draws from the “Gated Delta Networks: Improving Mamba2 with Delta Rule” paper, combining Mamba2’s gated decay with a delta rule for hidden state updates.
The 3:1 Hybrid Ratio
Qwen3.5 doesn’t abandon traditional attention entirely. Instead, it uses an elegant 3:1 hybrid design:
60 layers = 15 × (3 × [Gated DeltaNet → MoE] → 1 × [Full Attention → MoE])
- Every 4 Transformer blocks: 3 use Gated DeltaNet (linear), 1 uses full quadratic attention
- Linear attention handles efficient long-range context processing
- Full attention ensures precise information retrieval at critical positions
This design was first introduced in Qwen3-Next (September 2025) and validated by NVIDIA’s engineering team before full production deployment in Qwen3.5.
Sparse MoE: Doing 397B’s Work with 17B
| Parameter | Value |
|---|---|
| Total parameters | 397B |
| Active parameters | 17B (only 4.3%) |
| Total experts | 512 |
| Active experts | 10 routed + 1 shared |
| Expert intermediate dim | 1024 |
| Hidden dimension | 4096 |
| Layers | 60 |
The Shared Expert mechanism is distinctive to Qwen3.5’s MoE: a dedicated dense MLP processes every token to capture universal features, improving training stability and overall performance. The remaining 10 routed experts are dynamically allocated based on input, handling domain-specific knowledge.
Native Multimodal: End of the “Bolt-On” Era
For years, the industry standard was training a language model first, then attaching a vision encoder. Qwen3.5 completely abandons this approach — the model is trained from scratch on text, images, and video simultaneously.
Why Native Multimodal Is Superior
| Aspect | Bolt-on Multimodal | Native Multimodal (Qwen3.5) |
|---|---|---|
| Training | Language first, vision later | Text + image + video jointly |
| Modal fusion | Late alignment, information loss | Early fusion, seamless understanding |
| Training efficiency | Lower for vision | ~100% of text-only efficiency |
| Model count | Separate VL models needed | Single unified model |
This enables Qwen3.5 to comprehensively surpass the previous standalone Qwen3-VL model series on vision benchmarks.
Vision-Language Benchmark Results
| Benchmark | Qwen3.5 | Category |
|---|---|---|
| MMMU | 85.0 | STEM/General Reasoning |
| MMMU-Pro | 79.0 | Advanced Multimodal Reasoning |
| MathVision | 88.6 | Mathematical Visual Reasoning |
| MathVista (mini) | 90.3 | Math Visualization |
| OmniDocBench | 90.8 | Document Understanding |
| OCRBench | 93.1 | OCR Recognition |
| VideoMME (w/ subtitles) | 87.5 | Video Understanding |
| MLVU | 86.7 | Multilingual Video Understanding |
Language Benchmarks: Outperforms GPT-5.2 on 80% of Dimensions
Math & Reasoning
| Benchmark | Qwen3.5 | Description |
|---|---|---|
| AIME26 | 91.3 | Competition-level math |
| HMMT Feb 25 | 94.8 | Elite math competition |
| MMLU-Pro | 87.8 | Comprehensive knowledge |
| GPQA Diamond | 88.4 | Graduate-level QA |
| C-Eval | 93.0 | Chinese comprehensive |
Coding & Agents
| Benchmark | Qwen3.5 | Description |
|---|---|---|
| SWE-bench Verified | 76.4 | Real software engineering |
| LiveCodeBench v6 | 83.6 | Live coding evaluation |
| IFBench | 76.5 | Instruction following |
| LongBench v2 | 63.2 | Long context understanding |
According to reports, Qwen3.5 outperforms GPT-5.2 and Claude Opus 4.5 on 80% of evaluated categories, though Claude maintains an edge on complex software engineering tasks (SWE-bench 80%+).
201 Languages, 250K Vocabulary
| Comparison | Qwen3 | Qwen3.5 |
|---|---|---|
| Languages | 119 | 201 |
| Vocabulary size | 150K | 250K |
| Multi-token prediction | No | Yes (10-60% inference cost reduction) |
Multi-Token Prediction (MTP) allows the model to “guess” multiple subsequent tokens in a single step, combined with speculative decoding to dramatically reduce inference costs across all 201 supported languages.
Context Window: Native 262K, Extensible to 1M
| Feature | Specification |
|---|---|
| Native context | 262,144 tokens |
| Extended context | 1,010,000 tokens (YaRN scaling) |
| Qwen3.5-Plus | Default 1M token context |
Performance improvements over Qwen3-Max:
| Scenario | Speedup |
|---|---|
| 256K token long-context | 19x faster |
| Standard workflow (32K) | 8.6x faster |
| Memory usage (FP8) | 50% reduction |
| Operating cost | 60% lower |
Model Family: From 0.8B to 397B
Qwen3.5 follows a staggered release strategy:
| Release Date | Model | Type |
|---|---|---|
| 2026.02.16 | Qwen3.5-397B-A17B | MoE Flagship |
| 2026.02.24 | Qwen3.5-122B-A10B | MoE Medium |
| 2026.02.24 | Qwen3.5-35B-A3B | MoE Medium |
| 2026.02.24 | Qwen3.5-27B | Dense Medium |
| 2026.03.02 | Qwen3.5-9B | Small Flagship |
| 2026.03.02 | Qwen3.5-4B | Lightweight Agent |
| 2026.03.02 | Qwen3.5-2B | Edge Device |
| 2026.03.02 | Qwen3.5-0.8B | Mobile/IoT |
Notably, Qwen3.5-35B-A3B beats Qwen3-235B-A22B on core benchmarks — purely through better architecture, data, and RL training, not more parameters.
March 2 Highlight: Small Model Series Goes Open-Source
On March 2, 2026, the Qwen team open-sourced Qwen3.5-0.8B, 2B, 4B, and 9B under Apache 2.0. These aren’t dumbed-down versions of the big model — every one inherits the full Gated DeltaNet hybrid architecture, with 262K native context windows and native multimodal capabilities (text + images + video).
This is the first time in AI history that a 0.8B model can process video, a 4B model can serve as a multimodal agent, and a 9B model comprehensively outperforms previous-generation 30B models.
Qwen3.5-9B: Compact Powerhouse That Closes the Gap
The 9B is the small series flagship, and its headline story is how dramatically it has closed the gap with much larger models.
vs. GPT-OSS-120B (13.5x its parameter count):
| Benchmark | Qwen3.5-9B | GPT-OSS-120B | Gap |
|---|---|---|---|
| GPQA Diamond | 81.7 | 71.5 | +10.2 |
| HMMT Feb 2025 | 83.2 | 76.7 | +6.5 |
| MMLU-Pro | 82.5 | 80.8 | +1.7 |
| MMMLU (multilingual) | 81.2 | — | Leading |
vs. GPT-5-Nano:
| Benchmark | Qwen3.5-9B | GPT-5-Nano | Gap |
|---|---|---|---|
| MMMU-Pro | 70.1 | 57.2 | +12.9 |
| MathVision | 78.9 | 62.2 | +16.7 |
| OmniDocBench v1.5 | 87.7 | 55.9 | +31.8 |
vs. Previous-gen Qwen3-30B (3.3x its parameter count):
| Benchmark | Qwen3.5-9B | Qwen3-30B | Notes |
|---|---|---|---|
| GPQA Diamond | 81.7 | 77.2 | Surpassed |
| Instruction Following | 91.5 | 88.9 | Surpassed |
| LongBench v2 | 55.2 | 48.0 | Surpassed |
Hardware requirements: Runs on a single RTX 3090 (24GB) at BF16 precision. With 4-bit quantization, it drops to ~5GB — viable on an RTX 3060 12GB or M1 Mac with room to spare.
Qwen3.5-4B: Surprisingly Powerful Multimodal Base for Lightweight Agents
The 4B occupies a unique position — it’s the strongest multimodal agent base model at this parameter count.
With only 8GB of VRAM, Qwen3.5-4B supports:
| Capability | Details |
|---|---|
| 262K context | Rare ultra-long context for a small model |
| Native vision | Image recognition, document OCR, chart analysis |
| Video processing | Native video understanding |
| Tool calling | Function Calling support |
| Thinking mode | Deep reasoning with Thinking Mode |
Previously, 4B-parameter models were limited to basic text tasks. Qwen3.5-4B shatters that expectation — it approaches large-model feature completeness while maintaining a minimal hardware footprint. For deploying lightweight AI Agents on consumer GPUs, this may be the best option available today.
Qwen3.5-0.8B / 2B: Tiny, Fast, Built for the Edge
The 0.8B and 2B are purpose-built for edge devices and on-device deployment:
| Feature | Qwen3.5-0.8B | Qwen3.5-2B |
|---|---|---|
| Target devices | Smartphones, IoT, embedded | Edge servers, rapid prototyping |
| Multimodal | Native image + video | Native image + video |
| Context window | 262K tokens | 262K tokens |
| Offline video | Supported (60s / 8 FPS) | Supported |
| Spatial reasoning | Supported | Supported |
| Battery-friendly | Ultra-low power | Low power |
Key breakthrough: This is the first time a 0.8B-parameter model has native video processing capability. On-device offline video summarization (up to 60 seconds at 8 FPS) and spatial reasoning — without draining the battery.
Competitive Landscape
| Dimension | Qwen3.5 Small | Google Gemma 3 | Meta Llama 3.2 | Microsoft Phi-4-mini |
|---|---|---|---|---|
| Smallest model | 0.8B | 1B | 1B | 14B |
| Native multimodal | All models | Partial | Text-only at small sizes | Text-only |
| Context window | 262K | 128K | 128K | 128K |
| Video processing | All models | Partial | No | No |
| License | Apache 2.0 | Apache 2.0 | Llama License | MIT |
Agentic Capabilities: Built for the Agent Era
Qwen3.5 isn’t just a chat model — it’s designed for the agentic paradigm:
| Capability | Details |
|---|---|
| Tool calling | Native function calling support |
| MCP protocol | Model Context Protocol integration |
| Code execution | Built-in code interpreter |
| Agent orchestration | Multi-agent collaboration framework |
| Adaptive thinking | Thinking Mode for deep reasoning |
Thinking Mode (Default On)
Qwen3.5 uses thinking mode by default, generating in the format:
<think>
...reasoning process...
</think>
Final answer
Set max_tokens to 32,768+ for standard queries and 81,920 for complex math/coding tasks.
Deployment Guide
Recommended Frameworks
| Framework | Best For |
|---|---|
| SGLang | Production recommended |
| vLLM | High-throughput serving |
| KTransformers | CPU-GPU heterogeneous compute |
| Transformers | Development and debugging |
SGLang Deployment
python -m sglang.launch_server \
--model-path Qwen/Qwen3.5-397B-A17B \
--port 8000 \
--tp-size 8 \
--context-length 262144 \
--reasoning-parser qwen3
vLLM Deployment
vllm serve Qwen/Qwen3.5-397B-A17B \
--port 8000 \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--reasoning-parser qwen3
Text-Only Mode (Save VRAM)
vllm serve Qwen/Qwen3.5-397B-A17B \
--port 8000 \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--language-model-only
API Usage Example
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY"
)
# Text conversation
response = client.chat.completions.create(
model="Qwen/Qwen3.5-397B-A17B",
messages=[{"role": "user", "content": "Explain quantum computing basics"}],
max_tokens=81920,
temperature=0.6,
top_p=0.95,
extra_body={"top_k": 20},
)
# Image understanding
response = client.chat.completions.create(
model="Qwen/Qwen3.5-397B-A17B",
messages=[{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://example.com/image.jpg"}},
{"type": "text", "text": "Describe this image in detail"}
]
}],
max_tokens=81920,
)
Recommended Sampling Parameters
| Parameter | Thinking Mode | Non-Thinking Mode |
|---|---|---|
| Temperature | 0.6 | 0.7 |
| Top-P | 0.95 | 0.8 |
| Top-K | 20 | 20 |
| Presence Penalty | 0.0 | 1.5 |
Competitive Landscape
vs. GPT-5.2 & Claude Opus 4.5
| Benchmark | Qwen3.5 | Notes |
|---|---|---|
| AIME26 | 91.3 | Competition-level math |
| MMLU-Pro | 87.8 | Comprehensive knowledge |
| GPQA Diamond | 88.4 | Graduate-level QA |
| SWE-bench Verified | 76.4 | Claude leads at 80%+ |
| LiveCodeBench v6 | 83.6 | Strong coding |
Qwen3.5 reportedly outperforms GPT-5.2 and Claude Opus 4.5 on 80% of evaluated dimensions. Claude maintains a clear lead on complex software engineering tasks, while Gemini 3 holds advantages on some vision-specific benchmarks.
Architectural Significance: The New Attention Battlefield
Qwen3.5’s release marks several important inflection points in the AI landscape:
- Sparse MoE is the default scaling strategy — the era of training dense models with hundreds of billions of active parameters is ending
- Linear attention is production-ready — Gated DeltaNet proves linear attention can replace full quadratic attention for most layers without quality loss
- Attention is the new battleground — a year ago, the question was “MoE or Dense?” Now the divergence is in attention: Qwen chose Gated DeltaNet, DeepSeek chose MLA — two competing approaches
- Native multimodal is the future — separately training and bolting on vision modules will be phased out
- Small models aren’t “small” anymore — Qwen3.5-9B beating GPT-OSS-120B proves architectural innovation matters more than parameter stacking. The era of on-device AI has arrived
Open Source & Resources
- License: Apache 2.0 (fully open for commercial use)
- Model Download: Hugging Face
- GitHub: QwenLM/Qwen3.5
- Official API: Alibaba Cloud Model Studio
- Try Online: Qwen Chat
- Qwen-Agent: GitHub
Bottom Line: Qwen3.5 combines Gated DeltaNet hybrid attention with sparse MoE to achieve performance rivaling or exceeding trillion-parameter models while activating only 17B parameters. The March 2 small model release pushes this architectural advantage to the extreme — the 9B beats GPT-OSS-120B, the 4B becomes the strongest lightweight multimodal agent base, and the 0.8B brings video understanding to smartphones. From cloud to edge, Qwen3.5 is the most compelling open-source model family in the 2026 AI landscape.
More Articles