StableLearn Logo

Search Content

News 8 min read

Qwen3.5: 9B Beats 120B, 0.8B Runs Video on Phones—Full MoE Family

Qwen3.5 MoE series (0.8B-397B) with Gated DeltaNet hybrid attention. 9B outperforms 120B models, 0.8B processes video on smartphones. 1M context, native multimodal, 201 languages.

Cover image for Qwen3.5: 9B Beats 120B, 0.8B Runs Video on Phones—Full MoE Family

Published 210 days ago. Content may be outdated.

Alibaba’s Qwen team officially released the Qwen3.5 model family on February 16, 2026 — the latest flagship of the Qwen series. The headline model, Qwen3.5-397B-A17B, packs 397 billion total parameters while activating only 17 billion per forward pass, achieving new state-of-the-art results across reasoning, coding, and multimodal understanding as an open-source model.

On March 2, the Qwen3.5 Small Model Series went open-source — Qwen3.5-0.8B, 2B, 4B, and 9B, all under Apache 2.0. The 9B model scores 81.7 on GPQA Diamond, beating OpenAI’s GPT-OSS-120B (71.5) — a model 13.5x its size — while the 0.8B can process video on a smartphone. This is the Qwen3.5 architecture advantage playing out at every scale.

What makes Qwen3.5 remarkable isn’t incremental improvement — it’s fundamental architectural innovation. By introducing Gated DeltaNet hybrid attention and native multimodal training from scratch, Qwen3.5 matches or surpasses GPT-5.2 and Claude Opus 4.5 on 80% of evaluated dimensions while using a fraction of the active compute.

Architecture Innovation: Why Qwen3.5 Is Different

Gated DeltaNet: Breaking the Attention Bottleneck

Standard Transformer self-attention has a critical weakness: compute scales quadratically with sequence length. Processing 1 million tokens costs 100x more than processing 100K tokens — practically infeasible for real-world applications.

Qwen3.5’s solution is Gated DeltaNet — a linear attention variant that incorporates recurrent neural network concepts:

FeatureStandard AttentionGated DeltaNet
ComplexityO(n²) quadraticNear O(n) linear
1M token supportProhibitively expensiveFeasible and efficient
Long-range depsFull attention requiredGated state updates
Memory usageScales with sequence lengthNear-constant

The mechanism draws from the “Gated Delta Networks: Improving Mamba2 with Delta Rule” paper, combining Mamba2’s gated decay with a delta rule for hidden state updates.

The 3:1 Hybrid Ratio

Qwen3.5 doesn’t abandon traditional attention entirely. Instead, it uses an elegant 3:1 hybrid design:

   60 layers = 15 × (3 × [Gated DeltaNet → MoE] → 1 × [Full Attention → MoE])
  • Every 4 Transformer blocks: 3 use Gated DeltaNet (linear), 1 uses full quadratic attention
  • Linear attention handles efficient long-range context processing
  • Full attention ensures precise information retrieval at critical positions

This design was first introduced in Qwen3-Next (September 2025) and validated by NVIDIA’s engineering team before full production deployment in Qwen3.5.

Sparse MoE: Doing 397B’s Work with 17B

ParameterValue
Total parameters397B
Active parameters17B (only 4.3%)
Total experts512
Active experts10 routed + 1 shared
Expert intermediate dim1024
Hidden dimension4096
Layers60

The Shared Expert mechanism is distinctive to Qwen3.5’s MoE: a dedicated dense MLP processes every token to capture universal features, improving training stability and overall performance. The remaining 10 routed experts are dynamically allocated based on input, handling domain-specific knowledge.

Native Multimodal: End of the “Bolt-On” Era

For years, the industry standard was training a language model first, then attaching a vision encoder. Qwen3.5 completely abandons this approach — the model is trained from scratch on text, images, and video simultaneously.

Why Native Multimodal Is Superior

AspectBolt-on MultimodalNative Multimodal (Qwen3.5)
TrainingLanguage first, vision laterText + image + video jointly
Modal fusionLate alignment, information lossEarly fusion, seamless understanding
Training efficiencyLower for vision~100% of text-only efficiency
Model countSeparate VL models neededSingle unified model

This enables Qwen3.5 to comprehensively surpass the previous standalone Qwen3-VL model series on vision benchmarks.

Vision-Language Benchmark Results

BenchmarkQwen3.5Category
MMMU85.0STEM/General Reasoning
MMMU-Pro79.0Advanced Multimodal Reasoning
MathVision88.6Mathematical Visual Reasoning
MathVista (mini)90.3Math Visualization
OmniDocBench90.8Document Understanding
OCRBench93.1OCR Recognition
VideoMME (w/ subtitles)87.5Video Understanding
MLVU86.7Multilingual Video Understanding

Language Benchmarks: Outperforms GPT-5.2 on 80% of Dimensions

Math & Reasoning

BenchmarkQwen3.5Description
AIME2691.3Competition-level math
HMMT Feb 2594.8Elite math competition
MMLU-Pro87.8Comprehensive knowledge
GPQA Diamond88.4Graduate-level QA
C-Eval93.0Chinese comprehensive

Coding & Agents

BenchmarkQwen3.5Description
SWE-bench Verified76.4Real software engineering
LiveCodeBench v683.6Live coding evaluation
IFBench76.5Instruction following
LongBench v263.2Long context understanding

According to reports, Qwen3.5 outperforms GPT-5.2 and Claude Opus 4.5 on 80% of evaluated categories, though Claude maintains an edge on complex software engineering tasks (SWE-bench 80%+).

201 Languages, 250K Vocabulary

ComparisonQwen3Qwen3.5
Languages119201
Vocabulary size150K250K
Multi-token predictionNoYes (10-60% inference cost reduction)

Multi-Token Prediction (MTP) allows the model to “guess” multiple subsequent tokens in a single step, combined with speculative decoding to dramatically reduce inference costs across all 201 supported languages.

Context Window: Native 262K, Extensible to 1M

FeatureSpecification
Native context262,144 tokens
Extended context1,010,000 tokens (YaRN scaling)
Qwen3.5-PlusDefault 1M token context

Performance improvements over Qwen3-Max:

ScenarioSpeedup
256K token long-context19x faster
Standard workflow (32K)8.6x faster
Memory usage (FP8)50% reduction
Operating cost60% lower

Model Family: From 0.8B to 397B

Qwen3.5 follows a staggered release strategy:

Release DateModelType
2026.02.16Qwen3.5-397B-A17BMoE Flagship
2026.02.24Qwen3.5-122B-A10BMoE Medium
2026.02.24Qwen3.5-35B-A3BMoE Medium
2026.02.24Qwen3.5-27BDense Medium
2026.03.02Qwen3.5-9BSmall Flagship
2026.03.02Qwen3.5-4BLightweight Agent
2026.03.02Qwen3.5-2BEdge Device
2026.03.02Qwen3.5-0.8BMobile/IoT

Notably, Qwen3.5-35B-A3B beats Qwen3-235B-A22B on core benchmarks — purely through better architecture, data, and RL training, not more parameters.

March 2 Highlight: Small Model Series Goes Open-Source

On March 2, 2026, the Qwen team open-sourced Qwen3.5-0.8B, 2B, 4B, and 9B under Apache 2.0. These aren’t dumbed-down versions of the big model — every one inherits the full Gated DeltaNet hybrid architecture, with 262K native context windows and native multimodal capabilities (text + images + video).

This is the first time in AI history that a 0.8B model can process video, a 4B model can serve as a multimodal agent, and a 9B model comprehensively outperforms previous-generation 30B models.

Qwen3.5-9B: Compact Powerhouse That Closes the Gap

The 9B is the small series flagship, and its headline story is how dramatically it has closed the gap with much larger models.

vs. GPT-OSS-120B (13.5x its parameter count):

BenchmarkQwen3.5-9BGPT-OSS-120BGap
GPQA Diamond81.771.5+10.2
HMMT Feb 202583.276.7+6.5
MMLU-Pro82.580.8+1.7
MMMLU (multilingual)81.2—Leading

vs. GPT-5-Nano:

BenchmarkQwen3.5-9BGPT-5-NanoGap
MMMU-Pro70.157.2+12.9
MathVision78.962.2+16.7
OmniDocBench v1.587.755.9+31.8

vs. Previous-gen Qwen3-30B (3.3x its parameter count):

BenchmarkQwen3.5-9BQwen3-30BNotes
GPQA Diamond81.777.2Surpassed
Instruction Following91.588.9Surpassed
LongBench v255.248.0Surpassed

Hardware requirements: Runs on a single RTX 3090 (24GB) at BF16 precision. With 4-bit quantization, it drops to ~5GB — viable on an RTX 3060 12GB or M1 Mac with room to spare.

Qwen3.5-4B: Surprisingly Powerful Multimodal Base for Lightweight Agents

The 4B occupies a unique position — it’s the strongest multimodal agent base model at this parameter count.

With only 8GB of VRAM, Qwen3.5-4B supports:

CapabilityDetails
262K contextRare ultra-long context for a small model
Native visionImage recognition, document OCR, chart analysis
Video processingNative video understanding
Tool callingFunction Calling support
Thinking modeDeep reasoning with Thinking Mode

Previously, 4B-parameter models were limited to basic text tasks. Qwen3.5-4B shatters that expectation — it approaches large-model feature completeness while maintaining a minimal hardware footprint. For deploying lightweight AI Agents on consumer GPUs, this may be the best option available today.

Qwen3.5-0.8B / 2B: Tiny, Fast, Built for the Edge

The 0.8B and 2B are purpose-built for edge devices and on-device deployment:

FeatureQwen3.5-0.8BQwen3.5-2B
Target devicesSmartphones, IoT, embeddedEdge servers, rapid prototyping
MultimodalNative image + videoNative image + video
Context window262K tokens262K tokens
Offline videoSupported (60s / 8 FPS)Supported
Spatial reasoningSupportedSupported
Battery-friendlyUltra-low powerLow power

Key breakthrough: This is the first time a 0.8B-parameter model has native video processing capability. On-device offline video summarization (up to 60 seconds at 8 FPS) and spatial reasoning — without draining the battery.

Competitive Landscape

DimensionQwen3.5 SmallGoogle Gemma 3Meta Llama 3.2Microsoft Phi-4-mini
Smallest model0.8B1B1B14B
Native multimodalAll modelsPartialText-only at small sizesText-only
Context window262K128K128K128K
Video processingAll modelsPartialNoNo
LicenseApache 2.0Apache 2.0Llama LicenseMIT

Agentic Capabilities: Built for the Agent Era

Qwen3.5 isn’t just a chat model — it’s designed for the agentic paradigm:

CapabilityDetails
Tool callingNative function calling support
MCP protocolModel Context Protocol integration
Code executionBuilt-in code interpreter
Agent orchestrationMulti-agent collaboration framework
Adaptive thinkingThinking Mode for deep reasoning

Thinking Mode (Default On)

Qwen3.5 uses thinking mode by default, generating in the format:

   <think>
...reasoning process...
</think>

Final answer

Set max_tokens to 32,768+ for standard queries and 81,920 for complex math/coding tasks.

Deployment Guide

FrameworkBest For
SGLangProduction recommended
vLLMHigh-throughput serving
KTransformersCPU-GPU heterogeneous compute
TransformersDevelopment and debugging

SGLang Deployment

   python -m sglang.launch_server \
  --model-path Qwen/Qwen3.5-397B-A17B \
  --port 8000 \
  --tp-size 8 \
  --context-length 262144 \
  --reasoning-parser qwen3

vLLM Deployment

   vllm serve Qwen/Qwen3.5-397B-A17B \
  --port 8000 \
  --tensor-parallel-size 8 \
  --max-model-len 262144 \
  --reasoning-parser qwen3

Text-Only Mode (Save VRAM)

   vllm serve Qwen/Qwen3.5-397B-A17B \
  --port 8000 \
  --tensor-parallel-size 8 \
  --max-model-len 262144 \
  --language-model-only

API Usage Example

   from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="EMPTY"
)

# Text conversation
response = client.chat.completions.create(
    model="Qwen/Qwen3.5-397B-A17B",
    messages=[{"role": "user", "content": "Explain quantum computing basics"}],
    max_tokens=81920,
    temperature=0.6,
    top_p=0.95,
    extra_body={"top_k": 20},
)

# Image understanding
response = client.chat.completions.create(
    model="Qwen/Qwen3.5-397B-A17B",
    messages=[{
        "role": "user",
        "content": [
            {"type": "image_url", "image_url": {"url": "https://example.com/image.jpg"}},
            {"type": "text", "text": "Describe this image in detail"}
        ]
    }],
    max_tokens=81920,
)
ParameterThinking ModeNon-Thinking Mode
Temperature0.60.7
Top-P0.950.8
Top-K2020
Presence Penalty0.01.5

Competitive Landscape

vs. GPT-5.2 & Claude Opus 4.5

BenchmarkQwen3.5Notes
AIME2691.3Competition-level math
MMLU-Pro87.8Comprehensive knowledge
GPQA Diamond88.4Graduate-level QA
SWE-bench Verified76.4Claude leads at 80%+
LiveCodeBench v683.6Strong coding

Qwen3.5 reportedly outperforms GPT-5.2 and Claude Opus 4.5 on 80% of evaluated dimensions. Claude maintains a clear lead on complex software engineering tasks, while Gemini 3 holds advantages on some vision-specific benchmarks.

Architectural Significance: The New Attention Battlefield

Qwen3.5’s release marks several important inflection points in the AI landscape:

  1. Sparse MoE is the default scaling strategy — the era of training dense models with hundreds of billions of active parameters is ending
  2. Linear attention is production-ready — Gated DeltaNet proves linear attention can replace full quadratic attention for most layers without quality loss
  3. Attention is the new battleground — a year ago, the question was “MoE or Dense?” Now the divergence is in attention: Qwen chose Gated DeltaNet, DeepSeek chose MLA — two competing approaches
  4. Native multimodal is the future — separately training and bolting on vision modules will be phased out
  5. Small models aren’t “small” anymore — Qwen3.5-9B beating GPT-OSS-120B proves architectural innovation matters more than parameter stacking. The era of on-device AI has arrived

Open Source & Resources


Bottom Line: Qwen3.5 combines Gated DeltaNet hybrid attention with sparse MoE to achieve performance rivaling or exceeding trillion-parameter models while activating only 17B parameters. The March 2 small model release pushes this architectural advantage to the extreme — the 9B beats GPT-OSS-120B, the 4B becomes the strongest lightweight multimodal agent base, and the 0.8B brings video understanding to smartphones. From cloud to edge, Qwen3.5 is the most compelling open-source model family in the 2026 AI landscape.

Share Article

More Articles