Model
Model explorer

Qwen3.8-Flash-Next

OPEN
Alibaba · Qwen3.8 family · released Aug 26, 2026

Alibaba's August 26, 2026 open-weight architecture preview — explicitly an early look at the stack Qwen4 will be built on, released early so the community can examine it, exactly as Qwen3-Next was. The headline is efficiency: 6B activated parameters, 8.6x the prefill throughput of Qwen3.7-Plus at 1M context under a 90% prefix-cache hit rate, and $0.15/$0.47 per Mtok for the production `qwen3.8-flash` SKU on QwenCloud (which defaults to 1M context and built-in tools). The 51B N-gram embedding table is the notable trick — capacity added at almost no per-token compute, stored in host memory and prefetched asynchronously. Results are strong for the activated size (SWE-bench Pro 62.5 beats Claude Opus 4.6 Max's 53.4; SWE-bench Multilingual 81.0) but the ceiling shows on long-horizon work: DeepSWE 1.1 58.7 and OSWorld 2.0 19.4 are far off the frontier. Rows are taken from the post-trained model card at the native 262K context, not the 1M YaRN extension; the in-house CoWorkBench, RecreationBench and ClawEval-MM benchmarks are deliberately not filed as catalog benchmarks.

ReasoningCodingVisionFunction callingTool useAgentic
2659.1
Elo · rank #21
Parameters
125B
Active params
6B (MoE)
Context
262K tokens
Architecture
125B-parameter multimodal MoE with 6B activated per token, plus 51B of N-gram embedding parameters that are deterministically addressed and never enter the per-token matmul budget. GDN + Qwen Sparse Attention hybrid (3 of every 4 layers Gated DeltaNet), gated residual, Muon optimizer; 262,144-token native context extensible to 1M with YaRN
License
Apache 2.0
Languages
API price (in/out)
$0.15 / $0.47
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
Agents' Last ExamAgents51.2%#3
best: GPT-6 Astra · 59.3%
AndroidWorldAgents84.5%#1
best: this model · 84.5%
CharXivVision90.6%#1
best: this model · 90.6%
DeepSWECoding58.7%#12
best: Muse Spark 1.3 · 75.4%
GPQA DiamondReasoning91.7%#22
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning35.9%#37
best: Claude Opus 5 · 64.7%
IFBenchReasoning81.3%#6
best: MiniMax M3 · 83.0%
JobBenchAgents55.7%#5
best: Claude Opus 5 · 65.7%
LiveCodeBenchCoding91.9%#2
best: DeepSeek-V4-Pro (Think Max) · 93.5%
LVBenchVision76.6%#5
best: Gemini 3.8 Flash · 87.8%
MathVisionVision90.6%#2
best: Seed 2.1 Pro · 92.6%
NL2RepoCoding48.1%#3
best: GLM-5.3 · 58.0%
OSWorld 2.0Agents19.4%#9
best: Claude Fable 5.1 · 77.9%
best: Claude Opus 5 · 89.5%
SWE-bench ProCoding62.5%#12
best: Claude Fable 5.1 · 81.2%
best: GLM-5.3-Flash · 78.4%
Run it locally
VRAM @ Q4
VRAM @ FP16
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
Fine-tune it
Permissive
QLoRA85.0 GB1× H200 141GB
LoRA266.3 GB2× H200 141GB
Full fine-tune2006.3 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $64.65 (1× H200 141GB)
API price $0.15/$0.47 · each benchmark row carries its own source badge (see methodology)