Model
Model explorer

Step 3.7 Flash

OPEN
StepFun · Step family · released May 29, 2026

Newest StepFun release (successor to Step 3.5 Flash), adding native multimodal (image/GUI/document) support that the text-only Step 3.5 Flash lacked. Not previously in corpus (corpus's newest StepFun entry is Step-3 from 2025-07-31). No shipped 'Step 4' flagship found -- only reported as in training.

ReasoningCodingVisionFunction callingTool useAgentic
2212.7
Elo · rank #61
Parameters
198B
Active params
11B (MoE)
Context
262K tokens
Architecture
Sparse Mixture-of-Experts with 1.8B vision encoder + 196B language backbone, 198B total / 11B active
License
Apache 2.0
Languages
API price (in/out)
$0.2 / $1.15
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath95.0%#19
best: GPT-5.2 · 100.0%
BrowseCompAgents75.8%#22
best: Kimi K3 · 91.2%
GDPval-AAAgents1415.8#28
best: Claude Fable 5 · 1932
GPQA DiamondReasoning78.4%#95
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning49.7%#10
best: Claude Opus 5 · 64.7%
MMMU-ProVision75.0%#29
best: Claude Opus 4.7 · 85.5%
SWE-bench ProCoding56.3%#29
best: Claude Fable 5.1 · 81.2%
SWE-bench VerifiedCoding76.5%#37
best: Claude Opus 5 · 96.0%
Terminal-Bench 2.0Coding59.6%#44
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
VRAM @ FP16
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
Fine-tune it
Permissive
QLoRA134.6 GB1× H200 141GB
LoRA421.7 GB4× H200 141GB
Full fine-tune3177.9 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $5.69 (1× H200 141GB)
Step family
Elo progression across releases
API price $0.2/$1.15 · each benchmark row carries its own source badge (see methodology)