Model
Model explorer

Step-4

OPEN
StepFun · Step family · released Sep 23, 2026

StepFun's September 23, 2026 flagship — the first Step-4 release and a return to open-weights flagships after the May Step-3.7-Flash mid-tier (610B/38B MoE, Apache-2.0). The launch table is StepFun's broadest yet: every shared row lands 4-8 points over 3.7-Flash, with a first-time OSWorld 3.0 filing (71.4) alongside ByteDance on the new revision. Capability flags keep the Step convention — agentic stays false (no first-party agent harness), and the $0.40/$1.60 API price undercuts the open-weights cohort.

ReasoningCodingVisionFunction callingTool useAgentic
2461.7
Elo · rank #38
Parameters
610B
Active params
38B (MoE)
Context
512K tokens
Architecture
610B-total/38B-active sparse MoE, Apache-2.0 open weights; 512K-token context
License
Apache-2.0
Languages
—
API price (in/out)
$0.4 / $1.6
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath96.2%#10
best: GPT-5.2 · 100.0%
BrowseCompAgents82.6%#23
best: Kimi K3.1 · 92.4%
GDPval-AAAgents1489#29
best: Claude Fable 5 · 1932
GPQA DiamondReasoning84.6%#73
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning55.2%#7
best: Claude Opus 5 · 64.7%
MCP AtlasAgents79.8%#3
best: Kimi K3.1 · 85.6%
MMMU-ProVision81.3%#8
best: Claude Opus 4.7 · 85.5%
OSWorld 3.0Agents71.4%#2
best: Seed 2.2 Pro · 74.6%
SWE-bench ProCoding61.8%#16
best: Claude Fable 5.1 · 81.2%
SWE-bench VerifiedCoding81.4%#13
best: Claude Opus 5 · 96.0%
Terminal-Bench 2.0Coding68.4%#37
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
—
VRAM @ FP16
—
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
—
Fine-tune it
Permissive
QLoRA414.8 GB4× H200 141GB
LoRA1299.3 GB8× B200 192GB
Full fine-tune9790.5 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $19.65 (4× H200 141GB)
Step family
Elo progression across releases
API price $0.4/$1.6 · each benchmark row carries its own source badge (see methodology)