Benchmarks
Benchmarks

Terminal-Bench 3.0

Agentsunit % · normalized over [0, 45]

Terminal-Bench 3.0 — the third-generation terminal-agent suite, a large step up in difficulty from 2.1 (frontier models score in the 20-35% range where they clear 85% on 2.1). General agent capability in a real shell, not coding alone.

#ModelSourceScoreNormalized
Score distribution
8 tracked results across the normalization window
045
Score vs. parameters
Open-weights models, log-x params
1B10B100B1000B