Benchmarks
Benchmarks

Terminal-Bench 4.0

Agentsunit % · normalized over [0, 70]

Terminal-Bench 4.0 — 66-task successor suite spanning computational biology, physics simulation, CAD and formal proofs, run in a real terminal environment. Measures general long-horizon agent capability rather than coding specifically.

#ModelSourceScoreNormalized
Score distribution
8 tracked results across the normalization window
070
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.