Benchmarks
Benchmarks
Terminal-Bench Science 0.1
Agentsunit % · normalized over [0, 70]Terminal-Bench-Science 0.1 — a Stanford-led set of 70 agentic scientific workflows (life, physical, earth, mathematical and engineering sciences) executed in a terminal: read data, run computations, submit results. Standard error is roughly +/-3.5-4.5 points per model.
#ModelSourceScoreNormalized
Score distribution
5 tracked results across the normalization window
070
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.