Benchmarks
Benchmarks
JobBench
Agentsunit % · normalized over [0, 75]JobBench — 65 professional tasks across 35 white-collar occupations, each asking the agent to produce a practical work product scored against task-specific rubrics. Mean rubric score.
#ModelSourceScoreNormalized
Score distribution
9 tracked results across the normalization window
075
Score vs. parameters
Open-weights models, log-x params