Benchmarks
Benchmarks

AutomationBench

Agentsunit % · normalized over [0, 50]

Zapier's cross-application workflow-orchestration benchmark: whether an agent can complete real business tasks end to end via REST/tool calls. Scored on a held-out private set where every assertion must pass.

#ModelSourceScoreNormalized
Score distribution
4 tracked results across the normalization window
050
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.