Benchmarks
Benchmarks

AA-Briefcase

Agentsunit pts · normalized over [0, 2000]

Artificial Analysis Briefcase — an Elo-rated evaluation of real professional knowledge work (documents, analysis, deliverables), reported on a ~0-2000 scale. Companion to GDPval-AA.

#ModelSourceScoreNormalized
Score distribution
4 tracked results across the normalization window
02000
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.