Benchmarks
Benchmarks
GAIA2
Agentsunit % · normalized over [0, 60]GAIA2 — second-generation general AI assistant benchmark: multi-step, tool-using real-world assistant tasks with verifiable answers.
#ModelSourceScoreNormalized
Score distribution
2 tracked results across the normalization window
060
Score vs. parameters
Open-weights models, log-x params