Benchmarks
Benchmarks
ARC-AGI-3
Reasoningunit % · normalized over [0, 50]Interactive abstraction-and-reasoning benchmark (ARC Prize): agents explore novel turn-based environments with no instructions, infer goals on the fly, build world models, and adapt across increasingly hard levels. v3.
#ModelSourceScoreNormalized
Score distribution
3 tracked results across the normalization window
050
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.