Benchmarks
Benchmarks

HLE-Verified

Knowledgeunit % · normalized over [0, 70]

HLE-Verified — the verified subset of Humanity's Last Exam, filtered for unambiguous ground truth. Scores are not comparable with the full HLE set filed under `hle`.

#ModelSourceScoreNormalized
Score distribution
4 tracked results across the normalization window
070
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.