Benchmarks
Benchmarks
HLE-Verified
Knowledgeunit % · normalized over [0, 70]HLE-Verified — the verified subset of Humanity's Last Exam, filtered for unambiguous ground truth. Scores are not comparable with the full HLE set filed under `hle`.
#ModelSourceScoreNormalized
Score distribution
4 tracked results across the normalization window
070
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.