Benchmarks
Benchmarks
DeepSearchQA
Agentsunit % · normalized over [0, 100]DeepSearchQA — 900 agentic browsing questions whose answers are lists of items; the agent researches each with search/open/find tools and is graded by semantic set matching (F1).
#ModelSourceScoreNormalized
Score distribution
6 tracked results across the normalization window
0100
Score vs. parameters
Open-weights models, log-x params