Benchmarks
Benchmarks

DeepSearchQA

Agentsunit % · normalized over [0, 100]

DeepSearchQA — 900 agentic browsing questions whose answers are lists of items; the agent researches each with search/open/find tools and is graded by semantic set matching (F1).

#ModelSourceScoreNormalized
Score distribution
6 tracked results across the normalization window
0100
Score vs. parameters
Open-weights models, log-x params
1B10B100B1000B