Benchmarks
Benchmarks
ExploitBench
Agentsunit % · normalized over [0, 100]ExploitBench — turns known, disclosed vulnerabilities into working exploits end to end, the step above vulnerability discovery on the exploitation chain. Public-set scores are contamination-sensitive: labs increasingly report private internal ports alongside them.
#ModelSourceScoreNormalized
Score distribution
7 tracked results across the normalization window
0100
Score vs. parameters
Open-weights models, log-x params