NexBench provides signals of a model's ability to write, reason and execute real exploits against a real target in a messy environment, while measuring price to performance ratio, token efficiency, and more.
Best severity-weighted score across every reasoning-effort run= 61
Best severity-weighted score across every reasoning-effort run= 52
Best severity-weighted score across every reasoning-effort run= 45
Best severity-weighted score across every reasoning-effort run= 43
Best severity-weighted score across every reasoning-effort run= 42
Best severity-weighted score across every reasoning-effort run= 41
Best severity-weighted score across every reasoning-effort run= 36
Best severity-weighted score across every reasoning-effort run= 34
Best severity-weighted score across every reasoning-effort run= 34
Best severity-weighted score across every reasoning-effort run= 28
Best severity-weighted score across every reasoning-effort run= 24
Best severity-weighted score across every reasoning-effort run= 18
03365
61
52
45
43
42
41
36
34
34
28
24
18
Best Runtime
Best-run wall time · Lower is better
Kimi K3
Qwen3.8 27B
GPT-5.6 Terra
GPT-5.5
GLM-5.2
Opus 4.8
GPT-5.6 Luna
GPT-5.6 Sol
Kimi K2.7
Sonnet 5
Grok 4.5
Nemotron 3 Ultra
Wall-clock run time of the best-scoring attempt= 49.0m
Wall-clock run time of the best-scoring attempt= 59.0m
Wall-clock run time of the best-scoring attempt= 66.2m
Wall-clock run time of the best-scoring attempt= 78.8m
Wall-clock run time of the best-scoring attempt= 100.0m
Wall-clock run time of the best-scoring attempt= 110.5m
Wall-clock run time of the best-scoring attempt= 119.0m
Wall-clock run time of the best-scoring attempt= 139.1m
Wall-clock run time of the best-scoring attempt= 152.0m
Wall-clock run time of the best-scoring attempt= 269.7m
Wall-clock run time of the best-scoring attempt= 302.5m
Wall-clock run time of the best-scoring attempt= 302.7m
0.0m175.0m350.0m
49.0m
59.0m
66.2m
78.8m
100.0m
110.5m
119.0m
139.1m
152.0m
269.7m
302.5m
302.7m
Cost per Finding
Standard cost per accepted finding · Lower is better
Kimi K2.7
Qwen3.8 27B
Kimi K3
GLM-5.2
Grok 4.5
GPT-5.5
GPT-5.6 Luna
GPT-5.6 Terra
Opus 4.8
Nemotron 3 Ultra
Sonnet 5
GPT-5.6 Sol
$14 standard cost ÷ 14 accepted findings= $0.99
$51 standard cost ÷ 45 accepted findings= $1.14
$47 standard cost ÷ 41 accepted findings= $1.15
$38 standard cost ÷ 31 accepted findings= $1.22
$122 standard cost ÷ 51 accepted findings= $2.39
$120 standard cost ÷ 45 accepted findings= $2.67
$158 standard cost ÷ 47 accepted findings= $3.36
$215 standard cost ÷ 49 accepted findings= $4.38
$380 standard cost ÷ 68 accepted findings= $5.59
$92 standard cost ÷ 16 accepted findings= $5.77
$374 standard cost ÷ 54 accepted findings= $6.93
$1093 standard cost ÷ 87 accepted findings= $12.56
$0.00$7.50$15.00
$0.99
$1.14
$1.15
$1.22
$2.39
$2.67
$3.36
$4.38
$5.59
$5.77
$6.93
$12.56
Cost Efficiency
Accepted findings per dollar · Higher is better
Kimi K2.7
Qwen3.8 27B
Kimi K3
GLM-5.2
Grok 4.5
GPT-5.5
GPT-5.6 Luna
GPT-5.6 Terra
Opus 4.8
Nemotron 3 Ultra
Sonnet 5
GPT-5.6 Sol
14 accepted findings ÷ $14 standard cost= 1.01
45 accepted findings ÷ $51 standard cost= 0.88
41 accepted findings ÷ $47 standard cost= 0.87
31 accepted findings ÷ $38 standard cost= 0.82
51 accepted findings ÷ $122 standard cost= 0.42
45 accepted findings ÷ $120 standard cost= 0.37
47 accepted findings ÷ $158 standard cost= 0.30
49 accepted findings ÷ $215 standard cost= 0.23
68 accepted findings ÷ $380 standard cost= 0.18
16 accepted findings ÷ $92 standard cost= 0.17
54 accepted findings ÷ $374 standard cost= 0.14
87 accepted findings ÷ $1093 standard cost= 0.08
0.000.751.50
1.01
0.88
0.87
0.82
0.42
0.37
0.30
0.23
0.18
0.17
0.14
0.08
Token Efficiency
Accepted findings per 1M tokens · Higher is better
Qwen3.8 27B
Grok 4.5
Kimi K3
Opus 4.8
Kimi K2.7
GLM-5.2
GPT-5.6 Terra
GPT-5.5
GPT-5.6 Sol
GPT-5.6 Luna
Sonnet 5
Nemotron 3 Ultra
45 accepted findings ÷ 6.8M tokens= 6.57
51 accepted findings ÷ 49.6M tokens= 1.03
41 accepted findings ÷ 40.0M tokens= 1.02
68 accepted findings ÷ 77.8M tokens= 0.87
14 accepted findings ÷ 17.3M tokens= 0.81
31 accepted findings ÷ 42.7M tokens= 0.73
49 accepted findings ÷ 80.3M tokens= 0.61
45 accepted findings ÷ 88.0M tokens= 0.51
87 accepted findings ÷ 224.5M tokens= 0.39
47 accepted findings ÷ 127.8M tokens= 0.37
54 accepted findings ÷ 173.1M tokens= 0.31
16 accepted findings ÷ 261.5M tokens= 0.06
0.003.507.00
6.57
1.03
1.02
0.87
0.81
0.73
0.61
0.51
0.39
0.37
0.31
0.06
Full standings
Fig. 1
Model leaderboard
Every model at its best run across all reasoning efforts, ranked by severity-weighted score. Click any column header to re-sort.
Model
Best-run effort
GPT-5.6 Sol
X-High
61
2h 19m
87
40
$1,093
224.5M
0.080
0.387
Grok 4.5
High
52
5h 3m
51
26
$122
49.6M
0.419
1.028
Claude Opus 4.8
Max
45
1h 50m
68
38
$380
77.8M
0.179
0.875
GPT-5.6 Luna
X-High
43
1h 59m
47
21
$158
127.8M
0.297
0.368
Kimi K3
Default
42
49m
41
22
$47
40.0M
0.872
1.025
GPT-5.5
X-High
41
1h 19m
45
19
$120
88.0M
0.375
0.511
GPT-5.6 Terra
X-High
36
1h 6m
49
19
$215
80.3M
0.228
0.610
GLM-5.2
X-High
34
1h 40m
31
12
$38
42.7M
0.821
0.726
Claude Sonnet 5
Max
34
4h 30m
54
31
$374
173.1M
0.144
0.312
Qwen3.8 27B
Default
28
59m
45
9
$51
6.8M
0.876
6.569
Kimi K2.7
Default
24
2h 32m
14
2
$14
17.3M
1.005
0.809
Nemotron 3 Ultra
High
18
5h 3m
16
2
$92
261.5M
0.173
0.061
Fig. 2
Accepted findings by severity
Validator-accepted findings per model, sorted by total volume. Each model keeps its own hue; darker shades are higher-severity findings. Hover a bar for the low / medium / high split.
Hue = model · shade = severity
High / critical
Medium
Low
GPT-5.6 SolX-HighOpenAI
Opus 4.8MaxAnthropic
Sonnet 5MaxAnthropic
Grok 4.5HighxAI
GPT-5.6 TerraX-HighOpenAI
GPT-5.6 LunaX-HighOpenAI
GPT-5.5X-HighOpenAI
Qwen3.8 27BDefaultAlibaba
Kimi K3DefaultMoonshot AI
GLM-5.2X-HighZ.ai
Nemotron 3 UltraHighNVIDIA
Kimi K2.7DefaultMoonshot AI
87
high40
medium25
low22
87 accepted findings
68
high38
medium13
low17
68 accepted findings
54
high31
medium14
low9
54 accepted findings
51
high26
medium15
low10
51 accepted findings
49
high19
medium18
low12
49 accepted findings
47
high21
medium11
low15
47 accepted findings
45
high19
medium15
low11
45 accepted findings
45
high9
medium16
low20
45 accepted findings
41
high22
medium9
low10
41 accepted findings
31
high12
medium10
low9
31 accepted findings
16
high2
medium7
low7
16 accepted findings
14
high2
medium6
low6
14 accepted findings
0255075100
Put your security on autopilot
First Deployment
15 min
Agents Running
24/7
More Coverage
100x
Hours Saved
1,000+
No UI neededWe have a full MCP that controls our entire platform without a UI.