Skip to main content
← Back to Research

NexBench

NexBench provides signals of a model's ability to write, reason and execute real exploits against a real target in a messy environment, while measuring price to performance ratio, token efficiency, and more.

Last updated · 12 models

Read more on how we built NexBench

Highlights

Best Score

Best severity-weighted score · Higher is better

GPT-5.6 Sol
Grok 4.5
Opus 4.8
GPT-5.6 Luna
Kimi K3
GPT-5.5
GPT-5.6 Terra
GLM-5.2
Sonnet 5
Qwen3.8 27B
Kimi K2.7
Nemotron 3 Ultra
Best severity-weighted score across every reasoning-effort run= 61
Best severity-weighted score across every reasoning-effort run= 52
Best severity-weighted score across every reasoning-effort run= 45
Best severity-weighted score across every reasoning-effort run= 43
Best severity-weighted score across every reasoning-effort run= 42
Best severity-weighted score across every reasoning-effort run= 41
Best severity-weighted score across every reasoning-effort run= 36
Best severity-weighted score across every reasoning-effort run= 34
Best severity-weighted score across every reasoning-effort run= 34
Best severity-weighted score across every reasoning-effort run= 28
Best severity-weighted score across every reasoning-effort run= 24
Best severity-weighted score across every reasoning-effort run= 18
03365
61
52
45
43
42
41
36
34
34
28
24
18

Best Runtime

Best-run wall time · Lower is better

Kimi K3
Qwen3.8 27B
GPT-5.6 Terra
GPT-5.5
GLM-5.2
Opus 4.8
GPT-5.6 Luna
GPT-5.6 Sol
Kimi K2.7
Sonnet 5
Grok 4.5
Nemotron 3 Ultra
Wall-clock run time of the best-scoring attempt= 49.0m
Wall-clock run time of the best-scoring attempt= 59.0m
Wall-clock run time of the best-scoring attempt= 66.2m
Wall-clock run time of the best-scoring attempt= 78.8m
Wall-clock run time of the best-scoring attempt= 100.0m
Wall-clock run time of the best-scoring attempt= 110.5m
Wall-clock run time of the best-scoring attempt= 119.0m
Wall-clock run time of the best-scoring attempt= 139.1m
Wall-clock run time of the best-scoring attempt= 152.0m
Wall-clock run time of the best-scoring attempt= 269.7m
Wall-clock run time of the best-scoring attempt= 302.5m
Wall-clock run time of the best-scoring attempt= 302.7m
0.0m175.0m350.0m
49.0m
59.0m
66.2m
78.8m
100.0m
110.5m
119.0m
139.1m
152.0m
269.7m
302.5m
302.7m

Cost per Finding

Standard cost per accepted finding · Lower is better

Kimi K2.7
Qwen3.8 27B
Kimi K3
GLM-5.2
Grok 4.5
GPT-5.5
GPT-5.6 Luna
GPT-5.6 Terra
Opus 4.8
Nemotron 3 Ultra
Sonnet 5
GPT-5.6 Sol
$14 standard cost ÷ 14 accepted findings= $0.99
$51 standard cost ÷ 45 accepted findings= $1.14
$47 standard cost ÷ 41 accepted findings= $1.15
$38 standard cost ÷ 31 accepted findings= $1.22
$122 standard cost ÷ 51 accepted findings= $2.39
$120 standard cost ÷ 45 accepted findings= $2.67
$158 standard cost ÷ 47 accepted findings= $3.36
$215 standard cost ÷ 49 accepted findings= $4.38
$380 standard cost ÷ 68 accepted findings= $5.59
$92 standard cost ÷ 16 accepted findings= $5.77
$374 standard cost ÷ 54 accepted findings= $6.93
$1093 standard cost ÷ 87 accepted findings= $12.56
$0.00$7.50$15.00
$0.99
$1.14
$1.15
$1.22
$2.39
$2.67
$3.36
$4.38
$5.59
$5.77
$6.93
$12.56

Cost Efficiency

Accepted findings per dollar · Higher is better

Kimi K2.7
Qwen3.8 27B
Kimi K3
GLM-5.2
Grok 4.5
GPT-5.5
GPT-5.6 Luna
GPT-5.6 Terra
Opus 4.8
Nemotron 3 Ultra
Sonnet 5
GPT-5.6 Sol
14 accepted findings ÷ $14 standard cost= 1.01
45 accepted findings ÷ $51 standard cost= 0.88
41 accepted findings ÷ $47 standard cost= 0.87
31 accepted findings ÷ $38 standard cost= 0.82
51 accepted findings ÷ $122 standard cost= 0.42
45 accepted findings ÷ $120 standard cost= 0.37
47 accepted findings ÷ $158 standard cost= 0.30
49 accepted findings ÷ $215 standard cost= 0.23
68 accepted findings ÷ $380 standard cost= 0.18
16 accepted findings ÷ $92 standard cost= 0.17
54 accepted findings ÷ $374 standard cost= 0.14
87 accepted findings ÷ $1093 standard cost= 0.08
0.000.751.50
1.01
0.88
0.87
0.82
0.42
0.37
0.30
0.23
0.18
0.17
0.14
0.08

Token Efficiency

Accepted findings per 1M tokens · Higher is better

Qwen3.8 27B
Grok 4.5
Kimi K3
Opus 4.8
Kimi K2.7
GLM-5.2
GPT-5.6 Terra
GPT-5.5
GPT-5.6 Sol
GPT-5.6 Luna
Sonnet 5
Nemotron 3 Ultra
45 accepted findings ÷ 6.8M tokens= 6.57
51 accepted findings ÷ 49.6M tokens= 1.03
41 accepted findings ÷ 40.0M tokens= 1.02
68 accepted findings ÷ 77.8M tokens= 0.87
14 accepted findings ÷ 17.3M tokens= 0.81
31 accepted findings ÷ 42.7M tokens= 0.73
49 accepted findings ÷ 80.3M tokens= 0.61
45 accepted findings ÷ 88.0M tokens= 0.51
87 accepted findings ÷ 224.5M tokens= 0.39
47 accepted findings ÷ 127.8M tokens= 0.37
54 accepted findings ÷ 173.1M tokens= 0.31
16 accepted findings ÷ 261.5M tokens= 0.06
0.003.507.00
6.57
1.03
1.02
0.87
0.81
0.73
0.61
0.51
0.39
0.37
0.31
0.06

Full standings

Fig. 1

Model leaderboard

Every model at its best run across all reasoning efforts, ranked by severity-weighted score. Click any column header to re-sort.

ModelBest-run effort
GPT-5.6 SolX-High612h 19m8740$1,093224.5M0.0800.387
Grok 4.5High525h 3m5126$12249.6M0.4191.028
Claude Opus 4.8Max451h 50m6838$38077.8M0.1790.875
GPT-5.6 LunaX-High431h 59m4721$158127.8M0.2970.368
Kimi K3Default4249m4122$4740.0M0.8721.025
GPT-5.5X-High411h 19m4519$12088.0M0.3750.511
GPT-5.6 TerraX-High361h 6m4919$21580.3M0.2280.610
GLM-5.2X-High341h 40m3112$3842.7M0.8210.726
Claude Sonnet 5Max344h 30m5431$374173.1M0.1440.312
Qwen3.8 27BDefault2859m459$516.8M0.8766.569
Kimi K2.7Default242h 32m142$1417.3M1.0050.809
Nemotron 3 UltraHigh185h 3m162$92261.5M0.1730.061
Fig. 2

Accepted findings by severity

Validator-accepted findings per model, sorted by total volume. Each model keeps its own hue; darker shades are higher-severity findings. Hover a bar for the low / medium / high split.

Hue = model · shade = severity
High / critical
Medium
Low
GPT-5.6 SolX-HighOpenAI
Opus 4.8MaxAnthropic
Sonnet 5MaxAnthropic
Grok 4.5HighxAI
GPT-5.6 TerraX-HighOpenAI
GPT-5.6 LunaX-HighOpenAI
GPT-5.5X-HighOpenAI
Qwen3.8 27BDefaultAlibaba
Kimi K3DefaultMoonshot AI
GLM-5.2X-HighZ.ai
Nemotron 3 UltraHighNVIDIA
Kimi K2.7DefaultMoonshot AI
87
high40
medium25
low22
87 accepted findings
68
high38
medium13
low17
68 accepted findings
54
high31
medium14
low9
54 accepted findings
51
high26
medium15
low10
51 accepted findings
49
high19
medium18
low12
49 accepted findings
47
high21
medium11
low15
47 accepted findings
45
high19
medium15
low11
45 accepted findings
45
high9
medium16
low20
45 accepted findings
41
high22
medium9
low10
41 accepted findings
31
high12
medium10
low9
31 accepted findings
16
high2
medium7
low7
16 accepted findings
14
high2
medium6
low6
14 accepted findings
0255075100

Put your security on autopilot

First Deployment

15 min

Agents Running

24/7

More Coverage

100x

Hours Saved

1,000+