Skip to main content
← Back to Blog

Introducing NexBench: MindFort's Internal Model Evaluation

Akul Gupta, Co-Founder & CTO at MindFort

Written by

Akul Gupta

2026-07-22·Updated 2026-08-28·11 min read

Attackers now find and exploit 0-days faster than defenders can keep up, which is why continuous pen-testing against attack surfaces that change often has become table stakes. Better models have made that testing faster, but running the same agents has become token- and cost-inefficient: they run for 8-16 hours at a time, sometimes every day. That leaves a gap between how fast attackers find 0-days and how fast teams can fix them. We built MindFort to make security efficient, performant, and abundant.

This led us to become interested in how efficient agents were when executing offensive security tasks. For us to be able to create these offensive security agents, we needed them to be capable of running for multiple hours at a time, continuously through the year, which led us to create an internal benchmark measuring models' ability to balance token efficiency and performance.

Current cybersecurity evals share one flaw: they are boxed in. Real environments are messy and ambiguous, with layered vulnerabilities that can be chained together in different ways. They come in different deployment and authentication shapes and different sizes, which makes it hard for agents to even work their way in, and that is half the battle. So we built our own, modeled on the thousands of real environments we have watched our agents work through.

Introducing, NexBench. NexBench demonstrates the real world environments that agents find themselves in. By running our agents on NexBench, we aim to understand which models can find the most validated vulnerabilities, have the best price to performance ratios, and exhibit long running coherence and ability.

Our eval provides signals of a model's ability to write, reason and execute real exploits against a target, while measuring their price to performance ratio and how well the harness handles the complex environment. In our research, we have found that evaluations of agents in existing benchmarks far outperform the same models/agents in real world environments.

Setup

Our setup for NexBench uses a modified version of our harness that allows for our agent system to be more portable. The full test range is spun up in an isolated container, and the agents have full access to it, mimicking a web accessible environment.

Agents are scored on their ability to discover vulnerabilities in the following 3 severity levels: low, medium, and high. We dropped critical as its own tier as high vs. critical judgments became subjective and inconsistent over time, so we folded both into high and scored coarsely. Each finding was then re-validated by a separate judge agent, which scored the finding consistently via CVSS, and then re-exploited to confirm that the agent finding held. The vulnerabilities in the test range varied from trivial and low severity, to highly complex multi-stage vulnerabilities in atypical locations. One interesting observation that our bench captures is reviewing how an agent weaves vulnerabilities together. Since the benchmark does not have a concept of “100%” completion but rather non-deterministic scoring that follows a predictable trajectory, we are able to see how effective and how creative agents are in surprising us with newer and better ways to exploit multiple vulnerabilities together. Again, mirroring closely how in the wild agents would interact with complex environments.

Because of this, we score non-deterministically, and take the best score across multiple runs at every reasoning effort available for each model capped at a 5 hour runtime, per agent. We chose a cap time of 5 hours, as it marks a reasonable amount of time for daily pen testing.

We report 4 metrics: findings, validated findings, tokens used, and estimated token cost:

  • A finding is validated if the judge agent is able to successfully reproduce the exploit, or fails otherwise.
  • A model's score is a sum of all validated findings across a single run of the model in the harness.

Results

Fig. 1

Accepted findings by severity

Validator-accepted findings per model, sorted by total volume. Each model keeps its own hue; darker shades are higher-severity findings. Hover any bar for the exact low / medium / high breakdown.

Hue = model · shade = severity
High / critical
Medium
Low
GPT-5.6 SolX-HighOpenAI
Opus 4.8MaxAnthropic
Sonnet 5MaxAnthropic
Grok 4.5HighxAI
GPT-5.6 TerraX-HighOpenAI
GPT-5.6 LunaX-HighOpenAI
GPT-5.5X-HighOpenAI
Qwen3.8 27BDefaultAlibaba
Kimi K3DefaultMoonshot AI
GLM-5.2X-HighZ.ai
Nemotron 3 UltraHighNVIDIA
Kimi K2.7DefaultMoonshot AI
87
high40
medium25
low22
87 accepted findings
68
high38
medium13
low17
68 accepted findings
54
high31
medium14
low9
54 accepted findings
51
high26
medium15
low10
51 accepted findings
49
high19
medium18
low12
49 accepted findings
47
high21
medium11
low15
47 accepted findings
45
high19
medium15
low11
45 accepted findings
45
high9
medium16
low20
45 accepted findings
41
high22
medium9
low10
41 accepted findings
31
high12
medium10
low9
31 accepted findings
16
high2
medium7
low7
16 accepted findings
14
high2
medium6
low6
14 accepted findings
0255075100

On raw performance, GPT-5.6 performed best with 87 validated findings on its best run. Opus 4.8 came in second with 68 total findings (Figs. 1 and 2).

Note: The leaderboards below rank all twelve models by the highlighted column, each shown at its best run across every reasoning effort. No model or effort is dropped from the evaluation. Click any column header to re-sort.

Fig. 2

Leaderboard, sorted by accepted findings

All 12 models, ranked by total validator-accepted findings summed across every run. The sorted column is highlighted in orange; click any header to re-sort.

ModelBest-run effort
GPT-5.6 SolX-High612h 19m8740$1,093224.5M0.0800.387
Claude Opus 4.8Max451h 50m6838$38077.8M0.1790.875
Claude Sonnet 5Max344h 30m5431$374173.1M0.1440.312
Grok 4.5High525h 3m5126$12249.6M0.4191.028
GPT-5.6 TerraX-High361h 6m4919$21580.3M0.2280.610
GPT-5.6 LunaX-High431h 59m4721$158127.8M0.2970.368
GPT-5.5X-High411h 19m4519$12088.0M0.3750.511
Qwen3.8 27BDefault2859m459$516.8M0.8766.569
Kimi K3Default4249m4122$4740.0M0.8721.025
GLM-5.2X-High341h 40m3112$3842.7M0.8210.726
Nemotron 3 UltraHigh185h 3m162$92261.5M0.1730.061
Kimi K2.7Default242h 32m142$1417.3M1.0050.809

Where it starts getting interesting is when we start to implement scoring, assigned based on severity of the finding. We assigned a score of 1 to low findings, 2 for medium findings, and 3 for high/critical findings. In general, critical and high findings were harder to find, since they required chaining multiple vulnerabilities together, which makes this metric a good read on overall performance in offensive security workflows. Each score was calculated as shown below and the greatest score out of 3 runs per model was used.

(3 * #_of_high_vulnerabilities) + (2 * #_of_medium_vulnerabilities) + (1 * #_of_low_vulnerabilities)

When weighting severity against each finding, Grok 4.5 did impressively well, scoring just below GPT-5.6 Sol and beating out Opus 4.8. Grok did take the full runtime to get there, more than double that of Opus, Luna, Sol, and others, but within the time a pen test should reasonably take it outperformed most models at a very efficient token rate. Another fast model was Kimi K3, which scored a 42 in 49 minutes (Fig. 3).

Fig. 3

Leaderboard, sorted by best score

All 12 models, ranked by best severity-weighted score, each model's strongest single run across all reasoning efforts. The sorted column is highlighted in orange; click any header to re-sort.

ModelBest-run effort
GPT-5.6 SolX-High612h 19m8740$1,093224.5M0.0800.387
Grok 4.5High525h 3m5126$12249.6M0.4191.028
Claude Opus 4.8Max451h 50m6838$38077.8M0.1790.875
GPT-5.6 LunaX-High431h 59m4721$158127.8M0.2970.368
Kimi K3Default4249m4122$4740.0M0.8721.025
GPT-5.5X-High411h 19m4519$12088.0M0.3750.511
GPT-5.6 TerraX-High361h 6m4919$21580.3M0.2280.610
GLM-5.2X-High341h 40m3112$3842.7M0.8210.726
Claude Sonnet 5Max344h 30m5431$374173.1M0.1440.312
Qwen3.8 27BDefault2859m459$516.8M0.8766.569
Kimi K2.7Default242h 32m142$1417.3M1.0050.809
Nemotron 3 UltraHigh185h 3m162$92261.5M0.1730.061
Fig. 4

Cost per accepted finding

Total standard cost divided by accepted findings, cheapest model first. Hover a bar for the calculation.

Kimi K2.7DefaultMoonshot AI
Qwen3.8 27BDefaultAlibaba
Kimi K3DefaultMoonshot AI
GLM-5.2X-HighZ.ai
Grok 4.5HighxAI
GPT-5.5X-HighOpenAI
GPT-5.6 LunaX-HighOpenAI
GPT-5.6 TerraX-HighOpenAI
Opus 4.8MaxAnthropic
Nemotron 3 UltraHighNVIDIA
Sonnet 5MaxAnthropic
GPT-5.6 SolX-HighOpenAI
$14 standard cost ÷ 14 accepted findings= $0.99
$51 standard cost ÷ 45 accepted findings= $1.14
$47 standard cost ÷ 41 accepted findings= $1.15
$38 standard cost ÷ 31 accepted findings= $1.22
$122 standard cost ÷ 51 accepted findings= $2.39
$120 standard cost ÷ 45 accepted findings= $2.67
$158 standard cost ÷ 47 accepted findings= $3.36
$215 standard cost ÷ 49 accepted findings= $4.38
$380 standard cost ÷ 68 accepted findings= $5.59
$92 standard cost ÷ 16 accepted findings= $5.77
$374 standard cost ÷ 54 accepted findings= $6.93
$1093 standard cost ÷ 87 accepted findings= $12.56
$0.00$3.75$7.50$11.25$15.00
$0.99
$1.14
$1.15
$1.22
$2.39
$2.67
$3.36
$4.38
$5.59
$5.77
$6.93
$12.56

GPT-5.6 Sol also scored as the most expensive model when looking at cost to accepted finding in our evals, with Sonnet coming in second (Fig. 4). Surprisingly, Kimi K2.7, K3, and GLM-5.2 all scored the best in terms of cost per finding, which we will explore later in regards to overall efficiency.

Looking at token and cost efficiency, we can then deduce which models were able to make the most of token and dollar cost. Our x-axis measures validated points per USD, cost efficiency as a ratio, with our y-axis measuring performance per token, token efficiency. Kimi K2.7 ended up having the best dollar to finding ratio, but token use inefficiency ended up weighing down overall results (Fig. 5). Among the hosted models, Grok 4.5 performed in the inverse, with strong token efficiency, but was more expensive than Kimi K2.7, along with GLM-5.2. The standout on token efficiency, though, is Qwen3.8 27B, run locally on our own hardware: it validated 45 findings on just 6.85M total tokens. That works out to roughly 6.6 findings per million tokens, several times the rate of any hosted model, at about $1.14 per validated finding.

Kimi K3 ended up scoring similarly to Grok 4.5 in terms of token efficiency (1.03 findings per 1M tokens vs 1.02 findings per 1M tokens) but scored significantly better in terms of dollar efficiency, at more than double the findings per USD, making it a really cost efficient choice.

Fig. 5

Cost vs. token efficiency

Accepted findings per dollar (x) against accepted findings per million tokens (y), one point per model. Upper-right is best, hover any point for details.

0.001.402.804.205.607.000.000.240.480.720.961.20Findings per $Findings per 1M tokensGPT-5.6 SolGrok 4.5Claude Opus 4.8GPT-5.6 LunaKimi K3GPT-5.5GPT-5.6 TerraGLM-5.2Claude Sonnet 5Kimi K2.7Nemotron 3 UltraQwen3.8 27BMost cost & token efficient

Modeling score to cost in dollars, we can model the pareto frontier of the above models, based on available budgets. Grok on high ended up scoring best, with an additional $300 in spend needed by Sol on high effort to surpass Grok's performance on total points discovered (Fig. 6).

Fig. 6

Score vs. standard cost

Every model-effort cell. The dashed orange line marks the Pareto frontier, the highest score reached at or below each cost.

01326395265$0$110$220$330$440$550Standard cost ($)Score
Claude Opus 4.8
Claude Sonnet 5
GLM-5.2
GPT-5.5
GPT-5.6 Sol
GPT-5.6 Luna
GPT-5.6 Terra
Grok 4.5
Kimi K2.7
Kimi K3
Nemotron 3 Ultra
Qwen3.8 27B

A similar pattern can be observed when modeling the pareto frontier of token cost to findings, with each Grok model appearing along the frontier, no matter the reasoning effort. Additionally, K3 scored on the pareto frontier for both cost and token efficiency (Fig. 7).

Fig. 7

Score vs. total tokens

The same model-effort cells on a log-scale x-axis, token spend ranges from roughly 1.3M to over 150M across runs.

013263952651M10M100MTotal tokens (log scale)Score
Claude Opus 4.8
Claude Sonnet 5
GLM-5.2
GPT-5.5
GPT-5.6 Sol
GPT-5.6 Luna
GPT-5.6 Terra
Grok 4.5
Kimi K2.7
Kimi K3
Nemotron 3 Ultra
Qwen3.8 27B

Conclusion

As more software gets shipped, teams need to start thinking about how software can become self-securing. Long-running agents are the way forward, and to power them we have to figure out which models can do the work at an efficient cost without giving up performance. Open source models are already gaining fast on closed source ones like Opus and GPT, and as powerful models get cheaper, attackers get access to stronger weapons. Defenders have to pick up those same resources and catch vulnerabilities continuously, before an attacker gets the chance to exploit them.

About the author

Akul Gupta, Co-Founder & CTO at MindFort

Akul Gupta

Co-Founder & CTO · MindFort

AI researcher focusing on LLMs in cybersecurity. Red-teamed models for OpenAI and Anthropic as part of their safety programs. Published multiple conference papers. M.S. Computer Science, UIUC.

Put your security on autopilot

First Deployment

15 min

Agents Running

24/7

More Coverage

100x

Hours Saved

1,000+