Skip to main content
← Back to Research

Introducing NexBench: MindFort's Internal Model Evaluation

Today, we're introducing NexBench, our internal benchmark for measuring which models can lead an offensive-security harness while balancing validated findings, token efficiency, and cost across real-world environments.

Akul Gupta, Co-Founder & CTO at MindFort

Written by

Akul Gupta

2026-07-22·11 min read

In a time where 0-days can be found and exploited faster by attackers, it's more important than ever to run continuous pen-testing, against all attack surfaces that change often. With rapid advancements in model performance, this process has gotten faster, but running those same agents has also become token and cost inefficient, as they are running for 8-16 hours at a time, sometimes every day. There now becomes a disparity between the security risk of attackers finding these 0-days and the remediation needed to fix them. Efficient, performant, and abundant security became our end goal while creating MindFort.

This led us to become interested in how efficient agents were when executing offensive security tasks. For us to be able to create these offensive security agents, we needed them to be capable of running for multiple hours at a time, continuously through the year, which led us to create an internal benchmark measuring models' ability to balance token efficiency and performance.

Current evals for cybersecurity have a fatal flaw: they are boxed in. Real world environments are messy, ambiguous, and contain layered vulnerabilities that can be chained together in different ways. They also have various deployment and authentication shapes, as well as various sizes which make it challenging for agents to actually navigate through, which is half the battle. That is why we created our own, modeled off the thousands of real world environments we have observed our agents navigate through.

Introducing, NexBench. NexBench demonstrates the real world environments that agents find themselves in. By running our agents on NexBench, we aim to understand which models can find the most validated vulnerabilities, have the best price to performance ratios, and exhibit long running coherence and ability.

Our eval provides signals of a model's ability to write, reason and execute real exploits against a target, while measuring their price to performance ratio and how well the harness handles the complex environment. In our research, we have found that evaluations of agents in existing benchmarks far outperform the same models/agents in real world environments.

Setup

Our setup for NexBench uses a modified version of our harness that allows for our agent system to be more portable. The full test range is spun up in an isolated container, and the agents have full access to it, mimicking a web accessible environment.

Agents are scored on their ability to discover vulnerabilities in the following 3 severity levels: low, medium, and high. We dropped critical as its own tier as high vs. critical judgments became subjective and inconsistent over time, so we folded both into high and scored coarsely. Each finding was then re-validated by a separate judge agent, which scored the finding consistently via CVSS, and then re-exploited to confirm that the agent finding held. The vulnerabilities in the test range varied from trivial and low severity, to highly complex multi-stage vulnerabilities in atypical locations. One interesting observation that our bench captures is reviewing how an agent weaves vulnerabilities together. Since the benchmark does not have a concept of “100%” completion but rather non-deterministic scoring that follows a predictable trajectory, we are able to see how effective and how creative agents are in surprising us with newer and better ways to exploit multiple vulnerabilities together. Again, mirroring closely how in the wild agents would interact with complex environments.

Because of this, we score non-deterministically, and take the best score across multiple runs at every reasoning effort available for each model capped at a 5 hour runtime, per agent. We chose a cap time of 5 hours, as it marks a reasonable amount of time for daily pentesting.

We report 4 metrics: findings, validated findings, tokens used, and estimated token cost:

  • A finding is validated if the judge agent is able to successfully reproduce the exploit, or fails otherwise.
  • A model's score is a sum of all validated findings across a single run of the model in the harness.

Results

Fig. 1

Accepted findings by severity

Validator-accepted findings per model, sorted by total volume. Each model keeps its own hue; darker shades are higher-severity findings. Hover any bar for the exact low / medium / high breakdown.

Hue = model · shade = severity
High / critical
Medium
Low
GPT-5.6 SolX-HighOpenAI
Opus 4.8MaxAnthropic
Sonnet 5MaxAnthropic
Grok 4.5HighxAI
GPT-5.6 TerraX-HighOpenAI
GPT-5.6 LunaX-HighOpenAI
GPT-5.5X-HighOpenAI
Kimi K3DefaultMoonshot AI
GLM-5.2X-HighZ.ai
Nemotron 3 UltraHighNVIDIA
Kimi K2.7DefaultMoonshot AI
87
high40
medium25
low22
87 accepted findings
68
high38
medium13
low17
68 accepted findings
54
high31
medium14
low9
54 accepted findings
51
high26
medium15
low10
51 accepted findings
49
high19
medium18
low12
49 accepted findings
47
high21
medium11
low15
47 accepted findings
45
high19
medium15
low11
45 accepted findings
41
high22
medium9
low10
41 accepted findings
31
high12
medium10
low9
31 accepted findings
16
high2
medium7
low7
16 accepted findings
14
high2
medium6
low6
14 accepted findings
0255075100

On raw performance, GPT-5.6 performed best with 87 validated findings on its best run. Opus 4.8 came in second with 68 total findings (Figs. 1 and 2).

Note: The leaderboards below rank all eleven models by the highlighted column, each shown at its best run across every reasoning effort. No model or effort is dropped from the evaluation. Click any column header to re-sort.

Fig. 2

Leaderboard, sorted by accepted findings

All 11 models, ranked by total validator-accepted findings summed across every run. The sorted column is highlighted in orange; click any header to re-sort.

ModelBest-run effort
GPT-5.6 SolX-High612h 19m8740$1,093224.5M0.0800.387
Claude Opus 4.8Max451h 50m6838$38077.8M0.1790.875
Claude Sonnet 5Max344h 30m5431$374173.1M0.1440.312
Grok 4.5High525h 3m5126$12249.6M0.4191.028
GPT-5.6 TerraX-High361h 6m4919$21580.3M0.2280.610
GPT-5.6 LunaX-High431h 59m4721$158127.8M0.2970.368
GPT-5.5X-High411h 19m4519$12088.0M0.3750.511
Kimi K3Default4249m4122$4740.0M0.8721.025
GLM-5.2X-High341h 40m3112$3842.7M0.8210.726
Nemotron 3 UltraHigh185h 3m162$92261.5M0.1730.061
Kimi K2.7Default242h 32m142$1417.3M1.0050.809

Where it starts getting interesting is when we start to implement scoring, assigned based on severity of the finding. We assigned a score of 1 to low findings, 2 for medium findings, and 3 for high/critical findings. In general, critical and high findings were harder to find, since they required chaining multiple vulnerabilities together, which makes this metric a showcase of overall performance in offensive security workflows. Each score was calculated as shown below and the greatest score out of 3 runs per model was used.

(3 * #_of_high_vulnerabilities) + (2 * #_of_medium_vulnerabilities) + (1 * #_of_low_vulnerabilities)

When weighting severity against each finding, Grok 4.5 did impressively well, scoring just below GPT-5.6 Sol, beating out Opus 4.8. It is important to note that Grok did take the full runtime to be able to produce these results, more than double of Opus, Luna, Sol, etc, but given a reasonable amount of time where a pentest should have been completed, was able to outperform most models at a very efficient token rate. Another highly time efficient model was Kimi K3, scoring a 42 in 49 minutes (Fig. 3).

Fig. 3

Leaderboard, sorted by best score

All 11 models, ranked by best severity-weighted score — each model's strongest single run across all reasoning efforts. The sorted column is highlighted in orange; click any header to re-sort.

ModelBest-run effort
GPT-5.6 SolX-High612h 19m8740$1,093224.5M0.0800.387
Grok 4.5High525h 3m5126$12249.6M0.4191.028
Claude Opus 4.8Max451h 50m6838$38077.8M0.1790.875
GPT-5.6 LunaX-High431h 59m4721$158127.8M0.2970.368
Kimi K3Default4249m4122$4740.0M0.8721.025
GPT-5.5X-High411h 19m4519$12088.0M0.3750.511
GPT-5.6 TerraX-High361h 6m4919$21580.3M0.2280.610
GLM-5.2X-High341h 40m3112$3842.7M0.8210.726
Claude Sonnet 5Max344h 30m5431$374173.1M0.1440.312
Kimi K2.7Default242h 32m142$1417.3M1.0050.809
Nemotron 3 UltraHigh185h 3m162$92261.5M0.1730.061
Fig. 4

Cost per accepted finding

Total standard cost divided by accepted findings, cheapest model first. Hover a bar for the calculation.

Kimi K2.7DefaultMoonshot AI
Kimi K3DefaultMoonshot AI
GLM-5.2X-HighZ.ai
Grok 4.5HighxAI
GPT-5.5X-HighOpenAI
GPT-5.6 LunaX-HighOpenAI
GPT-5.6 TerraX-HighOpenAI
Opus 4.8MaxAnthropic
Nemotron 3 UltraHighNVIDIA
Sonnet 5MaxAnthropic
GPT-5.6 SolX-HighOpenAI
$0.99
$14 standard cost ÷ 14 accepted findings= $0.99
$1.15
$47 standard cost ÷ 41 accepted findings= $1.15
$1.22
$38 standard cost ÷ 31 accepted findings= $1.22
$2.39
$122 standard cost ÷ 51 accepted findings= $2.39
$2.67
$120 standard cost ÷ 45 accepted findings= $2.67
$3.36
$158 standard cost ÷ 47 accepted findings= $3.36
$4.38
$215 standard cost ÷ 49 accepted findings= $4.38
$5.59
$380 standard cost ÷ 68 accepted findings= $5.59
$5.77
$92 standard cost ÷ 16 accepted findings= $5.77
$6.93
$374 standard cost ÷ 54 accepted findings= $6.93
$12.56
$1093 standard cost ÷ 87 accepted findings= $12.56
$0.00$3.75$7.50$11.25$15.00

GPT-5.6 Sol also scored as the most expensive model when looking at cost to accepted finding in our evals, with Sonnet coming in second (Fig. 4). Surprisingly, Kimi K2.7, K3, and GLM-5.2 all scored the best in terms of cost per finding, which we will explore later in regards to overall efficiency.

Looking at token and cost efficiency, we can then deduce which models were able to make the most of token and dollar cost. Our x-axis measures validated points per USD, cost efficiency as a ratio, with our y-axis measuring performance per token, token efficiency. Kimi K2.7 ended up having the best dollar to finding ratio, but token use inefficiency ended up weighing down overall results (Fig. 5). Grok 4.5 performed in the inverse, with the best token efficiency, but was more expensive than Kimi K2.7, along with GLM-5.2.

Kimi K3 ended up scoring similarly to Grok 4.5 in terms of token efficiency (1.03 findings per 1M tokens vs 1.02 findings per 1M tokens) but scored significantly better in terms of dollar efficiency, at more than double the findings per USD, making it a really cost efficient choice.

Fig. 5

Cost vs. token efficiency

Accepted findings per dollar (x) against accepted findings per million tokens (y), one point per model. Upper-right is best, hover any point for details.

0.000.300.600.901.201.500.000.240.480.720.961.20Findings per $Findings per 1M tokensGPT-5.6 SolGrok 4.5Claude Opus 4.8GPT-5.6 LunaKimi K3GPT-5.5GPT-5.6 TerraGLM-5.2Claude Sonnet 5Kimi K2.7Nemotron 3 UltraMost cost & token efficient

Modeling score to cost in dollars, we can model the pareto frontier of the above models, based on available budgets. Grok on high ended up scoring best, with an additional $300 in spend needed by Sol on high effort to surpass Grok's performance on total points discovered (Fig. 6).

Fig. 6

Score vs. standard cost

Every model-effort cell. The dashed orange line marks the Pareto frontier, the highest score reached at or below each cost.

01326395265$0$110$220$330$440$550Standard cost ($)Score
Claude Opus 4.8
Claude Sonnet 5
GLM-5.2
GPT-5.5
GPT-5.6 Sol
GPT-5.6 Luna
GPT-5.6 Terra
Grok 4.5
Kimi K2.7
Kimi K3
Nemotron 3 Ultra

A similar pattern can be observed when modeling the pareto frontier of token cost to findings, with each Grok model appearing along the frontier, no matter the reasoning effort. Additionally, K3 scored on the pareto frontier for both cost and token efficiency (Fig. 7).

Fig. 7

Score vs. total tokens

The same model-effort cells on a log-scale x-axis, token spend ranges from roughly 1.4M to over 150M across runs.

013263952651M10M100MTotal tokens (log scale)Score
Claude Opus 4.8
Claude Sonnet 5
GLM-5.2
GPT-5.5
GPT-5.6 Sol
GPT-5.6 Luna
GPT-5.6 Terra
Grok 4.5
Kimi K2.7
Kimi K3
Nemotron 3 Ultra

Conclusion

As more software gets shipped, teams need to start thinking about how software can become self-securing. Long-running agents are the way forward, and to power these agents, we need to think about which models can execute tasks at an efficient cost while balancing the greatest performance. We're already seeing open source models quickly gain on closed source models like Opus, GPT, etc, and with powerful models getting cheaper, means that attackers get access to stronger weapons. Defenders must start using these same resources, to continuously catch vulnerabilities before attackers get a chance to exploit these vulnerabilities.

About the author

Akul Gupta, Co-Founder & CTO at MindFort

Akul Gupta

Co-Founder & CTO · MindFort

AI researcher focusing on LLMs in cybersecurity. Red-teamed models for OpenAI and Anthropic as part of their safety programs. Published multiple conference papers. M.S. Computer Science, UIUC.

Put your security on autopilot

First Deployment

15 min

Agents Running

24/7

More Coverage

100x

Hours Saved

1,000+