AI pentesting leaderboard, October 2026: rank the benchmark, not the brand

GPT-6 Astra leads OpenAI’s September exploitation comparison; GPT-5.4 with Codex leads the verified cost-capped CyberGym-E2E cohort. This reference separates those rankings from restricted new models, specialized agents, and web-app testing claims.

Read as MarkdownSubscribe via RSS
  1. Task and evidence
  2. Model plus harness
  3. Verified outcome
A score belongs to a task version, information policy, deployed model and harness, resource budget, and success criterion. Changing any of them changes the comparison.

An exploit that executes code, a sanitizer crash, a correct patch, and a captured CTF flag aren’t interchangeable successes. Before choosing a model, choose the task you need it to perform. Before comparing two scores, check whether both systems received the same information and resources.

The cutoff here is October 7, 2026. Results are reported as published, not locally rerun. Primary system cards, author papers, official board JSON/CSV, and original vendor accounts were inspected. Live-board dates identify submissions; they don’t prove immutable historical snapshots. “NR” means not reported or not recovered, never zero. Claim IDs map to docs/blog/claims-research.md.

The newest models are not all rankable against each other

There is no verified equal-budget cyber head-to-head covering all the newest frontier releases. The latest quantitative card inspected is OpenAI’s September 29 GPT-6.1 Sol addendum. Google’s September 30 Gemini 4 Argon announcement is newer, but mainly supplies defensive evidence and restricted-rollout information. Anthropic’s September releases describe cyber capability and safeguards without a comparable numeric cohort recovered here. Its October 6 access-condition experiment is discussed separately below.

Table 1. Newest-model evidence and availability ledger. This is the newest evidence recovered for these families, not proof of an exhaustive global release search. Launch access is preserved as stated; future expansion isn’t assumed completed.

Model / dated source Access or deployed behavior Rankable cyber evidence here
GPT-6.1 Sol · Sept 29 addendum [A01] Research/API evaluation may differ from ChatGPT; Astra safeguards and phased Daybreak cyber access Same-card ExploitGym and refreshed exploitation comparison below.
GPT-6 Astra · Sept 3 card [A02] Research configurations and sensitive targets aren’t fully public Highest in September OpenAI exploitation cohort; not equal-budget across providers.
Gemini 4 Argon · Sept 30 announcement [A03] Rolling out to trusted Fairwind defenders; broader release still future tense CWE-bench v1 68%, provider-reported tie for first; task n/budget NR in announcement. Unranked on exploitation cohorts.
Claude Sonnet 5.5 · Sept 28 launch [A04] Generally available; higher-risk tasks fall back to Sonnet 5; Oct 6 CVP expansion adds tiered cyber access Provider says cyber comparable to Opus 5. No comparable new exploitation score recovered: unranked.
Claude Opus 5.5 · Sept 22 launch [A05] Generally available; cyber fallback to Opus 4.8 at launch; Oct 6 CVP expands access Oct 6 CyScenarioBench result below; no same-protocol exploitation comparison: unranked across providers.
Claude Fable 5.1 / Mythos 5.1 · Sept 1 launch record; announcement [A06] Same underlying model; Fable general, Mythos trusted access; Fable redirects pentesting/exploit generation Strongest released Claude cyber model at that date is a provider statement. Not Mythos Preview’s old score.
Grok 4.7 early candidate · XBOW Sept 21 evaluation [A07] Early-access candidate; exact identifier NR; this isn’t xAI release verification Private harness evidence, including 50 exploit tasks ×10 repetitions; unranked across providers.
XekRung-1.5-27B-Preview · Alibaba Sept 13 report [A08] Self-hosted FP8, Qwen3.8-27B fine-tune; up to eight hours/3,000 turns CyberGym model-focused board 88.92%, report rounds to 88.9%; unequal budgets.
GLM-5.3 / DeepSeek-V4-Pro / Feyospace-v1.1 · official board [A09] Board dates Aug 14 / Aug 13 / Sept 23; complete release access not audited Board-listed reproduction scores, not normalized latest-provider comparison.

An unavailable score means the inspected sources don’t supply an auditable result for that exact model and protocol. It doesn’t mean the model failed or lacks capability. We also didn’t extract the full latest Anthropic cyber PDFs; the source notebook records oversized responses and CDN failures. Launch prose supports the access statements, not invented numeric results.

The latest access update is Anthropic’s October 6, 2026 expanded Cyber Verification Program. It includes Opus 5.5, Sonnet 5.5, and Mythos 5.1 in three tiers: Defense Access, Red Team Access, and Specialized Access. Authorized pentesting is explicitly in Red Team Access; that tier is organizations-only at publication. This supersedes the launches’ future-tense expansion language, but doesn’t establish that every applicant has access. [A31]

Its Opus 5.5 experiment uses 10 CyScenarioBench challenges × five attempts per tier. Generally available access blocked every task on the first prompt. Defense Access blocked 46/50 trials, with the remaining 4/50 succeeding. Red Team Access produced no blocks and 34/50 successes (68%, calculated). The provider compares that with 67.6% without safeguards, but doesn’t give the exact aggregation denominator for that separate rate. This is a useful same-model access-condition comparison, not a cross-provider leaderboard. Costs and time ceilings are NR; we didn’t reproduce the trials. [A31]

General Terminal-Bench 4.0 scores cannot fill this gap. They measure terminal-task completion, not a security-only population. Anthropic’s Opus 5.5 methodology explicitly says intervened cyber tasks were completed by Opus 4.8. A score can therefore belong to a routed deployment, even when the row carries a single model name. Availability, fallback, and effort level belong beside capability. [A05]

The September OpenAI same-card exploitation cohort

GPT-6 Astra leads the September 29 card’s intended ExploitGym success comparison; GPT-6.1 Sol follows. This is a defensible within-provider, within-card ordering of reported point estimates, not a statistical significance claim or an equal-cost competition with other providers. The card warns that older comparison values can reflect later snapshots than their launch cards. [A10]

Table 2. OpenAI September 29, 2026, §9.1.2: rank by intended ExploitGym v1 success per attempt. All rates below are from the same card. The public benchmark has 869 tasks. Exact seeds, integer numerators, token cap, and observed costs weren’t recovered. The inherited protocol is token-capped without a wall-clock cap.

Cohort rank Model Intended exploit / attempt SEC-Bench Pro pass@1
1 GPT-6 Astra 42.4% 85.4%
2 GPT-6.1 Sol 35.1% 78.8%
3 GPT-5.6 Sol 30.3% 79.1%
4 GPT-6 Sol 22.1% 66.3%

Source: OpenAI GPT-6.1 Sol addendum, §9.1.2.2 and Figures 36–37. SEC-Bench Pro uses a separate 183-vulnerability V8/SpiderMonkey May-2026 port with corrected root-cause grading, described in the Astra card. Peak effort and token consumption differ. Notice that the middle models reverse order on that metric. [A10–A11]

The same card’s private June–August 2026 ExploitBench refresh reports arbitrary-code-execution rates of 31.5%, 21.5%, 3.5%, and 5.5%, respectively, in the table’s model order. Its task count is NR; these are not rates over the public ExploitBench denominator. The private target set limits reproduction. [A12]

Public ExploitBench is different again: 41 V8 vulnerabilities, 16 capability flags, five seeds, and a capability percentage based on the union of flags across seeds, with arbitrary code execution earning full credit. Astra’s 100%, GPT-6.1 Sol’s 99.7%, and GPT-6 Sol’s 81.7% are not single-attempt exploit success rates. OpenAI explicitly warns of potential contamination from historical vulnerabilities. [A13]

That distinction prevents an attractive but false summary: “the newest model exploits almost every target.” Near-complete capability credit on an old, multi-seed evaluation is consistent with much lower actual code-execution success on a recent private refresh. Neither result should be averaged with the other.

CyberGym-E2E under both a cost and time cap

GPT-5.4 with Codex leads the verified 920-task, $10 and 90-minute cohort. This comparison controls the dataset, scoring, API-spend ceiling, and wall-clock ceiling. It remains a model-plus-harness comparison: Codex, Gemini CLI, and Claude Code manage context and tools differently. Tasks stop when either limit is reached. [A14]

The June 3, 2026 CyberGym-E2E paper, §3.4, defines cumulative stages:

  • S1: the agent’s input crashes the vulnerable program.
  • S2: its patch prevents that input’s crash.
  • S3: the patched project passes functionality tests.
  • S4: the patch also prevents the ground-truth vulnerability’s crash.

S3 is the authors’ main end-to-end outcome. S4 distinguishes the intended historical target from a different valid bug. A crash isn’t automatically weaponized code execution; passing existing tests isn’t proof of complete patch correctness. The dataset contains 920 historical vulnerabilities across 139 OSS projects, predominantly C/C++ memory-safety tasks. [A14]

Table 3. Strict budget-controlled ranking, paper Table 4, June 3, 2026. Every percentage is over the same 920-task cohort; exact per-row numerators and pinned CLI/API hashes are NR.

S3 rank / deployed system S1 S2 S3 S4
1 · GPT-5.4 / Codex 67.9% 66.2% 65.9% 22.2%
2 · Gemini 3.1 Pro / Gemini CLI 47.4% 44.3% 43.8% 20.5%
3 · Claude Opus 4.6 / Claude Code 39.7% 39.5% 37.9% 15.7%

Sources: CyberGym-E2E paper, Table 4 and official backing data, verified October 7. The separate patch-only scores are 87.1%, 83.0%, and 84.1%, respectively, when the original PoC/crash log is supplied. Those aren’t end-to-end discovery rates. [A15]

CyberGym-E2E: end-to-end success versus intended-target success On 920 tasks with a ten dollar API-spend and ninety minute time ceiling per task, GPT-5.4 with Codex scores 65.9 percent S3 and 22.2 percent S4. Gemini 3.1 Pro with Gemini CLI scores 43.8 and 20.5 percent. Claude Opus 4.6 with Claude Code scores 37.9 and 15.7 percent. Bars use a linear zero to one hundred percent scale. End-to-end isn't the intended bug 920 tasks · $10 API spend AND 90 minutes per task S3: crash → patch → tests pass S4: also blocks ground-truth crash GPT-5.4 / Codex 65.9% 22.2% Gemini 3.1 Pro / Gemini CLI 43.8% 20.5% Claude Opus 4.6 / Claude Code 37.9% 15.7% 0%50%100%
Figure 2. Published CyberGym-E2E Table 4, June 3, 2026: 920 historical OSS vulnerabilities, 139 projects, common $10/90-minute ceilings, no ground-truth bug information in E2E. S3 permits another valid discovered bug; S4 additionally tests the historical target. Source: author paper. Model+harness results, not a local rerun or production web-pentest recall. [A14–A15]

The newest E2E result inspected is AWS Continuum, published October 5, 2026: 819/920, 89.0% S3, and 37.8% S4, under 90 minutes without a cost cap. The board labels the underlying model Claude Opus 5; AWS describes a proprietary multi-agent, multi-model architecture. Total spend, concurrency, internal retries, and complete snapshots are NR. Its longer-running 93.7% result isn’t the official 90-minute score. [A16]

Continuum is therefore an uncapped system result, not first place in Table 3. The importance of the cost distinction is measurable: the same June Opus 4.6/Claude Code combination reaches 62.6% S3 without the cost ceiling, versus 37.9% with it, on 920 tasks and the same 90-minute ceiling. That’s spend sensitivity, not a model upgrade. [A17]

Harness matters too. In the separate 615-task cohort, Sonnet 4.5 reaches 10.6% S3 with Claude Code and 5.4% with OpenHands, both under $10/90 minutes. The authors discuss targeted reads and task tracking as explanations; this doesn’t isolate each mechanism causally. It does rule out treating a model name as the complete experimental specification. [A18]

ExploitGym v1: direct two-hour runs and a six-hour run’s two-hour checkpoint

GPT-5.6 Sol has the highest inspected intended-exploit count observed by two hours, but its result is a checkpoint from a six-hour run. It is not a separately verified two-hour-budget run. Unlike E2E’s sanitizer crashes, ExploitGym requires code execution or restricted flag recovery through the intended vulnerability, checked by an agent judge. Agents receive source, builds, a compiled target, the vulnerability description, and an existing crashing proof of vulnerability. [A19–A20]

Table 4. Descriptive ExploitGym v1 observations, June 16 and July 13, 2026 entries. N=869, comprising 502 userspace, 181 V8, and 186 kernel tasks. Rates are calculated as published intended-success count /869. The rows include direct two-hour runs and a two-hour checkpoint from a six-hour run; they are descriptive observations, not an established common-allocation comparison. Spend also differs. Mitigation-on counts are separate runs, not fractions of successful mitigation-off exploits.

Model + harness / run type Intended /869 Mitigation-on total Mean $ / attempted task
GPT-5.6 Sol max / Codex CLI · July 13, two-hour checkpoint of six-hour run 216 (24.86%) 84 $61.69
GPT-5.5 / Codex CLI · June 16, direct two-hour run 129 (14.84%) 22 $35.43
GPT-5.4 / Codex CLI · June 16, direct two-hour run 61 (7.02%) 3 $25.25
Opus 4.6 / Claude Code · June 16, direct two-hour run 16 (1.84%) 0 $21.76
Opus 4.7 / Claude Code · June 16, direct two-hour run 12 (1.38%) 0 $4.82
GLM-5.1 / Claude Code · June 16, direct two-hour run 4 (0.46%) 0 $6.39

Source: official ExploitGym JSON, verified October 7, 2026. GPT-5.6 Sol is the two-hour cutoff subrow of a six-hour run. A longer announced horizon could affect planning; a common stopping and planning policy is not established. Costs are observed means, not caps or costs per successful exploit. GPT-5.5’s row explicitly warns that default safety filters block attempts under default prompting; the evaluated configuration must not be assumed available unchanged. [A20]

The full six-hour GPT-5.6 Sol row reports 293/869 (33.72%), averaging $147.01 per attempted task. That endpoint and its two-hour checkpoint describe different elapsed times within the same longer-horizon run, not two separately verified allocation conditions. Old v0 Gemini/Mythos Preview entries, provider rows lacking full protocol fields, selected subsets, and GLM-5.3’s explicitly rescaled timeouts remain outside Table 4. The original May 11 paper abstract says 898 tasks; the current v1 data says 869. Version drift isn’t permission to choose whichever denominator produces the better headline. [A21]

September’s OpenAI token-capped rates in Table 2 also cannot be merged with these time-capped counts. The same model can have different results under a changed grader, package restrictions, effort, and stopping policy. A fair purchasing comparison needs observed cost and time alongside the score, not merely an API price list.

CyberGym’s high scores are reproduction claims, not general superiority

CyberGym Level 1 supplies a vulnerability description and unpatched codebase, then checks whether the final PoC reproduces the target on the vulnerable revision but not the patched revision. It contains 1,507 tasks from 188 projects. It is valuable large-scale evidence of reproduction, not blind discovery or complete application pentesting. [A22]

Table 5. Selected model-focused board entries, unranked across budgets and harnesses. Every result below is a board-listed one-trial claim over Level 1’s 1,507 instances; exact success numerators are NR. “Model-focused” still includes an agent harness. Source: official JSON, retrieved October 7.

Model / harness Reproduction Submission date / key caveat
XekRung-1.5-27B-Preview / XekRung Agent 88.92% Sept 13; eight-hour/3,000-turn ceiling, static target access.
Gemini 3.8 Flash Cyber / Google Antigravity 86.26% Sept 10; trusted cyber variant, budget NR.
GPT-5.5-Cyber / OpenAI Agent 85.6% June 22; harness snapshot and budget NR.
GLM-5.3 / Claude Code 84.5% Aug 14; budget NR; full provider report not audited.
Feyospace-v1.1 / Claude Code 84.41% Sept 23; budget NR; full report not audited.
DeepSeek-V4-Pro / DeepSeek Agent 83.3% Aug 13; provider submission, budget NR.
Claude Mythos Preview / Anthropic Agent 83.1% April 7; not Mythos 5.1.
GPT-5.5 / OpenAI Agent 81.8% April 23; not GPT-6.1 Sol.
Grok 4.6 / Grok Build 79.7% Aug 12; not Grok 4.7.

These catalog entries identify demonstrated configurations without awarding an equal-budget rank. ASL-Cyber-Flash’s 85.87%, dated September 3, is separately labeled dynamic vulnerable-image access. It shouldn’t be silently grouped with static-access entries. [A23]

The board’s above-90% tier is intentionally shown in random order with scores for reference. Preserve that policy. Examples include PwnBot, 93.1%, September 27 (dynamic execution plus cross-task test-time memory), Wiz Atlas, 90.9%, July 27 (multi-model/multi-stage), and Lyrie Agent, 99.20%, September 15 board date (multi-model plus dynamic execution). None is ranked here as generally superior. [A24]

Lyrie’s original report, metadata dated September 24, supplies 1,495/1,507 confirmed solves: 894 in a seed-guided libFuzzer lane and 601 in an agentic generator lane. Its three validation repetitions aren’t three independent solving attempts. Its quoted API economics exclude full self-hosted 16×B300 and CPU economics. The result belongs to the whole system, not to a bare model producing a final answer. [A25]

Web relevance and human comparisons need separate evidence

CVE-Bench is closer to web-app compromise than C/C++ fuzzing. Its original study has 40 critical web CVEs and automatic graders for outcomes such as database access, privilege escalation, and outbound requests. Success can be a permitted compromise objective rather than exploitation of the designated CVE. “Zero-day” withholds the description; it doesn’t mean the study discovered a new vulnerability. [A26]

The repository’s January 12, 2026 v2.1 release replaced arbitrary file upload with remote code execution; its current CLI example says 2.2.0. Historical 40-task scores cannot be advertised as current-version results. No normalized newest-frontier leaderboard was recovered. That leaves CVE-Bench unranked for current models, despite its stronger web relevance.

Cybench also remains a descriptive board rather than a unified latest-model league. Its official CSV reports Mythos Preview 100% over 35 tasks, Opus 4.7 96% over 35, and Opus 4.6 93% over 37, while the original suite has 40 CTF tasks. Subsets, provider harnesses, and rollout aggregation differ. These aren’t October’s Mythos 5.1 or Opus 5.5 scores. A competition’s first-human-solve time isn’t equal-budget professional labor. [A27]

XBOW’s August 5, 2024 human experiment is worth reading precisely because it describes the comparison. Five paid pentesters and XBOW attempted 104 newly commissioned web challenges. The best human and XBOW each scored a rounded 85%; the human union solved 87.5%, or 91/104. Humans had 40 hours; XBOW completed in 28 minutes wall-clock. Model/build, concurrency, hardware, and token budget are NR. The principal performed better on the hardest tasks. [A28]

That’s measured controlled-challenge evidence, not matched-resource proof that an agent replaces a production pentest. A real engagement includes scope, authentication setup, intended business rules, impact interpretation, partial failures, and reporting. A flag challenge supplies a crisp objective that production applications often don’t have.

Newer web-relevant evidence also shows harness fit. XBOW’s September 21, 2026 Grok candidate study reports Build-based correct findings rising from an equal-weight mean of 42 with Grok 4.6 to 68 with Grok 4.7, across three large targets, three workload sizes, and four systems. Its existing exploit harness slightly regressed. There’s no stable denominator of all true bugs, so those counts aren’t recall percentages or an open cross-provider rank. [A07]

Finding more isn’t enough if the evidence is wrong

The September 9, 2026 AWS Deception Benchmark tests single-turn vulnerability classification without agent scaffolding. It includes 14,822 samples, of which 9,695 are scored and 5,127 are unscored decoys, across 16 languages and over 70 CWEs. The scored set combines 6,988 code-level and 2,707 environment-gated examples. [A29]

Table 6. Leading three published Proof-of-Exploit-prompt accuracy point estimates, not a pentesting rank. Accuracy denominator: scored samples. False-positive rate denominator: safe scored samples; false-negative rate: vulnerable scored samples. Exact class counts and model snapshots are NR in the announcement.

PoE accuracy order Accuracy False positives False negatives
Claude Opus 5 79.3% 24.9% 16.8%
GPT-5.4 77.7% 10.1% 33.6%
Claude Opus 4.7 75.9% 32.0% 16.8%

Source: AWS original results table, September 9, 2026. PoE is a prompting strategy, not an executed exploit oracle. These adversarial classification errors are not commercial tool false-positive rates. No tested configuration had both error rates below 10% on this benchmark. [A30]

For a web SaaS buyer, ask for role-and-tenant coverage, reproducible request chains, independently checked impact, duplicates, uncertain findings, and human triage time. Evaluate the exact production deployment, including fallback models. Then preserve the original exploit and verify the fix with positive controls. Our multi-tenant authorization guide and website pentesting guide explain those application-specific checks.

If an agent can read repositories and run commands, its own permissions and untrusted-input boundaries also need review; see coding-agent security. If findings arrive faster than engineers close them, consult the separate 2026 penetration testing statistics. Discovery performance and verified remediation capacity should be measured together.

These are proposed evaluation criteria for Pensec, which is early access, not demonstrated scanner capabilities. We haven’t rerun these benchmarks, audited every submission, or established production-pentest recall. The ExploitGym two-hour checkpoint is not a verified equal-allocation run. The useful conclusion is concrete: Astra leads the verified September OpenAI exploitation cohort; GPT-5.4/Codex leads the strict E2E budget cohort; Table 4 supplies descriptive observations rather than a common-allocation ranking; newer unscored models remain unranked until the right evidence exists.

Source ledger and maintenance

All primary links above were retrieved 2026-10-07. The detailed claim register records dates, source locations, budgets, denominators, and fetch limitations. The June E2E paper and official JSON support Tables 3 and Figure 2; official ExploitGym JSON supports Table 4; CyberGym JSON supports the submission catalog; OpenAI’s September 29 §9.1.2 supports Table 2; AWS’s September 9 table supports Table 6.

Keep this reference’s comparisons separate when refreshing it. A new model release updates availability, not an old model’s score. A board submission updates a reported configuration, not every deployment of its underlying model. A new dataset or grader creates a new cohort. No weighted average of these benchmarks would repair the missing comparability.

About Gabe

@bucabay

Gabe writes about website pentesting, security research, and coding-agent security at Pensec.

More from Gabe