Penetration testing statistics in 2026: findings, breaches, and the fix backlog
The latest breach and pentest studies point to a practical problem for small SaaS teams: finding vulnerabilities is only useful when engineers can verify and close them. Here are the numbers, their denominators, and the decisions they support.
- Identify population
- Check denominator
- Apply finding
Start with the denominator. A breach report asks how recorded compromises happened. A pentest report asks what testers found inside selected engagements. A remediation study asks how those findings moved through a workflow. None tells you the probability that your particular website will be breached next month.
This reference uses the newest verified 2026 editions available on October 7, 2026. Their observation periods often end in 2025 or mid-2026. Numbers below are published findings, not experiments rerun by Pensec. Claim IDs connect the tables to the article’s evidence ledger and the repository’s detailed research register.
Breach statistics justify testing, but don’t measure pentest coverage
Vulnerability exploitation is a major recorded entry mechanism. That supports checking exposed software and application behavior; it doesn’t establish how many vulnerabilities a scanner will find, or whether an authenticated permissions test covers your deployment. For those scope distinctions, start with the website pentesting guide.
Verizon’s 2026 Data Breach Investigations Report executive summary, page 5, describes more than 31,000 incidents, more than 22,000 confirmed breaches, and organizations in 145 countries. These are contributor-supplied records, not randomly sampled websites. Its landing-page FAQ specifies incidents from November 1, 2024 through October 31, 2025. The summary instead says October 2024 through November 2025. That source discrepancy remains unresolved. The PDF’s June 5, 2026 metadata is a revision timestamp, not a verified launch date. [S01]
Table 1. Recorded breach attributes in Verizon’s 2026 executive summary. Page references are PDF pages. Different rows have different eligible populations; overlapping attributes must not be added.
| ID | Published statistic | Denominator and source location | Interpretation limit |
|---|---|---|---|
| S02 | Vulnerability exploitation: 31% | 19,905 non-Error, non-Misuse breaches with known initial-access vectors; p.9, Fig.4 | Includes organizational infrastructure, not just web application code. |
| S03 | Credential abuse: 13% | Same known-vector cohort, n=19,905; p.9, Fig.4 | Doesn’t include every later use of stolen credentials. |
| S04 | Ransomware: 48% | Reporting breach population; exact attribute-eligible n unavailable; p.10 | An attribute of recorded breaches, not ransomware probability per business. |
| S05 | Third-party involvement: 48% | Total breach population; exact eligible n unavailable; p.10 | Can overlap ransomware; isn’t necessarily a package supply-chain exploit. |
| S06 | Human element: 62% | Breach population; exact eligible n unavailable; p.11 | Doesn’t mean employee mistakes were the sole cause. |
| S07 | SMB exploitation: 26% | SMB breach slice: 7,152 confirmed disclosures from 7,256 incidents; exact known-vector subgroup n unavailable; p.16 | Broad SMB category, not a startup-SaaS census. |
| S08 | SMB credential abuse: 13% | Same SMB slice, metric-specific eligible n unavailable; p.16 | Not the percentage of all small businesses compromised. |
| S09 | SMB third-party involvement: 55% | Same SMB breach slice, eligible n unavailable; p.16 | No unbreached SMB population for estimating incidence. |
Source for all rows: Verizon, 2026 DBIR Executive Summary, pp.5, 9–11, 16, verified October 7, 2026. Overall corpus size is not substituted for an unavailable row-specific denominator.
The engineering implication is narrower than the headline. Maintain an inventory of reachable components, investigate known-exploited dependencies, and test the permissions behind valid sessions. DBIR does not show that tenant authorization is the largest breach cause, nor does it identify what fraction of exploitation came from AI-written code.
Mandiant’s M-Trends 2026, published March 24, 2026, supplies another view: its incident-response investigations conducted during 2025. The release describes over 500,000 investigation hours, but hours aren’t a case count. Organizations that call an incident-response provider are also a selected population. [S10]
Table 2. Mandiant’s 2025 caseload, reported in M-Trends 2026. The release does not disclose exact eligible case counts or unknown-vector exclusions.
| ID | Published statistic | Measured population | Source section / limitation |
|---|---|---|---|
| S11 | Exploits: 32% | Intrusions classified by initial infection vector | “By the Numbers”; not DBIR’s known-vector subset. |
| S12 | Voice phishing: 11% | Classified initial infection vectors | “By the Numbers”; not a survey of phishing recipients. |
| S13 | Prior compromise: 10% | Classified initial infection vectors | “Collapse of the Hand-Off Window”; measures reused access. |
| S14 | Email phishing: 6% | Intrusion initial infection vectors | “Voice Phishing and the SaaS Identity Crisis”; not email click rate. |
| S15 | Internal first detection: 52% | Investigations classified by detection source | “Detection by Source”; exact n unavailable. |
| S16 | Global median dwell time: 14 days | Incidents with measurable compromise-to-detection interval | “By the Numbers”; eligibility n unavailable; not patch time. |
Source: Mandiant, M-Trends 2026 release, March 24, 2026. Comparing its exploitation share with Verizon’s does not produce a pooled average: the providers, selection processes, and vector definitions differ.
Mandiant also describes stolen long-lived OAuth tokens, session cookies, hard-coded keys, and personal access tokens used to pivot into downstream SaaS customers. That’s qualitative evidence of mechanisms, without a frequency estimate. A stronger login does not invalidate an already stolen session. Test logout, suspension, membership removal, and integration revocation as server-side state changes. [S17]
Pentest statistics expose the gap between discovery and closure
The strongest practical finding is the unresolved tail. Cobalt’s State of Pentesting Report 2026, released April 21, 2026, analyzes over 16,500 pentests at nearly 3,000 organizations across five years, with Cyentia Institute analysis. Its separate double-blind survey covers 450 security professionals, evenly split between leaders and practitioners. Those survey respondents are not the denominator for vulnerability findings. [S18]
Cobalt defines high-risk finding half-life as the time to remediate half the findings, including still-unfixed findings. Its top organizational performance decile has a 10-day half-life; the bottom decile has 249 days. Exact group organization and finding counts aren’t disclosed on the public page. This is a survival-style measure, not the mean age of closed tickets. [S19]
Table 3. Cobalt’s 2026 finding-level results. Each percentage uses findings within its named category, not organizations or breaches.
| ID | Published statistic | Denominator | Limitation |
|---|---|---|---|
| S20 | AI/LLM high-risk share: 32% | Findings from AI/LLM pentests | Exact finding/engagement n unavailable; not the share of AI apps breached. |
| S21 | AI vulnerability resolution: 38% | AI-category findings | Eligible n, observation window, and retest/status rules unavailable. |
| S22 | Critical fixes under three days: 45% | Critical findings at programmatic teams | Group sizes unavailable; observational association. |
| S23 | Critical fixes under three days: 10% | Critical findings at compliance-driven teams | Different approach group; not randomized assignment. |
Source: Cobalt, State of Pentesting 2026, “AI risks” and “Programmatic advantage”, April 21, 2026. The publisher’s rounded comparative multipliers aren’t needed to make the operational point.
A small team doesn’t need an enterprise program to act on this. It needs an owner, an understandable reproduction, a deployment-specific impact judgment, and a reserved retest step. Scheduling more discovery while leaving every fix unowned can make the report queue larger without changing customer exposure.
Faster reported resolution can coexist with a growing backlog
HackerOne’s Security Research Report 2026, tenth edition, makes the distinction unusually visible. Its current public page describes 92,194 valid vulnerability reports and $89.1 million in bounty payouts for July 2025–June 2026. The exact launch day isn’t established; the 2026 edition was available when verified on October 7. [S24]
Table 4. HackerOne’s 2026 remediation and survey results. The current corpus is not automatically the denominator for historical comparisons.
| ID | Published result | Denominator / comparison | Limitation |
|---|---|---|---|
| S25 | Critical mean resolution time fell 61% | Platform critical-report resolution populations over two years | Exact eligible n, resolved-only eligibility, and censoring rules unverified. |
| S26 | Open valid backlog rose 131% | Open valid-report counts, 2024 versus 2026 | Absolute snapshot counts unavailable; discovery/program growth can contribute. |
| S27 | Open critical backlog rose 123% | Open critical-report counts over two years | Not a rate of newly vulnerable applications. |
| S28 | All-severity mean resolution: 135 → 62 days | Platform reports used for resolution averages | Exact eligible n, resolved-only eligibility, and censoring rules unverified; not Cobalt’s half-life. |
| S29 | Discovery outpaces fixes: 70% | Survey of 111 security leaders; question-level n unavailable | Self-report from $100M+ revenue organizations, not small SaaS. |
| S30 | Critical validation takes 30 hours, average | Same leader survey; per-question n unavailable | Validation is a separate stage before repair. |
| S31 | Valid AI report counts rose 222% | Counts in AI classifications over two years | Absolute category counts and tested-asset denominator unavailable. |
Source: HackerOne, Security Research Report 2026, highlights and methodology FAQ, available October 7, 2026. Surveys were fielded in summer 2026: 111 enterprise leaders across six countries and 408 active researchers. Independent yearly survey samples aren’t directly comparable.
A resolution-time average and an open-backlog count answer different questions. If the average covers resolved reports, it describes how quickly those represented reports closed; the publisher’s public summary does not establish the exact eligibility or censoring rules. The open backlog counts valid findings that remain unresolved.
Under a resolved-only calculation, a team could improve the average by fixing easy, newly discovered issues while older, difficult findings remain. That’s a hypothetical selection effect, not a verified explanation of HackerOne’s results. The published summaries don’t establish whether it occurred or quantify its contribution; the reported time and backlog trends nevertheless show why both metrics deserve attention.
Track open high-impact findings by age, new validated findings, verified closures, and reopened issues. Distinguish a duplicate report from a unique root cause. Count uncertain findings separately. A hundred scanner alerts aren’t a hundred exploitable defects, and a hundred closed alerts aren’t evidence that a tenant boundary now holds.
AI code statistics are a different experiment again
Veracode’s July 28, 2026 summary of its 2026 GenAI Code Security Report reports a 56% average security pass rate and roughly 44% risky outputs on selected generation tasks, not pentesting results. Its report page explicitly attaches the 56% average to more than 100 models over four years, separately describing 11 newly tested models across 80 tasks. Exact eligible-output counts and aggregation weights remain unavailable. The 56% isn’t established as an exclusive Summer 2026 cohort rate, and 11 × 80 isn’t its verified denominator. [S32]
Table 5. Security-task pass rates published by Veracode in 2026. Each weakness row measures the relevant task/model evaluations; exact per-category n is unavailable.
| ID | Metric | Published pass rate | What it doesn’t establish |
|---|---|---|---|
| S33 | GPT-5.5, Summer 2026 snapshot | 68% | Not a ranking of newer October models or real SaaS features. |
| S34 | SQL injection tasks | 83% | Not the safety rate of all generated database code. |
| S35 | Cryptographic-algorithm tasks | 87% | Not a guarantee of correct key lifecycle or protocol design. |
| S36 | Cross-site scripting tasks | 15% | Not the fraction of every AI-written web component that’s safe. |
| S37 | Log injection tasks | 12% | Selected test conditions, not incident prevalence. |
Sources: Veracode’s dated report article, “Where vulnerabilities persist” and Summer 2026 methodology outline. Published July 28, 2026; verified October 7.
These results motivate contextual review: where does untrusted text become HTML, a query, a log record, or a privileged tool argument? They don’t support multiplying a generation-task failure percentage by a breach percentage. Mandiant explicitly says it doesn’t consider 2025 the year breaches were directly caused by AI. Attacker use of AI and AI as the direct cause of a breach are different claims. [S38]
For the generating agent’s permissions and untrusted-input boundary, see coding-agent security. For discovery and exploitation performance, use the separate AI pentesting leaderboard. Writing secure code, identifying a bug, exploiting it, and repairing it are different tasks.
Turn the evidence into a small-SaaS remediation contract
The useful response is a testable workflow, not a universal budget percentage. Consider a hypothetical workspace app adding invitations, file exports, and customer API tokens. The risky question is whether the backend permits one tenant’s member to act on another tenant’s object. An anonymous homepage scan cannot answer that question because it never exercises the authorized session and wrong-object combination.
OWASP API1:2023 requires action-specific object authorization. WSTG 4.2, WSTG-ATHZ-04 recommends at least two users, often more, with different objects and privileges. These are standards and procedures, not prevalence measurements. [S39]
Use two synthetic tenants and the roles your application actually supports. Preserve an allowed read as a positive control. Attempt the prohibited read, update, export, download, and bulk operation using the other tenant’s known object identifiers. Include intended sharing exceptions. The multi-tenant authorization testing guide expands that matrix.
For each confirmed finding, retain the actor, target object, request sequence, observed response or state change, and deployment revision. Explain why the action should have been denied. Give the engineer a root-cause location and acceptance condition. If evidence isn’t reproducible, label the uncertainty rather than inflating severity to secure attention.
After the fix, replay the forbidden action and verify the legitimate action still works. Test adjacent entry points sharing the same data-access code. For a leaked credential, confirm revocation rather than merely deleting the string from a repository. For a session flaw, replay the old token after the policy’s revocation event. Closure is a behavioral claim.
This is an engineering recommendation derived from the evidence, not a published effectiveness percentage. Pensec is an early-access project; this article doesn’t claim a shipped scanner or completed customer assessments. The same contract is usable with a specialist pentester, your own regression suite, or an existing testing service.
Evidence ledger and refresh rules
The source families below keep the statistics auditable without pretending public summaries disclose every denominator. All were retrieved 2026-10-07. Exact claims, protocols, missing fields, and fetch notes are maintained in docs/blog/claims-research.md.
| IDs | Primary publication / location | Evidence type and date |
|---|---|---|
| S01–S09 | Verizon executive summary, pp.5, 9–11, 16; FAQ | 2026 recorded breach corpus; exact launch date unverified; PDF revision June 5. |
| S10–S17, S38 | Mandiant M-Trends 2026, numbers, SaaS identity, AI sections | March 24, 2026; 2025 incident-response observations. |
| S18–S23 | Cobalt State of Pentesting 2026, Key Findings and FAQ | April 21, 2026; five-year findings corpus and separate survey. |
| S24–S31 | HackerOne Security Research Report 2026, highlights and FAQ | July 2025–June 2026 platform data; summer surveys; launch day unknown. |
| S32–S37 | Veracode article and report | July 28, 2026; selected code-generation tasks. |
| S39 | OWASP API1:2023 and WSTG 4.2, linked above | Normative guidance; no statistical population. |
Refresh when an edition or methodology changes. Keep the original population beside every number, preserve unresolved source conflicts, and don’t derive a pentesting market size from vulnerability counts. For your own app, the next useful number is simpler: how many verified high-impact findings remain open, and who owns the next fix?