Penetration testing statistics in 2026: findings, breaches, and the fix backlog

The latest breach and pentest studies point to a practical problem for small SaaS teams: finding vulnerabilities is only useful when engineers can verify and close them. Here are the numbers, their denominators, and the decisions they support.

Read as MarkdownSubscribe via RSS
  1. Identify population
  2. Check denominator
  3. Apply finding
An evidence-reading workflow: identify who or what was observed, check the metric’s denominator and limitations, then decide what the finding supports for your application. These audit steps are not an empirical funnel linking breaches, pentest findings, and fixes.

Start with the denominator. A breach report asks how recorded compromises happened. A pentest report asks what testers found inside selected engagements. A remediation study asks how those findings moved through a workflow. None tells you the probability that your particular website will be breached next month.

This reference uses the newest verified 2026 editions available on October 7, 2026. Their observation periods often end in 2025 or mid-2026. Numbers below are published findings, not experiments rerun by Pensec. Claim IDs connect the tables to the article’s evidence ledger and the repository’s detailed research register.

Breach statistics justify testing, but don’t measure pentest coverage

Vulnerability exploitation is a major recorded entry mechanism. That supports checking exposed software and application behavior; it doesn’t establish how many vulnerabilities a scanner will find, or whether an authenticated permissions test covers your deployment. For those scope distinctions, start with the website pentesting guide.

Verizon’s 2026 Data Breach Investigations Report executive summary, page 5, describes more than 31,000 incidents, more than 22,000 confirmed breaches, and organizations in 145 countries. These are contributor-supplied records, not randomly sampled websites. Its landing-page FAQ specifies incidents from November 1, 2024 through October 31, 2025. The summary instead says October 2024 through November 2025. That source discrepancy remains unresolved. The PDF’s June 5, 2026 metadata is a revision timestamp, not a verified launch date. [S01]

Table 1. Recorded breach attributes in Verizon’s 2026 executive summary. Page references are PDF pages. Different rows have different eligible populations; overlapping attributes must not be added.

ID Published statistic Denominator and source location Interpretation limit
S02 Vulnerability exploitation: 31% 19,905 non-Error, non-Misuse breaches with known initial-access vectors; p.9, Fig.4 Includes organizational infrastructure, not just web application code.
S03 Credential abuse: 13% Same known-vector cohort, n=19,905; p.9, Fig.4 Doesn’t include every later use of stolen credentials.
S04 Ransomware: 48% Reporting breach population; exact attribute-eligible n unavailable; p.10 An attribute of recorded breaches, not ransomware probability per business.
S05 Third-party involvement: 48% Total breach population; exact eligible n unavailable; p.10 Can overlap ransomware; isn’t necessarily a package supply-chain exploit.
S06 Human element: 62% Breach population; exact eligible n unavailable; p.11 Doesn’t mean employee mistakes were the sole cause.
S07 SMB exploitation: 26% SMB breach slice: 7,152 confirmed disclosures from 7,256 incidents; exact known-vector subgroup n unavailable; p.16 Broad SMB category, not a startup-SaaS census.
S08 SMB credential abuse: 13% Same SMB slice, metric-specific eligible n unavailable; p.16 Not the percentage of all small businesses compromised.
S09 SMB third-party involvement: 55% Same SMB breach slice, eligible n unavailable; p.16 No unbreached SMB population for estimating incidence.

Source for all rows: Verizon, 2026 DBIR Executive Summary, pp.5, 9–11, 16, verified October 7, 2026. Overall corpus size is not substituted for an unavailable row-specific denominator.

The engineering implication is narrower than the headline. Maintain an inventory of reachable components, investigate known-exploited dependencies, and test the permissions behind valid sessions. DBIR does not show that tenant authorization is the largest breach cause, nor does it identify what fraction of exploitation came from AI-written code.

Mandiant’s M-Trends 2026, published March 24, 2026, supplies another view: its incident-response investigations conducted during 2025. The release describes over 500,000 investigation hours, but hours aren’t a case count. Organizations that call an incident-response provider are also a selected population. [S10]

Table 2. Mandiant’s 2025 caseload, reported in M-Trends 2026. The release does not disclose exact eligible case counts or unknown-vector exclusions.

ID Published statistic Measured population Source section / limitation
S11 Exploits: 32% Intrusions classified by initial infection vector “By the Numbers”; not DBIR’s known-vector subset.
S12 Voice phishing: 11% Classified initial infection vectors “By the Numbers”; not a survey of phishing recipients.
S13 Prior compromise: 10% Classified initial infection vectors “Collapse of the Hand-Off Window”; measures reused access.
S14 Email phishing: 6% Intrusion initial infection vectors “Voice Phishing and the SaaS Identity Crisis”; not email click rate.
S15 Internal first detection: 52% Investigations classified by detection source “Detection by Source”; exact n unavailable.
S16 Global median dwell time: 14 days Incidents with measurable compromise-to-detection interval “By the Numbers”; eligibility n unavailable; not patch time.

Source: Mandiant, M-Trends 2026 release, March 24, 2026. Comparing its exploitation share with Verizon’s does not produce a pooled average: the providers, selection processes, and vector definitions differ.

Mandiant also describes stolen long-lived OAuth tokens, session cookies, hard-coded keys, and personal access tokens used to pivot into downstream SaaS customers. That’s qualitative evidence of mechanisms, without a frequency estimate. A stronger login does not invalidate an already stolen session. Test logout, suspension, membership removal, and integration revocation as server-side state changes. [S17]

Pentest statistics expose the gap between discovery and closure

The strongest practical finding is the unresolved tail. Cobalt’s State of Pentesting Report 2026, released April 21, 2026, analyzes over 16,500 pentests at nearly 3,000 organizations across five years, with Cyentia Institute analysis. Its separate double-blind survey covers 450 security professionals, evenly split between leaders and practitioners. Those survey respondents are not the denominator for vulnerability findings. [S18]

Cobalt defines high-risk finding half-life as the time to remediate half the findings, including still-unfixed findings. Its top organizational performance decile has a 10-day half-life; the bottom decile has 249 days. Exact group organization and finding counts aren’t disclosed on the public page. This is a survival-style measure, not the mean age of closed tickets. [S19]

High-risk finding half-life: 10 days versus 249 days Cobalt's 2026 report compares the top and bottom organizational performance deciles. On a linear zero to 250 day scale, leaders take 10 days to remediate half their high-risk findings and laggards take 249 days. Still-open findings are included. Subgroup counts are not published. The unresolved tail changes the result High-risk finding half-life · days · lower is shorter Top performance decile 10 days Bottom performance decile 249 0100200250 Includes findings that haven't been fixed.
Figure 2. Cobalt, State of Pentesting 2026, April 21: five-year corpus of over 16,500 pentests at nearly 3,000 organizations; exact decile finding counts unavailable. Half-life includes open findings. Groups are defined by remediation performance, so this chart doesn't prove a process caused the difference. Source: Key Findings and Leaders/Laggards methodology FAQ. Published results, not a local rerun. [S18–S19]

Table 3. Cobalt’s 2026 finding-level results. Each percentage uses findings within its named category, not organizations or breaches.

ID Published statistic Denominator Limitation
S20 AI/LLM high-risk share: 32% Findings from AI/LLM pentests Exact finding/engagement n unavailable; not the share of AI apps breached.
S21 AI vulnerability resolution: 38% AI-category findings Eligible n, observation window, and retest/status rules unavailable.
S22 Critical fixes under three days: 45% Critical findings at programmatic teams Group sizes unavailable; observational association.
S23 Critical fixes under three days: 10% Critical findings at compliance-driven teams Different approach group; not randomized assignment.

Source: Cobalt, State of Pentesting 2026, “AI risks” and “Programmatic advantage”, April 21, 2026. The publisher’s rounded comparative multipliers aren’t needed to make the operational point.

A small team doesn’t need an enterprise program to act on this. It needs an owner, an understandable reproduction, a deployment-specific impact judgment, and a reserved retest step. Scheduling more discovery while leaving every fix unowned can make the report queue larger without changing customer exposure.

Faster reported resolution can coexist with a growing backlog

HackerOne’s Security Research Report 2026, tenth edition, makes the distinction unusually visible. Its current public page describes 92,194 valid vulnerability reports and $89.1 million in bounty payouts for July 2025–June 2026. The exact launch day isn’t established; the 2026 edition was available when verified on October 7. [S24]

Table 4. HackerOne’s 2026 remediation and survey results. The current corpus is not automatically the denominator for historical comparisons.

ID Published result Denominator / comparison Limitation
S25 Critical mean resolution time fell 61% Platform critical-report resolution populations over two years Exact eligible n, resolved-only eligibility, and censoring rules unverified.
S26 Open valid backlog rose 131% Open valid-report counts, 2024 versus 2026 Absolute snapshot counts unavailable; discovery/program growth can contribute.
S27 Open critical backlog rose 123% Open critical-report counts over two years Not a rate of newly vulnerable applications.
S28 All-severity mean resolution: 135 → 62 days Platform reports used for resolution averages Exact eligible n, resolved-only eligibility, and censoring rules unverified; not Cobalt’s half-life.
S29 Discovery outpaces fixes: 70% Survey of 111 security leaders; question-level n unavailable Self-report from $100M+ revenue organizations, not small SaaS.
S30 Critical validation takes 30 hours, average Same leader survey; per-question n unavailable Validation is a separate stage before repair.
S31 Valid AI report counts rose 222% Counts in AI classifications over two years Absolute category counts and tested-asset denominator unavailable.

Source: HackerOne, Security Research Report 2026, highlights and methodology FAQ, available October 7, 2026. Surveys were fielded in summer 2026: 111 enterprise leaders across six countries and 408 active researchers. Independent yearly survey samples aren’t directly comparable.

A resolution-time average and an open-backlog count answer different questions. If the average covers resolved reports, it describes how quickly those represented reports closed; the publisher’s public summary does not establish the exact eligibility or censoring rules. The open backlog counts valid findings that remain unresolved.

Under a resolved-only calculation, a team could improve the average by fixing easy, newly discovered issues while older, difficult findings remain. That’s a hypothetical selection effect, not a verified explanation of HackerOne’s results. The published summaries don’t establish whether it occurred or quantify its contribution; the reported time and backlog trends nevertheless show why both metrics deserve attention.

Track open high-impact findings by age, new validated findings, verified closures, and reopened issues. Distinguish a duplicate report from a unique root cause. Count uncertain findings separately. A hundred scanner alerts aren’t a hundred exploitable defects, and a hundred closed alerts aren’t evidence that a tenant boundary now holds.

AI code statistics are a different experiment again

Veracode’s July 28, 2026 summary of its 2026 GenAI Code Security Report reports a 56% average security pass rate and roughly 44% risky outputs on selected generation tasks, not pentesting results. Its report page explicitly attaches the 56% average to more than 100 models over four years, separately describing 11 newly tested models across 80 tasks. Exact eligible-output counts and aggregation weights remain unavailable. The 56% isn’t established as an exclusive Summer 2026 cohort rate, and 11 × 80 isn’t its verified denominator. [S32]

Table 5. Security-task pass rates published by Veracode in 2026. Each weakness row measures the relevant task/model evaluations; exact per-category n is unavailable.

ID Metric Published pass rate What it doesn’t establish
S33 GPT-5.5, Summer 2026 snapshot 68% Not a ranking of newer October models or real SaaS features.
S34 SQL injection tasks 83% Not the safety rate of all generated database code.
S35 Cryptographic-algorithm tasks 87% Not a guarantee of correct key lifecycle or protocol design.
S36 Cross-site scripting tasks 15% Not the fraction of every AI-written web component that’s safe.
S37 Log injection tasks 12% Selected test conditions, not incident prevalence.

Sources: Veracode’s dated report article, “Where vulnerabilities persist” and Summer 2026 methodology outline. Published July 28, 2026; verified October 7.

These results motivate contextual review: where does untrusted text become HTML, a query, a log record, or a privileged tool argument? They don’t support multiplying a generation-task failure percentage by a breach percentage. Mandiant explicitly says it doesn’t consider 2025 the year breaches were directly caused by AI. Attacker use of AI and AI as the direct cause of a breach are different claims. [S38]

For the generating agent’s permissions and untrusted-input boundary, see coding-agent security. For discovery and exploitation performance, use the separate AI pentesting leaderboard. Writing secure code, identifying a bug, exploiting it, and repairing it are different tasks.

Turn the evidence into a small-SaaS remediation contract

The useful response is a testable workflow, not a universal budget percentage. Consider a hypothetical workspace app adding invitations, file exports, and customer API tokens. The risky question is whether the backend permits one tenant’s member to act on another tenant’s object. An anonymous homepage scan cannot answer that question because it never exercises the authorized session and wrong-object combination.

OWASP API1:2023 requires action-specific object authorization. WSTG 4.2, WSTG-ATHZ-04 recommends at least two users, often more, with different objects and privileges. These are standards and procedures, not prevalence measurements. [S39]

Use two synthetic tenants and the roles your application actually supports. Preserve an allowed read as a positive control. Attempt the prohibited read, update, export, download, and bulk operation using the other tenant’s known object identifiers. Include intended sharing exceptions. The multi-tenant authorization testing guide expands that matrix.

For each confirmed finding, retain the actor, target object, request sequence, observed response or state change, and deployment revision. Explain why the action should have been denied. Give the engineer a root-cause location and acceptance condition. If evidence isn’t reproducible, label the uncertainty rather than inflating severity to secure attention.

After the fix, replay the forbidden action and verify the legitimate action still works. Test adjacent entry points sharing the same data-access code. For a leaked credential, confirm revocation rather than merely deleting the string from a repository. For a session flaw, replay the old token after the policy’s revocation event. Closure is a behavioral claim.

This is an engineering recommendation derived from the evidence, not a published effectiveness percentage. Pensec is an early-access project; this article doesn’t claim a shipped scanner or completed customer assessments. The same contract is usable with a specialist pentester, your own regression suite, or an existing testing service.

Evidence ledger and refresh rules

The source families below keep the statistics auditable without pretending public summaries disclose every denominator. All were retrieved 2026-10-07. Exact claims, protocols, missing fields, and fetch notes are maintained in docs/blog/claims-research.md.

IDs Primary publication / location Evidence type and date
S01–S09 Verizon executive summary, pp.5, 9–11, 16; FAQ 2026 recorded breach corpus; exact launch date unverified; PDF revision June 5.
S10–S17, S38 Mandiant M-Trends 2026, numbers, SaaS identity, AI sections March 24, 2026; 2025 incident-response observations.
S18–S23 Cobalt State of Pentesting 2026, Key Findings and FAQ April 21, 2026; five-year findings corpus and separate survey.
S24–S31 HackerOne Security Research Report 2026, highlights and FAQ July 2025–June 2026 platform data; summer surveys; launch day unknown.
S32–S37 Veracode article and report July 28, 2026; selected code-generation tasks.
S39 OWASP API1:2023 and WSTG 4.2, linked above Normative guidance; no statistical population.

Refresh when an edition or methodology changes. Keep the original population beside every number, preserve unresolved source conflicts, and don’t derive a pentesting market size from vulnerability counts. For your own app, the next useful number is simpler: how many verified high-impact findings remain open, and who owns the next fix?

About Gabe

@bucabay

Gabe writes about website pentesting, security research, and coding-agent security at Pensec.

More from Gabe