Security, examined.

Website pentesting, research, and agent security. Follow the evidence, inspect the boundaries, and explore the examples.

Subscribe via RSS

5 articles

  1. Identify population
  2. Check denominator
  3. Apply finding
An evidence-reading workflow: identify who or what was observed, check the metric’s denominator and limitations, then decide what the finding supports for your application. These audit steps are not an empirical funnel linking breaches, pentest findings, and fixes.
Research

Penetration testing statistics in 2026: findings, breaches, and the fix backlog

The latest breach and pentest studies point to a practical problem for small SaaS teams: finding vulnerabilities is only useful when engineers can verify and close them. Here are the numbers, their denominators, and the decisions they support.

Gabe13 min read
  1. Untrusted context
  2. Scoped tools
  3. Verified change
Repository content informs a task. It should not grant credentials, expand tool permissions, or approve its own release.
Agent security

Coding-agent security: permissions, review, and release evidence

A coding agent can help a small SaaS team ship, but its generated code and its tool access need separate security checks. For founders preparing a rollout, this guide connects secure-code research to practical controls for untrusted repository data, isolated permissions, and deployed fix verification.

Gabe9 min read
  1. Task and evidence
  2. Model plus harness
  3. Verified outcome
A score belongs to a task version, information policy, deployed model and harness, resource budget, and success criterion. Changing any of them changes the comparison.
Research

AI pentesting leaderboard, October 2026: rank the benchmark, not the brand

GPT-6 Astra leads OpenAI’s September exploitation comparison; GPT-5.4 with Codex leads the verified cost-capped CyberGym-E2E cohort. This reference separates those rankings from restricted new models, specialized agents, and web-app testing claims.

Gabe17 min read
  1. Verified membership
  2. Object and action
  3. Allowed or denied
Tenant selection, object ownership, and operation permission are separate inputs to one server-enforced decision.
Pentesting

Multi-tenant authorization testing: prove the boundary, then retest it

For a small B2B SaaS team shipping workspaces or permissions, multi-tenant authorization testing needs more than a tenant ID. Build a matrix of actors, objects, roles, and operations; prove both intended access and denied access; then repeat the checks against the deployed fix.

Gabe10 min read