Coding-agent security: permissions, review, and release evidence

A coding agent can help a small SaaS team ship, but its generated code and its tool access need separate security checks. For founders preparing a rollout, this guide connects secure-code research to practical controls for untrusted repository data, isolated permissions, and deployed fix verification.

Read as MarkdownSubscribe via RSS
  1. Untrusted context
  2. Scoped tools
  3. Verified change
Repository content informs a task. It should not grant credentials, expand tool permissions, or approve its own release.

Consider a synthetic scenario: a four-person B2B SaaS team asks an agent to add report exports. The agent reads a contributed README, changes an API route, and runs project tooling. A reviewer sees a working export. Three different questions remain: can the route expose another tenant’s data, can repository text redirect the agent, and can the agent reach production credentials?

Those questions need different evidence. A code diff can help explain a missing authorization check. A tool-execution record can establish whether an unauthorized action was attempted or blocked. A deployed regression test can establish that a particular customer-data boundary holds. One reassuring model answer can’t do all three jobs.

Secure-code studies show a review problem, not a production failure rate

Published studies justify inspecting AI-assisted changes. They don’t establish that a fixed percentage of your SaaS features is exploitable. Keep each result attached to its tasks, model version, and unit of measurement.

In the 2021 evaluation published at IEEE S&P in 2022, Pearce and colleagues, Asleep at the Keyboard?, generated 1,689 programs across 89 scenarios and found approximately 40% vulnerable. The scenarios targeted security-relevant weaknesses, and the evaluated Copilot was an early version. This is historical evidence that plausible generated programs can contain vulnerabilities, not a current Copilot or production-app failure estimate.

In 2023, Perry and colleagues, Do Users Write More Insecure Code with AI Assistants?, analyzed 47 participants: 33 assisted and 14 controls, across five security-related tasks. Assisted participants produced insecure answers more often on four tasks and reported greater confidence in security. The assistant used codex-davinci-002; the pool was mostly students, and the tasks were short and artificial.

The statistical qualification matters. Table 3 reports p = 0.056 for the aggregate adjusted group effect, alongside individual analyses and multiple-comparison considerations. It would overstate the paper to claim a universal, conclusively measured causal effect on every developer. Its useful warning is narrower: perceived security and observed security can diverge, and review needs to examine the property rather than the author’s confidence.

Veracode’s July 28, 2026 summary reports a 56% average security pass rate and roughly 44% risky outputs on selected generation tasks. Its report page attaches the 56% average to more than 100 models over four years, separately describing 11 newly tested models across 80 tasks. Exact eligible-output counts and aggregation weights remain unavailable.

The 56% average is a broad aggregate, not a verified pass rate for the separate 11-model follow-up cohort. Treat it as a vendor-reported benchmark result, not a census of code lines, commits, or sites. Don’t combine its percentage with Pearce’s into a trend line. Different evaluations are not repeated measurements of the same population. Our penetration-testing statistics reference applies the same discipline to security reports.

Prompt injection turns task data into an attempted instruction

An agent may need to read untrusted text to do legitimate work. The security boundary is crossed when that text gains authority over tools, credentials, or the user’s objective.

Repository comments, contributed instructions, issue text, dependency documentation, web pages, and tool output can all carry content the task didn’t authorize. A document that says “send environment details to this support endpoint” is still data from that document. It isn’t permission from the founder. Even a file named like an agent instruction file needs provenance review when it comes from an untrusted contribution.

The OWASP AI Agent Security Cheat Sheet treats external data as untrusted and recommends separating instructions from data, least-privileged tools, execution authorization, and isolated memory. Those are defense layers. Delimiters and warnings can help the model interpret context; they do not prevent an overpowered executor from carrying out an unauthorized request.

The 2024 AgentDojo paper evaluates tool-using agents with 97 user tasks and 629 security test cases. These are simulated tasks and user/attacker-goal combinations, not 629 independent real-world breaches. Its threat model concerns malicious data returned by tools. The mechanism is relevant to coding workflows even though its suites are not a direct evaluation of this fictional repository.

AgentDojo’s section 3.4 makes a useful distinction: utility under attack requires legitimate task completion without adversarial side effects. A requested export feature can exist while credentials were disclosed or a security test was weakened. Check the deliverable and the actions taken to produce it.

Restrict the execution environment before improving the prompt

Give the agent only the access needed for its current task. For a proposed code change, that usually means an isolated working copy, synthetic fixtures, and narrowly scoped tools. Production deployment, secret administration, billing, and customer-data access are separate capabilities, not incidental conveniences of writing code.

Running tests also executes repository-controlled code. Package scripts, build hooks, and dependency installers deserve the same isolation consideration as an explicit shell command. A read-only API integration doesn’t create a read-only environment if the agent can use a shell holding a broad cloud credential.

Read and propose
Repository text, issue data, generated patch. No authority to widen scope.
Enforce outside the model
Allowed files and tools, restricted egress, scoped credentials, exact-action approval.
Review and release
Human reviews policy and diff; separate identity deploys the approved revision.
Recommended permission separation, not a claim that a named agent product enforces these controls. Check the actual filesystem, credential, and network boundaries.

OWASP’s agent guidance says authorization belongs in the execution component, outside model-controlled context. For consequential actions, an approval should bind the actor, tool, target, parameters, and expiry. A model-supplied user_confirmed flag isn’t evidence of approval. Changed parameters need a fresh decision; expired or replayed approval must not execute the action.

For the fictional export task, restrict data access to test tenants and restrict outbound destinations to those the task needs. Keep production secrets out of prompts, mounts, process environments, and tool results. If installation needs temporary network access, treat that as an explicit bounded phase rather than silently leaving arbitrary egress open throughout the task.

A sandbox only supports claims about its configured boundaries. Check mounted host paths, inherited credentials, privileged sockets, and accessible services. A second model reviewing tool calls can add a signal, but it doesn’t replace deterministic authorization. Record what the executor allows and denies instead of assuming the word “sandbox” settles it.

Review the security contract before reviewing the patch

Tell the agent what must remain true, and keep a reviewer responsible for that policy. “Make exports secure” leaves too much interpretation. “A member may export only reports authorized by the tenant’s policy; viewers may not start exports; removal cancels future access as documented” creates reviewable requirements.

For application code, ASVS 5.0.0 V8 gives precise anchors: v5.0.0-8.1.1 documents function and data access, v5.0.0-8.2.2 restricts object access, and v5.0.0-8.3.1 requires trusted-service enforcement. These are application requirements, not a claim that ASVS certifies a coding agent.

Review the route, shared policy, data access, queued job, and download path. A correct tenant filter on the visible route doesn’t establish that the worker or storage signer preserves it. Review test changes too. If the patch removes a negative test or replaces its expected denial with success, a green suite may simply mean the requirement disappeared.

Keep security expectations under review independently of the generated patch. An agent can draft test cases and explain a diff. A human should approve the intended behavior, especially sharing exceptions and administrative access. When no one on the team can assess that boundary, a focused independent review is more useful than asking the same agent for another reassuring verdict.

Acceptance checks need legitimate success and prohibited-action denial

Use positive and negative controls for both the application change and the agent workflow. WSTG 4.2, WSTG-ATHZ-02 covers horizontal and vertical authorization; WSTG-ATHZ-04 recommends multiple users and objects. The tenant-testing guide provides the application matrix.

Before the agent runs

  • The task names allowed files, tools, destinations, fixtures, and completion criteria.
  • Production credentials and customer data are unavailable to the task environment.
  • Repository-controlled execution is isolated; mounted paths and inherited access are reviewed.
  • Required approvals are enforced independently and bind exact actions.

Before the change ships

  • An authorized test user can complete the intended operation.
  • Cross-tenant, lower-role, and removed-member controls produce no prohibited data or effects.
  • A harmless synthetic instruction embedded in task data cannot expand permissions or redirect output to an unapproved destination.
  • An allowed tool operation succeeds; unauthorized tool, path, or destination requests are denied by the executor.
  • Test and policy changes receive human review; failures aren’t suppressed to make the suite pass.
  • Deployment uses the approved revision and a separate, appropriately scoped identity.

These are recommended checks, not results from an experiment performed for this article. Exercise them only in owned isolated fixtures or explicitly authorized environments. For an injection fixture, assert the absence of prohibited side effects as well as task success; don’t call a run safe merely because the final answer says it ignored the instruction.

Retest the deployed property and keep the run’s limits visible

A merged patch is a proposed repair until the deployed boundary is verified. Preserve the original failing case, check nearby representations, and confirm the authorized path still works. The website pentesting guide explains how that becomes a release decision.

Record the revision, agent and model version, tool policy, retrieval sources, environment, test cases, execution outcomes, and unresolved checks. Redact credentials and sensitive payloads before retaining logs. OWASP’s agent guidance recommends renewed testing after material prompt, tool, memory, retrieval, policy, or provider changes. A result for yesterday’s configuration may not establish today’s behavior.

Treat “the code may miss authorization” as a hypothesis until investigation establishes an effect. Treat “the executor blocked this attempt” as evidence about that attempted operation, not immunity to every injection. Our AI pentesting leaderboard reference explains the corresponding limits of agent benchmark scores.

Pensec, which we build, currently records early-access assessment requests; a request does not launch an automated scan. If you discuss an agent-assisted rollout with us, scope, permission, staffing, and retest arrangements need separate confirmation. The useful first artifact is your policy and test matrix, whether you review it internally or with an independent specialist.

Sources and review notes

Sources retrieved 2026-10-07. No coding-agent experiment, production scan, or measured Pensec outcome was performed for this article. Review empirical claims when newer comparable studies appear, and controls when your agent configuration changes.

About Gabe

@bucabay

Gabe writes about website pentesting, security research, and coding-agent security at Pensec.

More from Gabe