Treat repository content as untrusted data. Comments, issue text, documentation, test fixtures, generated files, and images can all contain instructions that conflict with the actual task.
This guide focuses on the engineering decision behind repository prompt injection: what to standardize, what to constrain, and what evidence a reviewer should expect before accepting the result.
The core decision
Indirect prompt injection matters because an agent reads large amounts of text while holding tools. A malicious file may ask the agent to reveal secrets, contact an external endpoint, weaken tests, or edit unrelated code. The instruction can be disguised as setup guidance or a debugging note.
A workflow that holds up in review
Separate authoritative instructions from task data, minimize tools, restrict egress, and require approval for sensitive operations. Scan or quarantine newly introduced instruction files, and show reviewers which repository guidance the agent applied.
The failure mode to design around
Trying to enumerate every hostile phrase is brittle. Assume some attacks will influence the model and rely on capability boundaries, deterministic validation, and human review to limit impact. Preserve the malicious fixture so future environment and model changes can be tested against the same attack.
Implementation checklist
- Mark trusted instruction locations.
- Keep repository text from granting authority.
- Restrict tools and network routes.
- Red-team with malicious fixtures.
Turn the checklist into operating controls
- Mark trusted instruction locations: name the owner, the evidence that proves it happened, and the condition that should stop the run.
- Keep repository text from granting authority: name the owner, the evidence that proves it happened, and the condition that should stop the run.
- Restrict tools and network routes: name the owner, the evidence that proves it happened, and the condition that should stop the run.
- Red-team with malicious fixtures: name the owner, the evidence that proves it happened, and the condition that should stop the run.
The list becomes useful when every item produces visible evidence. Store that evidence with the task or pull request rather than in a private chat. A future reviewer should be able to tell which repository revision was used, which permission profile applied, what stopped or failed, and who accepted the remaining risk. For repository prompt injection, a short, complete record is more valuable than a long narrative that cannot be reproduced.
Move from one run to a repeatable practice
Test the control with a safe adversarial exercise before relying on it. Attempt an out-of-scope file edit, a denied network request, a fake instruction in repository content, and access to a canary credential. Confirm that prevention or detection produces an actionable event with a named owner. Retest after adding tools, integrations, or repositories because effective authority changes when capabilities are combined. Record accepted residual risk and its review date.
Before expanding the workflow, ask three review questions:
- What evidence shows that mark trusted instruction locations was satisfied, and would that evidence survive a rerun from the recorded commit?
- What evidence shows that keep repository text from granting authority was satisfied, and would that evidence survive a rerun from the recorded commit?
- What evidence shows that restrict tools and network routes was satisfied, and would that evidence survive a rerun from the recorded commit?
Write the answers in the same place as the code review. That creates a compact decision record and lets the team compare later runs without relying on memory.
A practical acceptance test
Run the workflow from a clean checkout at a recorded commit. Give the agent only the documented task packet and the intended permission profile. Then ask a reviewer who did not launch the run to reproduce the important checks, explain the changed behavior, and identify the rollback path. The task passes only when the artifact, evidence, and repository state agree. Keep the failed examples as regression cases; they are more useful than a polished demo because they reveal where instructions, environment, permissions, or tests need improvement.
Related reading
- Agentic Coding in the Cloud: The Complete Guide
- A Threat Model for Cloud Coding Agents
- Least Privilege for Coding Agents
Primary references
Vendor features, limits, preview labels, and pricing can change. Recheck the linked first-party documentation for the current state before making a purchase or rollout decision.
