So How Do We Secure AI Agents?
ABRI Systems ·
The first two articles examined different failure paths: an attacker obtaining an agent’s authority, and an authorized agent acting beyond its owner’s expectations. Both point to a practical design question: how do we control what an agent can access, decide and change?
ABRI’s working model is Identity → Authority → Action → Audit. It is an organizing framework for security reviews, not a certification or a guarantee. Use it to connect familiar access controls with the additional risks introduced by agents.
Identity: know which agent is acting
Give each production agent a distinct identity and an accountable owner. Record whose request it is carrying out. Avoid shared administrator accounts and credentials that cannot be revoked independently. Keep secrets in a protected credential service rather than prompts, conversation history or editable memory.
Authority: grant only what the task needs
Use least privilege across accounts, data and tools. An agent reading support tickets should not inherit the entire customer database. Prefer short-lived credentials and narrow scopes where the service supports them. Recheck permissions when the task changes, and remove access when it ends.
Choose small, defined tools such as searching approved documents or creating a draft. Restrict arbitrary command execution, available network destinations and connector methods. Sandbox necessary code execution and prevent the agent from modifying the policy that grants its own access.
Action: put approval at the point of consequence
Separate recommendations from execution. Require explicit human approval for high-impact actions such as transferring money, deleting records, deploying code, changing permissions or disclosing sensitive data. Bind approval to the concrete recipient, payload, amount and destination; recheck it immediately before execution.
Enforce these rules outside the model. OWASP’s excessive-agency guidance recommends limiting tool functionality and permissions, applying authorization in downstream systems, and requiring human approval for high-impact actions. Source: OWASP, LLM06:2025 Excessive Agency.
Treat external content and memory as security boundaries
A webpage, email or document can contain instructions planted by someone else. That is the indirect prompt-injection problem: task data can try to become authority. Preserve its provenance, label external content as untrusted, and prevent it from granting new permissions. Detection and model training can help, but neither removes the need for tool and data controls.
Apply the same discipline to persistent memory. Isolate users and tenants, limit who can write memories, retain their source and timestamp, and make suspicious entries removable. Stored text must not silently become an instruction that overrides policy. Protect memory with access controls, appropriate encryption and retention limits. After an incident, inspect it before allowing the agent to resume.
OWASP’s agentic threat guide provides a broader threat-model reference for reviewing agent systems. Meta’s published Muse design also illustrates separating runtime code, credentials and permission enforcement; it acknowledges that prompt injection remains an open problem. Source: OWASP, Agentic AI — Threats and Mitigations; vendor example: Meta AI Research. The operational checklist here is ABRI’s recommended application of these principles.
Audit: reconstruct what happened
Log the requester, agent identity, tool, permission decision, approval reference, affected resource and outcome. Use correlation identifiers to follow one task across services. Protect logs from alteration and restrict access to them. Record enough to investigate, while redacting tokens and avoiding unnecessary copies of private data.
Stop quickly and limit the blast radius
Build and test a kill switch that stops queued and running tasks, disables the agent identity and revokes connected tokens or sessions. Rotate exposed credentials and review persistent memory before restarting. Assign responsibility for each step so incident response does not depend on finding the person who first configured the assistant.
Set spending and transaction caps, rate limits, maximum tool calls, execution timeouts and outbound network restrictions. Separate development from production and keep backups for recoverable changes. These controls bound the damage from a compromised agent, an unintended action or a runaway loop.
Start with one task
Choose a limited workflow and document its identity, data, tools, permitted actions, approval points, logs and shutdown procedure. Test malicious instructions and unintended disclosures before expanding access. An agent is ready for more responsibility when its boundaries are observable and enforceable—not merely when its answers sound confident.
Sources checked October 1, 2026. Incident reporting is attributed above; recommendations and analysis are ABRI’s.
Share