AI Security

Securing AI Agents: Tool Permissions Are the Blast Radius

An AI agent that gets manipulated can only do what its tools allow. That makes the permission list the single most important security decision in any agent deployment.

Team reviewing an AI agent deployment

Most discussion of AI agent security focuses on whether the model can be tricked. It can. Treat that as settled and move to the question that actually determines your exposure: when it is tricked, what is it able to do?

The agent is not the vulnerability

A language model that produces a bad decision is harmless on its own. It becomes a security problem the moment it is wired to tools that act on the world: an API that moves money, a mailbox that sends messages, a database it can write to, a shell it can execute in.

So the security review that matters is not of the model. It is of the tool list, and of what each tool permits in the worst case rather than the intended case.

Audit each tool against the worst case

For every tool your agent can call, write down the answer to four questions:

  • What is the most damaging single call? Not the expected use — the worst one the parameters allow.
  • Is it reversible? Sending an email is not. Deleting a record may not be. Reading data you then leak is never reversible.
  • What does it have access to? A "read the database" tool that can read any table is a different risk from one scoped to three columns.
  • Would a human notice? If the answer is "eventually, in a monthly report", you have no detection.

The scoping mistakes we find most often

Agents are routinely given credentials far broader than their function needs, because narrowing them is fiddly and the broad credential already existed. Specifically:

  • Database users with write access to tables the agent only reads.
  • API keys with full account scope when the agent uses one endpoint.
  • Mail permissions that allow sending to any recipient, when the agent only ever replies within a thread.
  • Cloud roles inherited from the service the agent runs inside, which is usually far more privileged than the agent.

Approval gates, placed by consequence

Not everything needs a human in the loop, and demanding approval for everything trains people to click through without reading. Place gates by consequence:

  • Autonomous: reads, searches, draft generation, anything internal and reversible.
  • Logged and reviewable: writes to systems where an error is recoverable from backup.
  • Human approval required: anything that moves money, sends external communication, changes permissions, or deletes data.

Make the approval prompt show what will actually happen — the recipient, the amount, the record — rather than "the agent would like to use the payments tool". An approval the human cannot evaluate is not a control.

Assume the input is hostile

Any content the agent reads that originated outside your trust boundary should be treated the way you treat user input in a web application. A support ticket, a CV, an invoice, a web page, an email — all of them can carry instructions aimed at your agent rather than at your staff.

You cannot reliably sanitise natural language the way you escape SQL. What you can do is ensure that if the agent is persuaded, the tools cannot do anything that matters.

Log it like a user

Every tool call, with its parameters, the content that prompted it, and the outcome. When something does go wrong, the investigation needs the same evidence it would need for a compromised employee account — because functionally, that is what you have.