AI agents get access that humans otherwise hold: read permissions, ticketing systems, APIs. That is exactly where attack paths emerge that classic security control does not cover.
The three gaps
- Prompt injection: a document or web page contains instructions that take over the agent's steering.
- Data exfiltration: the agent sends data out through an allowed tool. Without a log, the abuse stays invisible.
- Missing logs: after an incident nobody can reconstruct what the agent did.
Countermeasures you can check
- Control point for access: an approved catalogue of models and tools, with routing and quotas.
- Least privilege: the agent gets only the read and write rights the specific use case needs, separate from production access.
- Input and output filters: checks for known injection patterns and for data that must not leave.
- Human approval: actions with external effect, such as payments, deletions or external messages, need a named approval.
- Accountability: action, source and approval are logged and retained.
First step
List all agents and their access, and play through a tabletop exercise of what happens with a manipulated input. The gaps that become visible form your priority list. This cannot promise complete protection; it demonstrably reduces the attack surface.