Skip to content Skip to footer

The 10 Domains of AI Agent Risk

Most conversations about AI agent risk start in the wrong place. They start with the model.

The model is rarely the problem. The problem is that an agent built for one purpose, by one team, six months ago, is still running — with credentials nobody has reviewed, against data nobody mapped, at a cost nobody tracks, and with no name attached to it.

That is not a model problem. It is an operations problem. And operations problems can be measured.

Below are the ten domains we score in an Agent Readiness Audit. They group into four layers. The order matters, and we will come back to why.

Layer one: control

Can you see your agents, and can you steer them?

1. Visibility and inventory

Can you produce one list of every AI agent running in your business that everyone agrees is correct?

Almost nobody can. Agents arrive through different doors — a vendor feature, a developer’s script, a department’s pilot — and no single system sees them all. Until this domain is solved, every other score is an estimate.

2. Lifecycle management

What happens to an agent between “someone had an idea” and “it is switched off for good”?

Most businesses have a strong process for one end and nothing for the other. Agents go live without a decision and stay alive without a review. The failure mode is quiet: an agent nobody owns, doing something nobody checks, long after the reason for it disappeared.

3. Identity and access

An agent is not a person, but it holds credentials like one — often broader ones, because narrowing them was harder at the time.

We score how agent identities are issued, scoped, rotated and revoked. The common finding is over-permissioning: an agent that reads a whole database because filtering to the three fields it needed would have taken another day.

Layer two: safety

Is the agent behaving, and would you know if it were not?

4. Security and attack surface

Agents introduce attack paths traditional software does not have. Prompt injection. Tool misuse. Poisoned retrieval sources. Compromised components in the supply chain.

We score against the OWASP Top 10 for Agentic Applications and check whether anyone has ever tested your agents adversarially — not whether they work, but whether they can be made to misbehave.

5. Runtime behaviour and safety

Does anyone watch what your agents actually do, while they do it?

This domain covers monitoring, drift detection, guardrails, escalation paths, and whether a kill switch exists that has been tested rather than assumed. Two independent ways to stop an agent, because the first one may be the thing that failed.

6. Data and context

What does the agent read, what does it remember, and what can it disclose?

Agents accumulate context. They cache, they retain, they build memory. We score whether data flowing into agents is classified, whether memory can be wiped and rebuilt from clean sources, and whether anyone has tested what an agent will say if asked the right wrong question.

Layer three: cost

7. Cost and value

Two questions, and most businesses can answer neither. What is each agent costing? And what is it worth?

The first is hard because model spend arrives as one undifferentiated bill. The second is harder, because value was rarely defined at go-live. This is where the recursion loop that quietly triples a monthly bill lives — and where the agent that costs real money to run something nobody uses hides in plain sight.

Layer four: accountability

When it goes wrong, whose name is on it?

8. Ownership and accountability

Every agent needs one accountable human. Not a team, not a function, not a shared mailbox.

We score whether ownership is recorded, current, and survives people leaving. The orphaned agent — still running, original owner gone two years — is one of the most common findings, and one of the easiest to fix once you can see it.

9. People and organisation

Do the people deploying agents know the rules?

This covers training, escalation paths, and whether there is a route for someone to raise a concern about an agent without it becoming a confrontation. Governance that exists only in a document is not governance.

10. Compliance, legal and vendor

Where do you stand against ISO/IEC 42001, the NIST AI Risk Management Framework, and whatever regulation applies to you?

This domain also covers vendor terms — what your model providers may do with your data — and contractual liability when an agent acts on a customer’s behalf.

How the scoring works

Each domain scores 1 to 5.

Score Meaning
1 Absent — nothing exists
2 Aware — the gap is known, nothing is in place
3 Partial — something exists, inconsistently applied
4 Managed — a defined process, followed
5 Optimised — measured and improving

Weighting is agreed with you before scoring starts, and published with the results. A regulated business weights security and compliance higher. A young AI team weights cost higher. Agreeing weights up front means nobody argues with the radar afterwards.

The order matters more than the scores

There is a dependency between these domains, and it defeats most improvement plans.

Fix visibility first. Always.

You cannot manage lifecycle for agents you have not found. You cannot review access for identities you have not listed. You cannot allocate cost to agents you cannot name. You cannot prove compliance for an estate you cannot enumerate.

Every domain past the first is partly a function of the first. Which is why an audit that returns ten scores and no sequence is not much use — and why our plans come in dependency order rather than in order of severity.

The most alarming score is not always the one to fix first.

Want to know how your business would score?

The Agent Readiness Audit takes four to six weeks and gives you an inventory, a score across all ten domains, and a plan in the order the work actually needs doing.

Author