Most conversations about AI agent risk start in the wrong place. They start with the model.
The model is rarely the problem. The problem is that an agent built for one purpose, by one team, six months ago, is still running — with credentials nobody has reviewed, against data nobody mapped, at a cost nobody tracks, and with no name attached to it.
That is not a model problem. It is an operations problem. And operations problems can be measured.
Below are the ten domains we score in an Agent Readiness Audit. They group into four layers. The order matters, and we will come back to why.
Layer one: control
Can you see your agents, and can you steer them?
1. Visibility and inventory
Can you produce one list of every AI agent running in your business that everyone agrees is correct?
Almost nobody can. Agents arrive through different doors — a vendor feature, a developer’s script, a department’s pilot — and no single system sees them all. Until this domain is solved, every other score is an estimate.
2. Lifecycle management
What happens to an agent between “someone had an idea” and “it is switched off for good”?
Most businesses have a strong process for one end and nothing for the other. Agents go live without a decision and stay alive without a review. The failure mode is quiet: an agent nobody owns, doing something nobody checks, long after the reason for it disappeared.
3. Identity and access
An agent is not a person, but it holds credentials like one — often broader ones, because narrowing them was harder at the time.
We score how agent identities are issued, scoped, rotated and revoked. The common finding is over-permissioning: an agent that reads a whole database because filtering to the three fields it needed would have taken another day.
Layer two: safety
Is the agent behaving, and would you know if it were not?
4. Security and attack surface
Agents introduce attack paths traditional software does not have. Prompt injection. Tool misuse. Poisoned retrieval sources. Compromised components in the supply chain.
We score against the OWASP Top 10 for Agentic Applications and check whether anyone has ever tested your agents adversarially — not whether they work, but whether they can be made to misbehave.
5. Runtime behaviour and safety
Does anyone watch what your agents actually do, while they do it?
This domain covers monitoring, drift detection, guardrails, escalation paths, and whether a kill switch exists that has been tested rather than assumed. Two independent ways to stop an agent, because the first one may be the thing that failed.
6. Data and context
What does the agent read, what does it remember, and what can it disclose?
Agents accumulate context. They cache, they retain, they build memory. We score whether data flowing into agents is classified, whether memory can be wiped and rebuilt from clean sources, and whether anyone has tested what an agent will say if asked the right wrong question.
Layer three: cost
7. Cost and value
Two questions, and most businesses can answer neither. What is each agent costing? And what is it worth?
The first is hard because model spend arrives as one undifferentiated bill. The second is harder, because value was rarely defined at go-live. This is where the recursion loop that quietly triples a monthly bill lives — and where the agent that costs real money to run something nobody uses hides in plain sight.
Layer four: accountability
When it goes wrong, whose name is on it?
8. Ownership and accountability
Every agent needs one accountable human. Not a team, not a function, not a shared mailbox.
We score whether ownership is recorded, current, and survives people leaving. The orphaned agent — still running, original owner gone two years — is one of the most common findings, and one of the easiest to fix once you can see it.
9. People and organisation
Do the people deploying agents know the rules?
This covers training, escalation paths, and whether there is a route for someone to raise a concern about an agent without it becoming a confrontation. Governance that exists only in a document is not governance.
10. Compliance, legal and vendor
Where do you stand against ISO/IEC 42001, the NIST AI Risk Management Framework, and whatever regulation applies to you?
This domain also covers vendor terms — what your model providers may do with your data — and contractual liability when an agent acts on a customer’s behalf.
How the scoring works
Each domain scores 1 to 5.
| Score | Meaning |
|---|---|
| 1 | Absent — nothing exists |
| 2 | Aware — the gap is known, nothing is in place |
| 3 | Partial — something exists, inconsistently applied |
| 4 | Managed — a defined process, followed |
| 5 | Optimised — measured and improving |
Weighting is agreed with you before scoring starts, and published with the results. A regulated business weights security and compliance higher. A young AI team weights cost higher. Agreeing weights up front means nobody argues with the radar afterwards.
The order matters more than the scores
There is a dependency between these domains, and it defeats most improvement plans.
Fix visibility first. Always.
You cannot manage lifecycle for agents you have not found. You cannot review access for identities you have not listed. You cannot allocate cost to agents you cannot name. You cannot prove compliance for an estate you cannot enumerate.
Every domain past the first is partly a function of the first. Which is why an audit that returns ten scores and no sequence is not much use — and why our plans come in dependency order rather than in order of severity.
The most alarming score is not always the one to fix first.
Want to know how your business would score?
The Agent Readiness Audit takes four to six weeks and gives you an inventory, a score across all ten domains, and a plan in the order the work actually needs doing.
