Every vendor selling AI for finance right now is making the same promise: agents will see every leak and fix it, so leakage disappears. One markets that 80% of ledgers reconcile autonomously. Another says 87% of accounts payable flows through untouched.
That story is half right, and dangerously incomplete. Autonomous finance won’t end leakage, it will relocate it, from the seams between people to the seams between agents, where it moves faster and looks cleaner on paper.
The Assumption Nobody Is Testing
The agentic-finance pitch treats leakage as a visibility problem: humans are slow, agents are fast, faster wins. But leakage was never only about speed. It happens wherever a process makes a judgment without accountability, an exception approved, a match accepted, a tolerance stretched.
Move that judgment to an agent and the judgment doesn’t disappear. What disappears is the human who might have caught a bad one, and it happens at machine speed, across every transaction, not just the ones large enough for a person to review.
Why “87% Touchless” Is Also “87% Unreviewed”
Agents are tuned for throughput, because throughput is what “autonomous” sells. Throughput rewards approving, not questioning. A model that drifted last Tuesday can post cleanly to the general ledger for months afterward, and nobody notices, because the system reports success, not error.
Human leakage was slow and legible. You could pull a sample and find it. Agent leakage is fast, high-volume, and self-assured. It doesn’t look like a mess. It looks like a closed period.
Outcome-based pricing compounds this. Point a recovery agent at “dollars recovered” and it optimizes that number, not your actual economics. It will chase the claims that pay it, and “recover” things that were never leaks to begin with, and the pricing model quietly manufactures its own leakage while reporting it as a win.
The Honest Objection, And Why It Falls Short
The strongest pushback here is a fair one: agents are more auditable than humans. Every action gets logged. You can replay any decision, sample at 100% instead of 5%. Governance is a feature, not an afterthought.
Logs aren’t controls. A perfect, 100% record of decisions nobody independently checked is just a well-documented set of approved leaks. Observability was never the hard part. The hard part is having an independent standard, the encoded contract, the policy, the ground truth to judge the agent’s judgment against.
An audit trail nobody audits against a source of truth is theater. It’s the comfort of a camera with no one watching the feed.
The early evidence backs this up. Gartner expects more than 40% of agentic-AI projects to be cancelled by 2027, and has warned openly about “agent washing.” MIT research found 95% of enterprise GenAI pilots returning nothing to the P&L. Those aren’t just adoption stumbles, they’re the first signal that ungoverned agents can quietly destroy value while the dashboard reports success.
Why This Happened: The Contract Could Never Read Itself, Until Now
For thirty years, checking a promise against a payment took a human mind. Contracts, deal sheets, and penalty schedules lived as prose – PDFs, emails, memory. Payments lived as data – fast, structured, automated. There was no affordable way to hold the slow, written promise against the fast, moving payment. So businesses rationed human attention: analysts fought the big deductions, and everything small drained away.
What changed is narrower than the industry’s framing suggests. It isn’t that agents can now act. It’s that a promise written in prose can finally be turned into something a machine can check continuously, at near-zero cost per check, call it a truth layer. Not a smarter dashboard. The actual contracted terms, encoded, running against every transaction the moment it happens.
That’s genuinely new. But it cuts both ways: the same capability that lets an agent catch a bad deduction also lets an ungoverned agent auto-approve ten thousand bad ones before lunch, with total confidence.
What CFOs Should Check Before Trusting an Agent’s Numbers
- Re-read autonomy metrics as risk metrics. “87% touchless” is also “87% unreviewed.” Stand up a sampled, independent check of agent decisions against the encoded terms, and watch exception-approval rates the way you’d watch DSO, drift there is leakage forming.
- Govern the incentive, not just the output. If a vendor prices on outcomes, define the outcome as your net economics, not their recovery count, or you’ll pay a machine to optimize the wrong number.
- Ask what the agent is checked against. A log of decisions is not evidence they were correct. Ask specifically what independent, encoded ground truth every agent decision gets measured against, and who owns updating that ground truth when a contract changes.
- Assume your counterparties are automating too. Their agents learn what yours auto-approve, and route to it. Leakage becomes an exploit surface where the side with weaker controls funds the other.
The New Finance Control Function
The shift isn’t reconciling transactions faster. It’s governing the things that reconcile them. That means fewer people doing manual matching, and a smaller, more senior group doing three things instead: owning the encoded truth the agents run on, judging the genuinely ambiguous cases that deserve a human, and independently auditing the agents from the original facts, not trusting the machine’s own account of what it did.
We were promised autonomous finance would end the leak. It’s more accurate to say it moves the leak, into the machines, where it runs faster, sounds more confident, and arrives pre-reconciled. The CFOs who win the next decade won’t be the ones with the most agents. They’ll be the ones who can answer who audits the agents.
At Iksula, this is why our Leakage & Recovery Engine runs a fourth agent whose only job is to independently re-check the other three, never taking the AI’s own account of its work at face value. If you’re evaluating any agentic finance vendor, that’s the first question worth asking them.
Want to pressure-test your own agentic finance setup? and we’ll show you where an unaudited agent could already be costing you.
Frequently Asked Questions
Technically, yes — every action can be logged and replayed. But a log only proves a decision was made, not that it was correct. Real auditability requires checking agent decisions against an independent, encoded standard of what was actually agreed to, not just reviewing the log.
Because agents act at scale and speed, a single misconfigured rule or model drift can produce thousands of confidently wrong decisions before anyone notices, often without triggering any alert, since the system still reports high throughput and clean closes.
An audit trail documents what happened. A control catches whether what happened was correct. Without an independent check against encoded ground truth, a 100% audit trail is just a detailed record of approved mistakes.
When an AI vendor is paid based on a metric like “dollars recovered,” the agent optimizes for that metric, not necessarily for your actual financial accuracy. This can incentivize agents to flag non-issues as recoverable, inflating apparent value.
Ask what independent ground truth agent decisions are checked against, who owns updating that ground truth when terms change, and whether pricing incentives are aligned with actual net financial outcomes rather than a single output metric.
Adoption is real, but outcomes are mixed, Gartner projects over 40% of agentic AI projects will be cancelled by 2027, and research has found the large majority of enterprise GenAI pilots deliver no measurable financial impact, often due to missing governance rather than model quality.
Iksula’s Leakage & Recovery Engine includes a dedicated agent whose sole function is to independently re-verify the other agents’ decisions against encoded contract terms, rather than relying on self-reported logs alone.
