Home  /  Articles  /  The Sixth Lesson About Artificial Intelligence: Why It Keeps Failing at the One Job You Hired It For
6 Fintechminute read

The Sixth Lesson About Artificial Intelligence: Why It Keeps Failing at the One Job You Hired It For

The first time I watched an "AI-powered" reconciliation tool hand back a number that was confidently, cleanly, professionally wrong, I assumed I had configured it badly.

A photorealistic close-up of a financial auditor's desk in a quiet corporate office at…

The first time I watched an "AI-powered" reconciliation tool hand back a number that was confidently, cleanly, professionally wrong, I assumed I had configured it badly. The second time, I assumed the data was dirty. The third time — same ledger, same prompt, a different total — I understood that the problem was not my configuration or my data. The problem was that I had asked a tool built to be plausible to be correct, and those are not the same job.

If you work anywhere near regulated finance, you have probably asked yourself a version of this question: why does artificial intelligence that can draft a flawless client email keep getting my numbers wrong? You are not being unfair to the technology. You are noticing something real about it, and almost nobody selling it to you wants to name what you noticed.

The question under the question

The way the question usually gets asked in a boardroom is "Is AI ready for compliance work yet?" — as if there is one substance called AI that ripens over time like fruit, and we are all waiting for the same harvest. That framing is the source of the confusion. It treats a category as a product.

So let me grant the part that is true, because it is genuinely true. The generative models everyone has been talking about for the last few years are extraordinary at a specific kind of work: drafting, summarizing, translating, brainstorming, rewriting a contract clause in plainer English, turning a messy support ticket into a structured one. Anywhere the goal is a good answer rather than the answer — where a human will read the output and judge it — these systems are a leap, not a gimmick. I am not here to take that away from anyone.

But "a good answer" and "the correct answer" are different requirements, and most of the disappointment I have watched finance teams live through comes from buying a tool optimized for the first and assigning it the second.

What the auditor actually sees

Here is the distinction in plain language, and it is the most useful thing in this piece, so I will go slowly.

A generative model produces output by estimating what is likely to come next, given everything it has seen. That is not a flaw to be patched out in the next version; it is the entire mechanism. It is a probability engine. Ask it to summarize a quarterly report and probability serves you beautifully, because there are many acceptable summaries and you, the reader, are the judge. Ask it to total a column of 4,000 transactions and produce the figure that goes on a regulatory filing, and probability is exactly the wrong instrument — because there is only one acceptable answer, the answer is not a matter of taste, and the cost of being plausibly-close is a penalty, a restatement, or a fine.

When an auditor looks at a number, they are not asking whether it reads well. They are asking: can you show me, step by step, how you got here, and will the same inputs produce the same output every single time? A system that might return a slightly different total on Tuesday than it did on Monday fails that test before the conversation starts. Not because it is bad. Because it was never built to answer that kind of question.

The opposite kind of system — call it rule-bound, or deterministic, or just "the boring kind that accountants have quietly trusted for forty years" — does not estimate anything. Given the same inputs, it returns the same output, and it can show its work. It cannot write you a charming email. It can tell you, with a straight face and an audit trail, that the number is the number.

The mistake of the last few years was assuming the charming one had quietly absorbed the boring one's job. It hasn't.

So which one do I need? It depends — and here is on what

I want to be honest about the "it depends," because the people selling certainty in either direction are both wrong.

The line is not "creative work versus financial work." Plenty of financial work is judgment work. The line is about what happens when the answer is wrong, and who has to defend it.

The task What it actually demands The class of tool that fits
Draft a first-pass response to a customer complaint A good answer, human-reviewed Generative
Summarize 40 pages of a credit agreement for a meeting A good answer, human-reviewed Generative
Categorize a transaction so a human can spot-check it A good-enough answer, supervised Generative, with review
Calculate the figure on a regulatory filing The exact answer, reproducible, auditable Deterministic
Decide whether a wire breaches a sanctions rule A defensible answer with a traceable reason Deterministic
Reconcile two ledgers to the cent One answer, identical every run Deterministic

The overlap is real and I will not pretend it away. A lot of modern systems use a generative model to read the messy input — to figure out which field is the invoice date when every vendor formats it differently — and then hand the actual arithmetic and rule-checking to deterministic machinery. That is not a contradiction. That is the design that works. The error is letting the probability engine do the part where there is only one right answer.

Why the market keeps buying the wrong tool anyway

If the distinction is this clear, why do so many regulated teams keep signing contracts for the mismatched thing?

Three reasons, and none of them are stupidity.

The first is the demo. Generative systems demonstrate spectacularly. They produce something on screen in seconds, and the something looks finished. A deterministic reconciliation engine demos like watching paint dry correctly. Procurement decisions get made in the room where the impressive demo happens.

The second is that "AI" became a single line item in the budget and a single box on the strategy slide. When the board asks "what is our AI plan," the easiest answer is to buy the thing everyone is talking about and point at it. Naming two different categories with two different jobs is a harder slide to build.

The third is the most human: the people who understand the compliance requirement and the people who are excited about the technology are frequently not in the same meeting. The wedge gets driven in there.

A short checklist before you assign a task to a model

Before you hand any task to an AI system, regulated or not, ask these four questions in order:

If a vendor cannot tell you which of these their product is built for, that is the answer.

Not two pillars. A handoff.

I have heard the truce described as "two equal halves of intelligence working in harmony," and that framing is too tidy to be useful. What actually happens in a working system is more like a relay. The generative side does what it is genuinely best at — making sense of unstructured, human-shaped mess. Then it hands the baton to something that does not guess, for the part where guessing is not allowed. The skill is knowing exactly where on the track the handoff happens, and refusing to let either runner carry the leg it will lose.

What this piece did not answer

I have stayed inside one question — which kind of system fits which kind of task — and I have left three larger ones untouched, on purpose, because they deserve their own honesty.

I did not tell you what any of this costs, or how vendor lock-in changes the math when the deterministic layer is proprietary and your audit trail lives inside someone else's product. I did not resolve who is liable when the rule-bound system is correct about the rule but the rule itself was coded wrong — a question that is going to define the next decade of compliance law more than any model release will. And I did not touch the governance problem of who, inside your organization, is allowed to decide where the handoff happens.

If you want the next thread to pull, pull that last one. The technology question is mostly settled once you stop treating artificial intelligence as a single substance. The question of who owns the line between plausible and correct is the one still wide open — and it is the one that will actually keep you up at night.

Fintech Compliance Ai