The AI Maturity Ladder for Finance: From Pasted Prompts to Governed Autonomy
Most vendors mention AI ethics in passing. Here is a four-stage maturity ladder for finance functions, with what breaks at each stage and what earns the next.
By The Rexfin team
Ask a finance leader how mature their AI adoption is and you will usually get a tools list: “we have Copilot, an analyst uses ChatGPT for drafts, we’re piloting an agent for reconciliations.” That is an inventory, not a maturity assessment. Maturity is not which tools are installed. It is whether you know, for any AI-touched number in a board deck, who checked it, what the model was allowed to see, and whether you could reproduce the answer if an auditor asked. Most finance organizations cannot answer that today, and most vendor content will not tell you, because gating that research behind a sales call is easier than admitting the honest answer is usually “no.”
This is a four-stage ladder for where a finance function actually sits, and what has to be true structurally to move up a rung. Ethics shows up in each stage as something you can point to, not a values statement: who is told a number came from a model, what the model was permitted to touch, who owns the resulting decision, and who is watching the model itself over time.
Stage 1: Ad hoc
This is where almost every finance team starts, whether they admit it or not. An analyst pastes a revenue table into a chatbot to draft commentary. A controller asks an AI tool to sanity-check a variance. Nobody approved the tool for this use, there is no record that AI touched the output, and the underlying numbers being pasted in have already been copied out of three systems by hand, so nobody can say with confidence they were right to begin with.
What breaks here is attribution and disclosure at the same time. If the board narrative reads better this quarter, nobody can say whether that is because the business improved or because a model chose flattering language for a number nobody re-verified. There is also a real data governance gap: sensitive customer or deal data ends up in a general-purpose consumer tool with no contract governing retention or training use. This stage is not “no AI.” It is AI with zero visibility, which is worse than no AI because leadership believes it has more control than it does.
Getting to Stage 2 does not take a policy memo. It takes picking one or two sanctioned tools, defining what may and may not be pasted into them, and requiring a named human to check every output before it moves. That is a low bar and it is still a real step up, because it converts invisible risk into visible, bounded risk.
Stage 2: Supervised
Now the function has approved tools, a data-handling rule (no customer PII into consumer-tier chat, for instance), and a working norm that a human reviews every AI-touched output before it ships. This is where most well-run finance teams sit today, and it is a legitimate, defensible place to be, not a stage to rush past for its own sake.
What breaks here is scale and consistency. Supervision works when volume is low. It buckles when a team wants AI drafting twenty variance explanations a week or checking a hundred vendor invoices, because “a human reviews everything” quietly becomes “a human skims most things,” and nobody decided that trade-off on purpose. The other failure mode is that review has no memory: the same category of error can get caught and waved through a dozen times because there is no log connecting this month’s mistake to last month’s near-identical one.
Reaching Stage 3 means defining, explicitly, which categories of AI output actually need full human review versus a lighter check, and starting to log every AI-assisted decision: not just the number, but which model, which inputs, which human signed off. Without that log you cannot tell reviewers apart from rubber stamps, and you cannot yet write rules based on evidence of where errors actually cluster.
Stage 3: Governed
This is the stage most finance organizations describe wanting and few have actually built. It has three concrete features. First, defined autonomy tiers: not every AI task gets the same scrutiny, because not every task carries the same risk. A draft commentary paragraph and a wire approval are not the same category, and treating them as if they were either over-controls the trivial or under-controls the dangerous. Second, decision logging as infrastructure, not an afterthought: every AI-touched figure carries a record of its source, the model version, and the accountable human, queryable later, not reconstructed from memory during an audit. Third, source-verification gates: the AI is not permitted to answer with a number that cannot be traced to a reconciled figure in the underlying system of record. This is the stage where the ethics questions stop being abstract and become engineering requirements: disclosure is a field in the log, not a disclaimer in an email footer, and consent and data governance are enforced by what the model’s connection is actually scoped to see, not by a policy nobody re-reads.
What breaks at this stage, if you try to reach it early, is that governance of outputs is structurally impossible if the inputs are not reconciled. You cannot build a source-verification gate on top of a spreadsheet with three conflicting versions of “revenue,” because there is no single reconciled figure to gate against: the AI has no ground truth to verify against, so the gate becomes theater. This is the load-bearing reason governance keeps stalling in finance functions that have not fixed their data layer first: the maturity ladder and the data foundation are the same problem viewed from two angles. Our pillar on AI governance and model risk goes deeper on this, and three lines of defense for AI covers who actually owns validation once the model becomes an author of numbers rather than just a drafting tool.
Earning Stage 4 means proving, over a real stretch of time, that the autonomy tiers hold under volume: the low-risk tier genuinely stays low-risk, and the logs catch the edge cases they were designed to catch, before any tier extends toward unsupervised action.
Stage 4: Autonomous-within-bounds
At this stage, agents act: initiating a reconciliation, flagging and routing an exception, drafting and filing a variance explanation, all inside hard guardrails set in advance (dollar thresholds, reversibility requirements, categories explicitly out of scope). Every action is attributable to a specific agent run and replayable, meaning you can reconstruct exactly what data it saw and why it acted, not just what it concluded. This is close to what our autonomy tiers framework and governed autonomy vs. human-in-the-loop describe in more detail: autonomy that scales past the ceiling of manual approval without giving up the audit trail approval was there to provide.
Few finance functions are genuinely here yet, and that is fine: this is not a stage to fake. The mistake we see is teams marketing themselves as Stage 4 while still operating Stage 1 or 2 controls underneath: an “autonomous” agent bolted onto ungoverned data with no real decision log is not more mature: it is Stage 1 with better branding. Real Stage 4 requires everything below it to already hold.
What breaks, and what earns the next stage
| Stage | What breaks | What earns the next stage |
|---|---|---|
| Ad hoc | No attribution, no disclosure, unverified inputs pasted freely | Sanctioned tools, a data rule, mandatory human check |
| Supervised | Review quality erodes at volume, no memory of past errors | Autonomy tiers by risk, decision logging begins |
| Governed | Gates require reconciled inputs that may not exist yet | Proven, logged performance of tiers under real volume |
| Autonomous-within-bounds | Guardrails drift if inputs or thresholds go unmonitored | Ongoing model governance: revalidation, monitoring, retirement |
That last row matters even once you have reached the top rung. Model governance does not end at deployment. Models get revalidated against new data, monitored for drift as the business changes shape, and retired when a better-fit model exists or when the questions being asked of it have outgrown what it was built to answer. Skipping this is how a Stage 4 system quietly decays back into something closer to Stage 2 with more confident branding.
The takeaway
Maturity is not a tools list and ethics is not a values statement. Maturity is whether disclosure, consent, attribution, and model governance are structural features of how AI touches your numbers, not habits you hope hold under pressure. And none of the higher stages are reachable without the one precondition vendors rarely say out loud: you cannot govern outputs built on unreconciled inputs. A source-verification gate needs a source of truth to verify against.
If you want to know honestly which stage your finance function is actually in, that conversation starts with the data layer underneath the AI, not the AI itself. Book a demo and we will show you what a reconciled foundation looks like against your own numbers.
Part of Governing AI in Finance: Model Risk, Controls, and Validation for the LLM Era