Skip to content
New: ask the Rexfin Analyst Agent about your model. Every figure comes back cited.
· 7 min read

The 95% Problem: What MIT's GenAI Divide Means for Finance

MIT found 95% of GenAI pilots return nothing. In finance the culprit isn't the model, it's unreconciled data. Here's the fix.

MIT found 95% of GenAI pilots return nothing. In finance the culprit isn't the model, it's unreconciled data. Here's the fix.

By The Rexfin team

A number has been making the rounds in board decks for the better part of a year: 95% of enterprise generative-AI pilots produce no measurable return. It comes from MIT’s NANDA initiative, in a 2025 report called The GenAI Divide: State of AI in Business. The researchers looked at hundreds of deployments and found that the vast majority never moved a P&L line. A small minority did.

That stat gets quoted to justify caution, or used as a punchline. Both miss the point. The interesting question isn’t why 95% failed. It’s what the other 5% did differently. And in finance specifically, the answer is uncomfortable, because it has almost nothing to do with which model you picked.

The divide isn’t about model quality

When a finance team’s AI pilot dies, the post-mortem usually blames the wrong thing. “The model wasn’t smart enough.” “Hallucinations.” “We needed GPT-5, not GPT-4.” So the team waits for the next release, runs the same pilot on a better model, and gets the same disappointing result. The model was never the bottleneck.

MIT’s framing is that the gap sits between the tools and the workflows. Generic chat assistants get adopted easily because they’re forgiving: draft an email, summarize a doc, no harm done if it’s slightly off. But systems meant to do real operational work, where the output feeds a decision, stall. They don’t retain context, they don’t integrate with how the team actually works, and crucially, nobody can trust the output enough to act on it without re-checking everything by hand. At which point, what did the AI save you?

In finance the trust problem is sharper than almost anywhere else. A summary that’s “mostly right” is fine for a meeting recap. A gross margin that’s “mostly right” is a wrong number in a board deck, and your name is on it.

Why finance fails for a different reason

Outside finance, the failure mode is often workflow friction. Inside finance, there’s a more specific blocker underneath that one: the data the AI is reasoning over was never trustworthy to begin with.

Walk through what actually happens. A team points an AI tool at their QuickBooks or NetSuite data, or uploads a few exports, and starts asking questions. The accounting system says one thing. The CRM says another about the same customers. Last month’s board deck has a third version of the revenue figure because someone made a manual adjustment in a spreadsheet that never made it back to the ledger. The AI doesn’t know any of this. It confidently averages over the mess and hands back an answer that ties out to nothing.

This isn’t a model defect. It’s a data defect that the model faithfully inherits. FP&A teams already know this pain intimately: survey after survey finds analysts spending a large share of their week, often something like a third to nearly half, just gathering and reconciling data before any analysis happens. That manual reconciliation is the work that makes a number trustworthy. Skip it, and the AI is fast and wrong instead of slow and right.

So the real precondition for AI ROI in finance is not a better model. It’s a reconciled single source of truth: one financial model where the numbers actually tie out to the ledger, before the AI ever touches them. This is the core of why trustworthy finance AI needs a reconciled model underneath.

What the successful 5% have in common

Strip away the hype and the deployments that worked share a boring trait. They were grounded. The AI wasn’t free-associating over a pile of documents; it was retrieving from a defined, governed data source, and the people using it could see where each answer came from.

That’s the whole game. Grounding turns a plausible-sounding system into a defensible one. When a controller can click a figure and trace it back to the source transaction, the AI stops being a liability and starts being leverage. When they can’t, every output needs a manual audit, and the productivity case collapses: you’ve added a step, not removed one.

There’s a related trap worth naming. Even with good grounding, asking a language model to compute the answer reintroduces error, because LLMs are genuinely bad at arithmetic. (We get into the mechanics in why AI gets financial math wrong.) The fix is to let AI retrieve and interpret, and hand the actual calculation to a deterministic engine. Retrieval grounds the inputs. The engine guarantees the math. Neither alone is enough.

How to be in the 5%, in order

If you’re a CFO deciding whether this year’s AI budget repeats last year’s disappointment, the sequence matters more than the tooling. Roughly:

  • Reconcile first. Get to one model where revenue, margin, and cash tie out to the ledger across every connected source. No AI step is worth running on numbers you wouldn’t sign.
  • Ground the AI in that model, not in raw exports or a document dump. The AI should retrieve from the reconciled source, with every figure traceable back to where it originated.
  • Offload the math to a deterministic layer. Let the model decide what to calculate; let a real engine do the calculating, so the same question gives the same answer every time.
  • Then automate. Once you trust the numbers and the path to them, expand into reporting, variance analysis, and scenarios. Automation is the reward for trust, not a substitute for it.

Most failed pilots invert this. They automate first, discover the numbers can’t be trusted, and quietly shut it down. The 95% isn’t a verdict on AI. It’s a verdict on starting in the wrong place.

The takeaway

MIT didn’t prove AI doesn’t work in finance. It proved that AI built on unreconciled data doesn’t work, which any controller could have told you for free. The model is a commodity now. The reconciled, traceable financial layer underneath it is the thing that’s hard, and the thing that decides which side of the divide you land on.

If you’d rather your next AI initiative be in the 5%, the unglamorous first move is fixing the numbers. That’s the part Rexfin is built to do. Book a demo and we’ll show you what AI looks like when every figure ties back to source.

Part of The Reliability Layer AI Needs Before It Touches Your Numbers

Keep reading

Book a demo

See your numbers tie out.

Book a 30-minute demo. Bring a question you can never answer fast enough, and we will model it live against real financial data.