Skip to content
New: ask the Rexfin Analyst Agent about your model. Every figure comes back cited.
· 7 min read

Why GCC CFOs Don't Trust AI With Their Numbers

Most Gulf CFOs say AI delivers no real value. The problem isn't the model. It's the data underneath. Here's how to fix the trust layer.

Most Gulf CFOs say AI delivers no real value. The problem isn't the model. It's the data underneath. Here's how to fix the trust layer.

By The Rexfin team

A finance director in Riyadh runs a board pack through a shiny new AI assistant. It returns a clean revenue figure, confident and well-formatted. She checks it against the ledger. It’s off by a few hundred thousand riyals. Not catastrophically wrong, just wrong enough that if she’d trusted it, she’d have stood in front of her board defending a number that didn’t exist.

She stops using the tool that week. And she’s not alone.

The Gulf is, by most measures, ahead of the rest of the world on AI ambition. Government mandates, sovereign AI programs, real budgets behind them. And yet survey after survey lands on the same uncomfortable finding: a large majority of CFOs in the region report that AI has produced no significant value in their finance function. One widely cited figure puts it around 86 percent. Read that again. A region that leads on adoption, stalled on outcomes.

The usual explanation is that the models aren’t good enough yet. That’s the wrong diagnosis.

The model is rarely the problem

When a CFO says “I don’t trust AI with my numbers,” she almost never means the language model can’t write a coherent sentence. The models are remarkable at language. The problem shows up the moment language meets arithmetic over real, messy financial data.

Two things break.

First, the math. Large language models don’t calculate the way a spreadsheet does. They predict the next token based on patterns in text, which means a question like “what’s our gross margin for Q3” gets pattern-matched, not computed. Sometimes the pattern lands on the right answer. Sometimes it confidently doesn’t. There’s no internal ledger, no double-entry check, nothing that ties the output back to a source.

Second, the data. Even a model that calculated perfectly would still produce garbage if it’s reading from unreconciled inputs. And in most Gulf mid-market and enterprise finance teams, the inputs are exactly that: scattered across an ERP, a few spreadsheets, a bank portal, an e-invoicing system, maybe a regional subsidiary on a different chart of accounts. The numbers don’t agree with each other yet. Asking AI to reason over that is asking it to average two contradictory truths.

So the AI gives you a number that is fluent, fast, and unverifiable. For a CFO who has to personally stand behind every figure that reaches the board or the regulator, fluent-but-unverifiable is worse than no answer at all.

Why trust is the actual bottleneck

Trust isn’t a soft word here. It’s the gate everything else waits behind.

A CFO will not automate a forecast she can’t audit. She won’t let an agent reforecast cash if she can’t trace where the cash number came from. She won’t put an AI-generated figure into a regulator-facing document if she can’t show the calculation path. Adoption stalls not because the technology is incapable, but because the output isn’t defensible.

In this region, defensibility carries extra weight. Between ZATCA e-invoicing in Saudi Arabia, UAE e-invoicing rolling out, and regulators increasingly asking for explainability around automated decisions, “the AI said so” is not an answer you can give. You need to point at a source.

That’s the real divide. Not capable AI versus incapable AI. Auditable numbers versus confident guesses.

What actually fixes it: a reconciled layer underneath

The fix isn’t a smarter model. It’s putting a reliable layer between your raw financial data and whatever AI you point at it.

Concretely, three things have to be true before the AI ever speaks:

  • One reconciled model. Your accounting platforms, bank data, and any uploaded statements get connected and tied out into a single financial model that agrees with the ledger. Not a copy, a reconciliation. If the model says revenue is X, X matches the books.
  • Calculation done deterministically. When someone asks for a margin, a runway, a variance, the math runs through a real calculation engine, not the language model’s pattern-matching. Same inputs, same answer, every time. Reproducible, not improvised.
  • Every figure traces to source. Any number the AI reports can be clicked back to the document, the transaction, the line it came from. That’s what turns an answer into something a CFO can sign.

This is the order that matters: trust first, then automate. Get the numbers verifiable, then let AI retrieve them, run scenarios against them, and surface insight. The intelligence sits on top of a foundation that already ties out.

It’s the same principle we cover across the broader pillar on building a trusted, compliant numbers layer for GCC finance teams. The data layer is the unglamorous part nobody markets, and it’s the part that decides whether any of the AI on top is usable. If you want to see the mechanism in detail, our how it works page walks through it.

A practical way to test it

Before you trust any AI finance tool, ask it one thing: show me where this number came from.

If it can show you the source document, the calculation path, and confirm the figure reconciles to your ledger, you have something you can build on. If it just restates the number more confidently, you have a liability with good UX.

This is also why the regional specifics matter so much. A tool that handles ZATCA and UAE compliance with audit-ready, source-traceable models is solving the trust problem at the level regulators actually inspect. And if your team works in Arabic, the language layer adds its own traps, which is why even the strongest Arabic models still need a deterministic numbers layer underneath to do the actual math.

The takeaway

The Gulf’s AI problem in finance was never a model problem. It’s a data-trust problem wearing a model costume. The 86 percent who see no value aren’t using bad AI. They’re pointing good AI at numbers that don’t yet agree with each other, getting confident answers they can’t defend, and quietly walking away.

Fix the layer underneath and the same models suddenly become useful, because now their outputs trace to source and tie to the ledger. Trust isn’t the reward for adopting AI. It’s the precondition.

If you want to see what AI looks like when every figure reconciles and traces back to source, book a demo and bring a number you’ve never trusted from a tool.

Part of AI in Finance for the GCC: A Trusted Numbers Layer

Keep reading

Book a demo

See your numbers tie out.

Book a 30-minute demo. Bring a question you can never answer fast enough, and we will model it live against real financial data.