Skip to content
New: ask the Rexfin Analyst Agent about your model. Every figure comes back cited.
· 8 min read

Explainability vs. Verifiability: The Trust Word Every AI Finance Vendor Uses (And What It's Missing)

Every FP&A vendor promises a 'glass box.' Explainability shows how the AI guessed; verifiability proves it against the ledger, or refuses to answer.

By The Rexfin team

Open the trust page of nearly any FP&A vendor right now and you will read some version of the same sentence: our AI is fully observable, nothing is a black box, you can right-click any cell and see the logic. There are CTO video series with titles like “Escaping the Black Box.” There are glass-box diagrams with little magnifying icons over every output. The entire category converged on the same pitch within about eighteen months of each other.

None of them will tell you the one thing that actually matters: what happens when the AI cannot prove a number. Does it say so? Or does it show you a very convincing explanation of a figure it made up?

That gap is the difference between explainability and verifiability, and right now almost the whole market is selling you the first one while implying the second.

Explainability shows you a reason. Verifiability shows you a source.

Explainability means the system can narrate its own reasoning. Ask it why Q3 marketing spend looks the way it does, and it walks you through the steps: it pulled the GL account, applied a filter, summed a range, adjusted for an accrual. That narration is genuinely useful. It is also, on its own, no guarantee the number is right.

Here is the part the glass-box demos gloss over: a large language model can produce a fluent, structured, entirely plausible explanation for a number it hallucinated. Explainability is a property of the narrative. It tells you the shape of the reasoning the model claims to have followed. It does not tell you whether that reasoning ever touched the actual ledger. You can right-click a cell and see a story. The story can be wrong and still read as confident, well-organized, internally consistent, because generating a plausible-sounding derivation is exactly the kind of task language models are good at, whether or not the underlying number is real.

Verifiability is a different claim entirely: the system cannot produce an answer unless it can prove that answer against source data. Not narrate a plausible path to it: prove it, the way an auditor proves it, by tracing the figure back to a specific transaction, statement line, or reconciled entry. A verifiable system doesn’t show you its reasoning. It shows you the receipt. And critically, if no receipt exists, a verifiable system says so instead of reasoning its way to a number anyway.

The two are not opposites, exactly. A good system should have both. But they solve different failure modes, and vendors are almost universally marketing the one that’s easier to build.

ExplainabilityVerifiability
What it showsThe reasoning path the model claims it tookThe source record the number ties to
What it provesThat the output sounds derivedThat the output is derived
Fails howConfidently narrates a wrong numberCannot fail silently: no source, no answer
Easy to fake?Yes, fluent narration is cheap for an LLMNo, either the citation resolves or it doesn’t
What you’re told when the AI is unsureUsually nothing; it explains anywayIt refuses

Why explainability won and became the whole vocabulary

It’s not hard to see why “glass box” ate the category. Explainability is achievable with the model you already have. You prompt it to show its work, format the chain-of-thought as a click-through UI, and you have a demo-ready feature in a sprint or two. It photographs well. It answers the objection buyers actually voice out loud (“I don’t trust a black box”) without requiring any change to how the number was produced underneath.

Verifiability is a harder engineering problem, because it requires the number to have been computed somewhere deterministic and traceable in the first place, before the AI ever touches it. You can’t retrofit a citation onto a number the model calculated in its own head. If the arithmetic happened inside the LLM, which is often exactly what’s happening beneath a fluent explanation, there is no ledger entry to point to, only a narrative that sounds like one. Our pillar piece on AI hallucinations in financial data goes deeper on why models get the arithmetic wrong even when the retrieval is right, which is precisely the failure a good explanation can paper over.

So the market did the achievable thing at scale, gave it a warm, transparent-sounding vocabulary, and the vocabulary became the trust signal, regardless of whether the underlying pipeline can actually refuse.

The two-question demo test

You don’t need to audit anyone’s architecture to tell these apart. You need two questions, asked in order, in any vendor demo.

Question 1: “Show me exactly where this number comes from.”

Almost every vendor will pass this one now. You’ll get a highlighted cell, a hover card, a click-through to “reasoning steps.” This tests explainability, and by 2026 it’s table stakes: it no longer discriminates between vendors.

Question 2: “Now make it refuse to answer. Ask it something it cannot verify, and show me it says so instead of guessing.”

This is the question that actually separates the field. Ask for a number the system has no clean source for: a metric that spans two unreconciled systems, a figure that requires an assumption nobody entered, a period with a data gap. A system built on explainability alone will usually still answer. It will produce a number, wrap it in a fluent explanation, and the explanation will look exactly as polished as the ones for numbers that were right. That’s the trap: you cannot tell the difference from the outside, because the narration quality doesn’t degrade when the underlying number is fabricated.

A verifiable system does something different: it stops. It tells you it can’t produce that figure and tells you why: which source is missing, which reconciliation didn’t close, which assumption was never confirmed. Refusal, on command, in front of you, is the tell. If a vendor can’t make their own product refuse, they’ve built an explainer, not a verifier.

What refusal actually requires

For a system to refuse credibly, refusal has to be a structural property, not a courtesy prompt someone added. It means every number the AI is allowed to state has to resolve to a citation against a reconciled source before the AI is allowed to state it, not after, as a decoration on an answer it already generated. If the citation can’t be resolved, the answer doesn’t get produced. That’s the cite-or-refuse mechanism: the model reasons and retrieves, but a deterministic layer underneath checks that every figure ties to source before it’s allowed to leave the building. No source, no output, regardless of how confident the model’s internal reasoning felt.

It also means the refusal has to survive the moment that matters most: the export. A demo that refuses politely in a chat window is not the same guarantee as a system that will not let an unverified number reach a board pack. That’s what an export gate is for: a hard technical block on producing a document, not a warning label a user can click past. And it’s worth being specific about why a confidence score isn’t a substitute for either of these: a percentage still lets a low-confidence number through with a caveat attached, which is a softer version of the same failure explainability has. Our piece on why confidence scores aren’t enough covers where that approach breaks down on real financial tables.

The takeaway

Explainability answers “how did you get this.” Verifiability answers “can you prove this is real,” and it’s the only one of the two that’s willing to say no. Every FP&A vendor can now show you a reasoning trail: that arms race is over, and it never actually protected you from a hallucinated figure reaching a board deck with a confident explanation attached. The test that still separates the field is whether the vendor’s AI can be made to refuse, on demand, when the number isn’t there to prove. Ask for that in the demo. If they can’t produce a refusal, you’ve learned what their glass box is actually made of.

Rexfin is built around the refusal, not just the reasoning trail. Book a demo and ask it to refuse.

Part of AI Hallucinations in Financial Data: Stop AI Inventing Numbers

Keep reading

Book a demo

See your numbers tie out.

Book a 30-minute demo. Bring a question you can never answer fast enough, and we will model it live against real financial data.