Skip to content
New: ask the Rexfin Analyst Agent about your model. Every figure comes back cited.
· 8 min read

The Citation Test for AI-Written Finance Commentary

AI board narratives sound authoritative whether or not the numbers inside them are sourced. Here's the one-click test that tells you which you're getting.

By The Rexfin team

Pick any AI-drafted variance paragraph in your board deck. Find a number inside it: a percentage, a dollar figure, a basis-point move. Click it. Does anything happen?

For most AI commentary tools on the market today, the answer is no. There’s nothing to click. The number is just text, generated the same way the sentence around it was generated: as the statistically likely next tokens, not as a value retrieved from a ledger. That gap is easy to miss, because prose doesn’t announce its own uncertainty the way a spreadsheet cell does. A red triangle in the corner of a cell says “something’s off here.” A fluent sentence says nothing of the kind, whether the number inside it is right or invented.

Why narrative is the easiest place for a wrong number to hide

A table forces scrutiny. Rows and columns invite the eye to check a total, spot a gap, notice a number that doesn’t match the one two rows up. Prose does the opposite. A well-formed sentence reads as a single unit of meaning, and readers extend trust to the whole sentence based on how well it’s written, not on how each embedded fact was derived. “Gross margin slipped 180 basis points, driven primarily by a mix shift toward lower-margin enterprise deals” sounds like an analyst’s conclusion. Whether the 180 is real or a plausible-sounding guess, the sentence carries the same tone of authority either way.

This is a known failure mode for language models specifically. A model asked to “explain the variance” from a budget file is not retrieving a number: it’s predicting the most likely sequence of words given the pattern of similar sentences in its training data. Most of the time that produces something close to right, because financial commentary follows recognizable patterns and the model has seen a lot of it. But “most of the time” is not a standard anyone would accept from a formula in a spreadsheet, and there’s no principled reason to accept it from a sentence instead just because the sentence reads smoothly.

The one buyer test that exposes the gap

You don’t need a benchmark or a vendor’s word for it. You need thirty seconds with the product you’re evaluating.

  1. Generate an AI variance or board commentary paragraph inside the tool.
  2. Pick any number embedded in the prose.
  3. Try to click it, or ask the tool to show its source.
  4. See what happens.

There are three possible outcomes, and they sort vendors cleanly:

OutcomeWhat it means
The number links to a specific ledger row, account, or filed statementThe figure was retrieved or calculated, then inserted into prose written around it
The tool shows a general “source: your ERP” note with no line-level traceProvenance exists at the system level but not the number level, you can’t prove this figure
Nothing happens; the number is just textThe figure has no traceable origin. It may be right. You have no way to know without re-deriving it by hand

The third outcome is more common than buyers expect, because commentary generation is often bolted onto a chat or analytics layer as a convenience feature, with the underlying model given a rough dataset and a prompt to “write commentary,” not a citation obligation. The tool’s own documentation may even acknowledge the gap indirectly, framing “finance edit for accuracy” as the safeguard: a human is expected to catch anything wrong before it ships. That’s not nothing, but it quietly transfers the entire verification burden back onto the person the tool was supposed to save time for, and it depends on that person noticing a plausible-looking error in a paragraph they didn’t write and don’t have time to re-derive from scratch.

Why “have finance review it” isn’t the same as being sourced

Human review is a real control, but it’s a weak substitute for structural citation, for a simple reason: review effort scales with suspicion, and fluent prose doesn’t trigger suspicion. A reviewer scanning a paragraph for wrongness is looking for something that reads off: an implausible number, a claim that contradicts what they remember. A wrong figure that’s merely a little off, sitting inside otherwise-correct prose about a variance the reviewer half-expected anyway, is exactly the kind of error that a confident sentence is built to slide past a tired reader at 11pm before a board meeting.

Compare that to how a spreadsheet forces verification: a formula either references a cell or it’s hardcoded, and anyone auditing the sheet can trace the reference in seconds. Commentary generated without embedded citations removes that structural check and replaces it with “someone should notice.” That’s a downgrade in verification standard dressed up as a writing feature, and it’s worth naming directly rather than assuming it’s obviously understood, because from the outside, a well-formatted AI commentary panel and a properly cited one look identical until you click.

What citation-per-line commentary actually requires

Fixing this isn’t a prompting trick. It requires restructuring what the model is allowed to do, not just what it’s asked to do.

The figures in a variance narrative (the percentage move, the dollar delta, the period comparison) have to be computed by a deterministic engine before any language model sees them, from a model that already ties to the ledger. The reliable modeling layer underneath rexfin.ai exists specifically so that variance math, YoY deltas, and threshold flags are arithmetic over reconciled data, not language generation. Once that evidence block exists, a model can write the connective prose (“driven primarily by,” “partially offset by”) but only from the numbers already computed, never inventing its own. A citation-truthfulness check then rejects any draft that introduces a number not present in the evidence block, the same discipline that governs Rexfin’s product-tour writeup on cited AI narratives and the pinpoint page citations that let a reader click any figure back to its source page in a filing.

This is the same split that makes automated variance commentary genuinely useful rather than genuinely risky: the model writes, the engine counts, and the two jobs never trade places. A vendor whose commentary feature can’t describe this split, who can only point to a downstream human edit as the safeguard, hasn’t built citation into the generation process. They’ve built a paragraph and hoped someone catches the mistake.

What to ask before you buy

If a vendor’s commentary feature is central to your evaluation, four questions separate structural citation from a hope-based safety net:

  • Can I click any number in a generated paragraph and see the exact source row or filed figure it came from? Not “the system,” the row.
  • What happens if the model can’t find a citable number for a claim it wants to make? The honest answer is that it refuses or falls back to a template, not that it writes the sentence anyway with the best guess it has.
  • Does the same citation check run when commentary is exported to a deck, or only in the live app? A number that’s clickable on screen and un-sourced in the PDF board pack has only solved half the problem.
  • Is a plan or budget figure inside the commentary labeled differently from an audited actual? Unaudited inputs feeding a narrative should read differently than filed ones, not blend into the same confident sentence.

If a demo can’t answer these live, on your own numbers, the “finance edit for accuracy” line in the documentation is doing more work than it should.

The takeaway

Fluent prose is not evidence of a sourced number: it’s evidence of a fluent model. The test for whether AI-written commentary is trustworthy isn’t how well it reads; it’s whether the figures inside it survive a click. Narrative generation needs the same citation discipline as a financial statement, not a lighter one, because the risk isn’t that the prose is unconvincing. It’s that it’s convincing regardless of whether it’s right.

Part of AI Hallucinations in Financial Data: Stop AI Inventing Numbers

Keep reading

Book a demo

See your numbers tie out.

Book a 30-minute demo. Bring a question you can never answer fast enough, and we will model it live against real financial data.