Skip to content
New: ask the Rexfin Analyst Agent about your model. Every figure comes back cited.
· 8 min read

The FP&A AI Buyer's Checklist: 12 Questions to Avoid a Black-Box Forecast

A skeptic's procurement checklist for AI FP&A tools: traceability, deterministic math, reconciliation, scenario reproducibility, and audit trails that survive a board question.

A skeptic's procurement checklist for AI FP&A tools: traceability, deterministic math, reconciliation, scenario reproducibility, and audit trails that survive a board question.

By The Rexfin team

A demo is the worst place to evaluate an FP&A tool. The data is clean, the scenarios are rehearsed, and the number on screen always ties out because someone made sure it would. The real test happens five months later, in a board meeting, when a director points at the revised forecast and asks: “Why is Q3 down 11 percent from what you showed us in March, and where did that gross-margin assumption come from?” If your tool can answer that in front of the room, you bought the right one. If it can’t, you bought a confident narrator.

Most AI FP&A pitches are built to win the demo, not the board question. So the procurement job is to interrogate the second scenario while you still have leverage. Below are twelve questions, grouped by the four failure modes that actually sink these tools: untraceable numbers, math the model makes up, foundations that don’t reconcile, and scenarios you can’t reproduce. Ask them in the order that exposes the weakest link first.

Traceability: can every number walk back to a source?

A forecast figure is only as good as your ability to reconstruct it. If the tool produces a number and the honest answer to “where did this come from” is a vector search and a language model’s best guess, you have bought a black box with good manners.

1. Click any figure in the forecast. Can it show the exact source rows, the formula path, and the inputs that produced it? Not a citation to a document. The actual lineage: ledger lines, the driver, the calculation. “It cites the source” is a lower bar than people think. A citation says where a number was mentioned; lineage says how it was derived.

2. When the AI gives a narrative (“revenue grew because of seasonality”), is the claim linked to figures that exist in the model, or generated from the prompt? This is the difference between commentary grounded in data and commentary that sounds grounded. Ask to see the same explanation regenerated twice. If the supporting numbers drift, the narrative was never tied to anything.

3. Can a non-modeler (your auditor, a board member, a new analyst) follow the trace without you in the room? Traceability that only the builder can read is not an audit trail. It’s a diary.

Deterministic calculation: who actually does the arithmetic?

This is the question that separates serious tools from impressive ones. Language models are pattern engines, not calculators. They are fluent at producing numbers that look right and are wrong, and the failure is invisible because the output is grammatically perfect. Research on AI-generated financial figures has documented exactly this: a quarterly figure read as an annual one inflates a metric by 400 percent, and the model states it with total confidence. The fix is not a better model. It’s removing arithmetic from the model’s job entirely.

4. When a number is calculated, does an LLM compute it, or does a deterministic engine compute it and the LLM only describe the result? There is no acceptable middle answer here. If the model is doing math, every figure carries hallucination risk that no confidence score can price.

5. Run the same query twice with identical inputs. Do you get a bit-for-bit identical number? A deterministic engine returns the same answer every time. A model sampling its way to an answer does not. This is a thirty-second test and it is brutally diagnostic. Make them run it live.

6. How does the tool handle the calculations finance actually argues about: adjusted EBITDA, ARR bridges, multi-entity eliminations, currency translation? Generic arithmetic is easy. The defensible-in-front-of-an-auditor calculations are where black boxes fall apart, because those definitions are opinionated and have to be applied the same way every period.

The distinction matters enough that we wrote a full piece on why driver-based forecasting should keep the math deterministic while the AI handles the levers and the language.

Source reconciliation: does the foundation tie to the ledger?

A forecast built on a foundation that doesn’t reconcile is a guess dressed as a model. Before you evaluate any forecasting cleverness, check what the cleverness sits on.

7. After connecting your accounting system or uploading statements, does the tool produce one reconciled model that ties out to the general ledger, and show you the tie-out? “It connects to QuickBooks” means it can read data. It says nothing about whether the resulting numbers agree with your books. Ask for the reconciliation report, not the connector logo.

8. When the same metric exists in two systems (say revenue in the CRM and revenue in the ERP), which one wins, and is that rule explicit or improvised per query? If the tool silently picks, you have a different number depending on what you asked and when. A single source of truth is a deliberate decision the tool should enforce, not an accident of retrieval order.

9. What happens when the underlying ledger is restated or a prior period closes differently? The forecast has to know it’s now sitting on changed ground. Tools that snapshot once and forget are quietly stale by the second close.

Scenario reproducibility and audit trail: can you defend it in six months?

The board question at the top of this article is fundamentally about time. Numbers change, and you have to explain why. That requires the tool to remember.

10. Pull up a forecast version from last quarter. Can you see exactly what inputs, drivers, and assumptions produced it, and diff it against the current version? Without versioned, reproducible scenarios, “why did this change” has no honest answer. You’re reconstructing from memory, which is how credibility erodes.

11. Is there an immutable log of who changed which assumption, when, and what the figure was before and after? This is table stakes for anything that touches the numbers you report. It’s also increasingly a compliance requirement, not a nice-to-have. A tool without a real audit trail will eventually cost you in front of an auditor or a regulator.

12. Can you re-run a past scenario today and reproduce the exact number you showed the board then? Not approximately. Exactly. If you can’t replay it, you can’t defend it, and a forecast you can’t defend is a liability with a dashboard.

How to actually run the evaluation

Don’t accept the vendor’s data. Bring a messy extract of your own: a real month, with the inter-company quirks and the one account everybody misclassifies. Ask for the deterministic re-run test (question 5) and the click-to-source test (question 1) live, on your data, in the meeting. Most tools that fail do so quietly here, and the failure is obvious the moment you ask them to do the same thing twice.

A useful mental model: the AI should be the analyst who explains the numbers, never the accountant who computes them. The moment those roles blur, every number on the screen inherits the model’s uncertainty. The whole point of a defensible forecast is that the figures come from a deterministic engine tied to your ledger, and the AI’s job is retrieval, scenario framing, and plain-language explanation that traces back to source.

If you want to see what passing all twelve looks like in practice (reconciled foundation, deterministic math, replayable scenarios, full lineage), that’s exactly what Rexfin is built to demonstrate. Book a demo and bring your worst dataset. The board question is the only test that matters; everything else is rehearsal.

For the broader picture of how this fits a forecasting process you can stand behind, start with the pillar on AI FP&A automation you can defend in the board room, or read how teams move to continuous forecasting without losing control.

Part of AI FP&A Automation: Forecasting You Can Defend in the Board Room

Keep reading

Book a demo

See your numbers tie out.

Book a 30-minute demo. Bring a question you can never answer fast enough, and we will model it live against real financial data.