Skip to content
New: ask the Rexfin Analyst Agent about your model. Every figure comes back cited.
· 7 min read

Scenario Analysis and Rolling Forecasts With AI: What-If Simulations Your CFO Will Actually Trust

AI can run thousands of what-if scenarios in seconds. But a forecast is only useful if it's reproducible and tied to a reconciled baseline. Here's how to get both.

AI can run thousands of what-if scenarios in seconds. But a forecast is only useful if it's reproducible and tied to a reconciled baseline. Here's how to get both.

By The Rexfin team

Run the question through a language model twice and you can get two different answers. That is fine when you’re asking it to summarize an email. It is a disqualifying flaw when you’re asking it what your cash position looks like in March if a key customer pays 45 days late and steel prices rise 12%.

This is the quiet tension underneath every AI-powered forecasting pitch. The demo is genuinely impressive: type a question in plain German, watch a thousand what-if scenarios compute in seconds, get a chart. The problem shows up the second time you run it, or the moment your CFO asks, “show me how you got that number.” A forecast that can’t reproduce itself isn’t a forecast. It’s a one-time hallucination with good production values.

The good news is that the speed is real and worth having. The fix isn’t to give up on AI scenario modeling. It’s to be precise about which part of the system does the thinking and which part does the math.

What AI is genuinely good at here

Scenario analysis used to be gated by labor. Building a single new what-if case in a spreadsheet (change three assumptions, re-link the formulas, check nothing broke) could eat an afternoon. So teams modeled three scenarios: base, upside, downside. Reality has more than three.

AI collapses that cost. Describe the driver you want to flex (DSO, churn, a raw-material input, headcount timing) and the system can generate dozens of permutations and a rolling forecast that updates as actuals land. For liquidity planning in particular, where the difference between “fine” and “covenant breach” can be a few weeks of timing, having many scenarios instead of three is a real upgrade. That part of the promise holds up.

The trouble is never the speed of the calculation. It’s whether you can stand behind the result.

Two requirements a CFO won’t waive

A forecast a finance leader will actually put in front of a board or a bank has to clear two bars. Most AI tools clear neither by default.

It has to be reproducible. Same inputs, same answer, every time. This is exactly where a raw LLM fails, because it generates plausible text rather than computing a fixed result. Two runs, two numbers, and now you’re explaining a discrepancy you can’t account for. Reproducibility isn’t a luxury feature in finance: it’s the price of entry, and in Germany it’s tangled up with GoBD Nachvollziehbarkeit, the requirement that any figure be traceable by an expert third party.

It has to be tied to a baseline that’s actually true. A scenario is just a transformation of a starting point. If the starting point (your current revenue, your real cash position, your committed costs) doesn’t agree with the ledger, every scenario built on it is wrong in the same direction. Fast, confident, and wrong. We make the full version of this argument in Data Quality Beats the Model, but the one-line version: garbage baseline, garbage scenarios, no matter how clever the model.

The fix: deterministic scenarios on a reconciled model

Here’s the architecture that makes AI forecasting defensible. It rests on a clean division of labor.

The AI interprets the request and orchestrates. You ask, in plain language, “what happens to runway if we delay the Q3 hires by two months and our largest customer slips to net-60?” The model figures out which drivers you mean and which scenario to assemble. That’s language work, and language models are good at it.

The calculation runs through a deterministic engine, not the model. The actual arithmetic (applying the delayed timing, re-rolling the forecast, recomputing the cash curve) happens in a calculation engine that produces the same output for the same input, every single time. The LLM never does the math. It hands the parameters to the engine and reports back what the engine returned. That single design choice is what converts “different answer each run” into “reproducible result, on demand.”

And it all sits on a reconciled baseline: one financial model built from your connected accounting and financial data that ties out to the ledger. The base case isn’t a number the AI remembered. It’s your real position, reconciled to source.

The payoff is that every scenario figure traces back to two things: the verified baseline it started from, and the documented calculation path that produced it. Run it again next week with the same assumptions and you get the same answer. Hand it to an auditor and they can follow the chain. That’s the difference between a number you defend and a number you hope nobody questions. It’s the same principle behind keeping AI accounting work GoBD-traceable, covered in AI in Accounting: Keeping GoBD Traceability Despite Black-Box AI.

Where the honest limits are

This approach doesn’t make forecasts correct. It makes them reproducible and traceable, which is a different, more achievable claim. Your assumptions can still be wrong. If you assume 5% churn and it’s 9%, a deterministic engine will faithfully compute the wrong answer. What you get is the ability to see exactly which assumption drove the result, change it, and re-run, instead of arguing about whether the tool computed the math right.

There’s also a discipline cost. A reconciled baseline means the upstream connection to your accounting data has to be maintained, not set up once and forgotten. That’s real work. But it’s work you were already doing badly in spreadsheets; here it’s done once and reused across every scenario.

The takeaway

AI scenario analysis is worth adopting: the ability to model dozens of what-if cases and roll your forecast forward automatically is a genuine gain over the base/upside/downside ritual. But speed without reproducibility is a trap, and reproducibility comes from architecture, not from a better prompt. Put the math in a deterministic engine, anchor everything to a reconciled baseline, and the AI on top becomes something a CFO will sign off on.

This article is part of our pillar on AI in finance for German teams: GoBD- and AI-Act-ready models. If you want to see deterministic scenarios run against your own reconciled numbers, book a demo.

Part of AI in Finance, Audit-Proof: GoBD- and AI-Act-Defensible Models With Traceable Numbers

Keep reading

Book a demo

See your numbers tie out.

Book a 30-minute demo. Bring a question you can never answer fast enough, and we will model it live against real financial data.