Skip to content
New: ask the Rexfin Analyst Agent about your model. Every figure comes back cited.
· 7 min read

IDW PS 861: How Your Auditor Will Test AI-Generated Financial Figures in 2026, and What Gets Rejected

Germany's first audit standard for AI systems demands traceability, reproducibility, and validation. Non-deterministic models without a reconciled number base fail. Here's the mechanism.

Germany's first audit standard for AI systems demands traceability, reproducibility, and validation. Non-deterministic models without a reconciled number base fail. Here's the mechanism.

By The Rexfin team

Picture the moment a partner at your audit firm asks one question: “Show me how the model produced this number, and produce it again.” If the answer is a chat transcript and a shrug, the figure is not an audit result. It is unaudited working material. That distinction is now written down.

The IDW, Germany’s institute of public auditors, published IDW PS 861 in March 2023 as the first formal audit standard for AI systems. It was built on the international assurance framework ISAE 3000 and originally scoped to AI assurance engagements outside the financial statement audit. But the standard says something the profession has since extended into the year-end audit itself: when an AI application produces output that is relevant to the financial statements, the auditor has to assess it, and the same minimum requirements apply. By 2026, that is no longer theoretical. AI is in close-to-the-ledger workflows, and auditors are arriving with a checklist.

What IDW PS 861 actually requires

Strip the standard to its load-bearing parts and three demands stand out.

Traceability (Nachvollziehbarkeit). The auditor must be able to follow how a result was produced, and where the boundary sits between the model’s logic and the data it consumed. A number that appears with no path back to its inputs is not traceable, and an untraceable number is not assessable.

Reproducibility. This is the one that breaks most generative-AI deployments. The standard’s own guidance concedes a hard limit: for a single transaction or a handful, you can verify by reproducing the result; across the volume and variation of a real ledger, you cannot manually reproduce every case. So the system itself has to be deterministic enough that the same inputs yield the same outputs. Run the question twice, get the same figure twice. A large language model does not do this by design. Ask it to total a column today and tomorrow and you may get two answers, both plausible, neither anchored.

Validation, IT security, and performance. Beyond the headline two, PS 861 sets minimum requirements covering governance, change management, the security of the surrounding IT, and demonstrated performance against defined criteria. The point is not that AI is forbidden. The point is that an AI result earns the status of “audit-relevant” only when it can be governed, secured, and shown to perform.

There is also a competency obligation worth noting: since early 2025, organizations are expected to ensure that staff using AI systems have appropriate AI literacy. The auditor will assume the people relying on the output understand its limits.

Why a chatbot over your ledger fails the test

Here is the uncomfortable part for anyone who has wired a language model directly to their accounting data. The failure is not a tuning problem. It is structural.

A general-purpose model answers a question about revenue by predicting tokens. It is, at its core, a probability machine. Two things follow. First, it is non-deterministic, which collides head-on with the reproducibility requirement. Second, even when it lands on the right figure, it usually cannot tell you which journal entries, which period cutoffs, which intercompany eliminations produced it. That collides with traceability. You end up with a number that is sometimes right, occasionally wrong, and never provably tied to the ledger. Under PS 861, that is the definition of working material, not a result.

The mistake is asking the model to be the calculator. Auditors do not object to AI reading and summarizing. They object to AI inventing arithmetic that no one can reproduce or trace.

The mechanism that survives the standard

This is where the architecture matters more than the model. The fix is to separate the two jobs the LLM is currently doing badly at once: understanding the question, and computing the answer. Let AI handle language. Make a deterministic engine handle math.

That separation is exactly how Rexfin is built. The platform connects to your accounting and financial-data sources, the integrations into QuickBooks, Xero, NetSuite, Sage, SAP, Oracle, or a data warehouse, or to uploaded statements, and builds one reconciled financial model. That model is a single source of truth that ties out to the ledger. When someone asks a question, AI interprets the request and retrieves figures, but the calculation runs through a deterministic engine, not the language model. Same inputs, same output, every time. And every figure traces back to source: the entry, the account, the period it came from.

Read against PS 861, the boxes line up. Reproducibility comes from the deterministic engine. Traceability comes from a model that ties out to the ledger and keeps the lineage. The boundary between “the AI” and “the data” is explicit, because the AI never touches the arithmetic. What-if scenarios run against the same reconciled base, so a sensitivity analysis is as auditable as the closing number. That is the difference between handing your auditor a transcript and handing them a working paper.

I won’t oversell it. PS 861 covers governance, change management, and IT security too, and no architecture audits those for you, they are organizational disciplines you still have to run. A reconciled, deterministic model does not exempt you from controls. What it does is remove the single most common reason AI output gets rejected at the door: the inability to reproduce and trace a figure.

What to do before the 2026 audit

Two practical moves. First, inventory every place AI output flows into something financially relevant, forecasts, reconciliations, board figures, disclosure inputs, and ask of each one: can I reproduce this exactly, and can I trace it to the ledger? If the honest answer is no, you have a finding waiting to happen. This connects directly to the broader compliance picture, including the GoBD requirement for machine-evaluable results and the human-oversight duty under Article 14 of the EU AI Act. PS 861 is one column in a wider table.

Second, fix the architecture, not the prompt. No amount of prompt engineering makes a probability machine deterministic. The reliable path is to give AI a reconciled model to read from and a deterministic engine to compute with.

The takeaway is blunt: in 2026 your auditor is no longer asking whether you used AI. They are asking whether your AI figures can be reproduced and traced. Build the layer that answers yes before they ask. If you want to see a deterministic, ledger-tied model in action, book a demo.

Part of AI in Finance, Audit-Proof: GoBD- and AI-Act-Defensible Models With Traceable Numbers

Keep reading

Book a demo

See your numbers tie out.

Book a 30-minute demo. Bring a question you can never answer fast enough, and we will model it live against real financial data.