AI in bookkeeping: keeping GoBD traceability despite the black box
An AI tool that books or summarizes your numbers but can't show its work fails GoBD. Here's how to keep the audit trail intact.
By The Rexfin team
A Steuerberater in Munich told us about a client review that still bothers him. The bookkeeping looked clean. The numbers tied out. Then he asked a junior how a particular batch of receipts had been categorized, and the answer was: “We pasted them into ChatGPT and it suggested the accounts.” Nobody had logged the prompts. Nobody could reproduce the suggestions. The booking existed, but the reasoning behind it had evaporated the moment the chat window closed.
That is the GoBD problem with AI in one scene. Not that the AI was wrong. That nobody could prove it was right.
What GoBD actually demands, in plain terms
The GoBD are the German tax authority’s principles for properly keeping and storing books and records in electronic form. They are not new, and they are not vague about the parts that matter here. Three requirements collide directly with how most generative AI works.
Nachvollziehbarkeit: traceability. A knowledgeable third party must be able to follow how a business transaction became a booking, within a reasonable time, without your help. The entry has to lead back to the source document, and the source document has to lead forward to the entry. A progressive and a retrograde check both have to work.
Nachpruefbarkeit: verifiability. The records have to be testable. An auditor should be able to take a number and check it against the underlying evidence and the rules that produced it.
Unveraenderbarkeit: immutability. Once a booking is recorded, you cannot silently overwrite it. Changes must be logged so the original state and the corrected state both survive.
There is a fourth piece people forget: the Verfahrensdokumentation. You are expected to document the process by which data enters your books and gets transformed along the way. If a tool touches the data, the tool is part of the process you have to describe.
Read those four together and the question answers itself. Does a free-text chat assistant that “suggests” account assignments leave a reproducible trail? No. Does it produce a process you could write down and a third party could re-run? Not unless you built that scaffolding yourself.
Why a standard LLM breaks all four at once
A large language model generates the most probable next token given the context. It does not retrieve a fixed figure and it does not run arithmetic in a way you can inspect. Ask the same question twice and you can get two different answers, both stated with the same calm confidence. That is fine for drafting an email. It is a problem when the output becomes part of a tax-relevant record.
Run the test yourself. Paste a profit-and-loss extract into a general assistant and ask for the gross margin. You will usually get a number. Ask again in a fresh session and you may get a slightly different one, or a different rounding, or a figure that quietly assumed VAT one way the first time and the other way the second. The model is not retrieving your margin. It is pattern-matching toward something that looks like a margin.
Now layer GoBD on top:
- Traceability fails because the answer can’t be tied to a specific source line.
- Verifiability fails because re-running the prompt may not reproduce the figure.
- Immutability fails because there is no record of what the model saw or said.
- The Verfahrensdokumentation can’t be written because the process isn’t deterministic enough to describe.
None of this means AI is unusable in German bookkeeping. It means the AI cannot be the system of record, and it cannot be the thing that does the math.
The real risk: Schatten-KI
The dangerous version of this isn’t a sanctioned tool. It’s Schatten-KI: shadow AI. Staff quietly pasting ledger extracts, invoices, and customer data into whatever consumer chatbot is open in another tab. It feels harmless because it’s just “asking for help.” But every paste is an undocumented step in your bookkeeping process, and some of those pastes are also a data-protection incident waiting to be found.
Surveys across German finance teams keep landing on the same picture: a large share of employees already use AI tools at work, and a meaningful chunk do it without IT or management knowing. You don’t get to opt out of that reality by writing a policy. People will reach for the tool that saves them twenty minutes. The job is to give them a sanctioned path that happens to be the GoBD-safe one.
The exposure is concrete. If a Betriebspruefer concludes the bookkeeping isn’t nachvollziehbar, the worst case isn’t a stern letter. It’s a Verwerfung: the records get rejected, and the tax base can be estimated instead. An estimate is rarely in your favor.
How to keep the audit trail when AI is in the loop
You don’t fix this by banning AI or by trusting it more. You fix it by changing what the AI is allowed to be responsible for. Three principles, in order.
1. The numbers live in a reconciled model, not in the chat
Before any AI touches a figure, the underlying data has to be assembled into one reconciled financial model that ties back to the ledger. Accounting data, bank data, sub-ledgers, uploaded statements: pulled together, matched, and tied out to source. This is the single source of truth, and it exists independently of any question anyone asks. The AI does not “know” your revenue. It looks it up in a model that already agrees with your books.
This is the heart of what we build at Rexfin: one reconciled model that every answer traces back to. If you want the mechanics of why reconciliation has to come before the AI, the pillar on GoBD- and AI-Act-safe finance models lays out the full argument.
2. The math runs in a deterministic engine, not in the model
The LLM is good at understanding a question in German and at explaining an answer in plain language. It is bad at arithmetic, and worse at arithmetic you can audit. So the calculation has to happen somewhere deterministic. Ask for a gross margin and the same inputs produce the same output every time, by the same documented formula. The AI routes the question. The engine does the sum. That split is what makes the result reproducible, which is exactly what Nachpruefbarkeit requires. There’s a deeper version of this point in the sibling piece on why data quality beats the model.
3. Every figure carries its provenance and gets logged
When the AI reports a number, you should be able to click it and see: which source documents fed it, which calculation produced it, and when. And the interaction itself should be written to an immutable log. That covers traceability and immutability in one move. A figure that can name its source and survive in a tamper-evident record is a figure you can defend in a Pruefung.
Put those three together and the Verfahrensdokumentation almost writes itself, because now there is a stable, describable process: data in, reconciled, calculated deterministically, answered with a citation, logged. You can hand that description to a third party and they can follow it.
A short reality check
This approach has limits worth stating plainly. It does not turn AI into the bookkeeper. A human still owns the bookings and the judgment calls. It does not absolve you of the data-protection obligations that come with feeding the model your records. And no architecture removes the duty to actually maintain the Verfahrensdokumentation: the tooling makes it possible, you still have to keep it current.
What it does buy you is the thing GoBD is really after. Not a smarter machine. A defensible one.
The test for any AI you let near your books is simple, and you can apply it in a demo: pick a number it produced and ask it to show its work. If it can name the source document, reproduce the calculation, and prove it hasn’t been altered, it belongs in a German finance function. If it can only sound confident, it doesn’t.
If you want to see what that looks like with your own ledger, book a demo and bring a figure you’d normally have to defend in front of an auditor.
Part of AI in Finance, Audit-Proof: GoBD- and AI-Act-Defensible Models With Traceable Numbers