Skip to content
New: ask the Rexfin Analyst Agent about your model. Every figure comes back cited.
· 8 min read

Which AI Models Can You Trust With Financial Data in 2026? Reading the Leaderboards Critically

Hallucination rates now span under 2% to over 25% depending on the model and the task. Here's why that headline number won't tell you whether your AI is safe on your financials.

Hallucination rates now span under 2% to over 25% depending on the model and the task. Here's why that headline number won't tell you whether your AI is safe on your financials.

By The Rexfin team

A controller forwarded us a vendor pitch last quarter. The headline: “Our model hallucinates only 1.8% of the time.” The implied promise: pick the right model and the numbers problem goes away. It does not. That 1.8% came from a document-summarization benchmark, measured on clean prose, scored by an automated judge. It says almost nothing about whether the same model will get your Q3 gross margin right when the source is a consolidated trial balance pulled from two ledgers that don’t agree.

Hallucination rates in 2026 are real, public, and useful. They are also routinely misread. If you are choosing an AI tool to put figures in front of your board, you need to know what the leaderboards measure, what they don’t, and why the gap between the two is exactly where finance errors live.

What the numbers actually say

Take Vectara’s Hallucination Evaluation Leaderboard, the most-cited public benchmark. As of its most recent refresh, the top models cluster impressively tight. A handful sit under 2%. Several mainstream models from OpenAI, Google, and Anthropic land in the low single digits on the original test. The spread across all evaluated models, from best to worst, runs from under 2% to well above 20%.

Here is the part the pitch decks skip. Vectara measures one narrow thing: how often a model introduces unsupported claims when summarizing a short document it was handed. That is a faithfulness test on text. It is not an arithmetic test, not a multi-document reconciliation test, and not a test of whether the model can read a financial table without transposing a column.

When Vectara rebuilt the benchmark in late 2025 with harder inputs, including longer documents up to roughly 32,000 tokens and a wider spread of domains that included finance, the same reasoning models that scored in the low single digits jumped above 10%. Same models. Harder, more realistic task. The error rate multiplied. That single fact should reset how you read any leaderboard: the score is a property of the test, not just the model.

Why a leaderboard score doesn’t transfer to your books

Three things separate a benchmark document from your actual financials, and each one degrades model performance in ways the headline number never captures.

First, your data is messy and multi-source. A summarization benchmark gives the model one clean document. Your real question, “what was adjusted EBITDA last quarter,” touches a general ledger, a few manual adjustments living in a spreadsheet, and maybe a subledger that posts on a lag. The model has to find the right figures across all of them before it does anything. Retrieval errors here look identical to reasoning errors in the output, and no document-faithfulness score predicts them.

Second, your data has tables and PDFs. Model accuracy on financial figures drops sharply when the source is a scanned statement or an image rather than clean text. A model that summarizes an article faithfully can still misread a cell in a balance sheet because the column header sat two rows up and the model guessed the alignment.

Third, and most important, finance is arithmetic, and arithmetic is not what these benchmarks score. A model can retrieve revenue and cost of goods sold perfectly, then compute gross margin as a difference instead of a ratio, or invert it, or apply last year’s figure. None of that registers as a “hallucination” on a faithfulness leaderboard. The claim is grounded in the source. The math is just wrong. We’ve written before about this hidden class of errors in AI got the numbers right and the math wrong, and it is precisely the category leaderboards are blind to.

The model-agnostic case

So which model should you trust with financial data? The honest answer for a CFO is: the question is shaped wrong.

If your accuracy depends on which model you picked this quarter, you have built a fragile system. Models change constantly. The leaderboard leader in February is rarely the leader by August. Providers ship new versions, deprecate old ones, and silently re-tune behavior. Pinning your reporting integrity to a specific model means re-validating your entire finance workflow every time a version number ticks up, and most teams don’t, which means they’re running on a model they never tested.

There is a stronger position. Make the model the part that can change, and put the things that must not change underneath it where no model can touch them.

That means a layer that does the reconciliation before any AI sees the data, so every figure ties back to the ledger. It means routing every calculation, margin, growth rate, variance, ratio, through a deterministic engine rather than letting the LLM do the arithmetic. The model decides what to compute; the engine computes it the same way every time. And it means every answer carries a trace back to its source, so a controller can prove a number the way they’d prove a formula, not take it on faith.

When that layer exists, the model’s hallucination rate stops being a load-bearing number. A weaker model produces clumsier prose; it does not produce a wrong gross margin, because it never did the math. You can swap GPT for Gemini for Claude on a Tuesday and your reported figures don’t move. That is what model-agnostic actually buys you: resilience to the one variable you can’t control.

How to read the next pitch deck

When a vendor leads with a hallucination rate, ask three questions. What task was it measured on, summarization or arithmetic? On what data, clean text or messy multi-source tables? And what happens to your numbers when the underlying model changes next quarter? If the honest answer to the last one is “we’d have to re-test everything,” the accuracy lives in the model, and the model is sand.

The leaderboards are worth reading. They are a genuine signal that models are improving and that some are clearly better than others at staying grounded. Just don’t mistake a faithfulness score on someone else’s documents for a guarantee on your books. The trustworthy figure isn’t the one from the best model. It’s the one that was reconciled before the model saw it and calculated by something deterministic after.

If you want to see what model-agnostic accuracy looks like on your own financials rather than a benchmark’s, book a demo and bring a number you’ve never trusted an AI to get right. For the broader argument that even strong benchmark scores are a failing grade for unsupervised finance, the verification checklist for board-deck numbers is a useful next read, as is the pillar on AI hallucinations in financial data.

Part of AI Hallucinations in Financial Data: Stop AI Inventing Numbers

Keep reading

Book a demo

See your numbers tie out.

Book a 30-minute demo. Bring a question you can never answer fast enough, and we will model it live against real financial data.