Skip to content
New: ask the Rexfin Analyst Agent about your model. Every figure comes back cited.
· 7 min read

You Can't Validate the Weights: Output Validation for Vendor LLMs

You don't get the weights of a vendor LLM. Here's how output validation against a reconciled ground-truth benchmark satisfies examiners and makes quarterly testing feasible.

You don't get the weights of a vendor LLM. Here's how output validation against a reconciled ground-truth benchmark satisfies examiners and makes quarterly testing feasible.

By The Rexfin team

Ask your LLM vendor for the model weights and you’ll get a polite no. Ask for the training data, the architecture details, the fine-tuning corpus, and you’ll get the same answer wrapped in an NDA. The thing your model risk team is supposed to validate is, by contract, a sealed box. And yet the model produces numbers that end up in a board deck, a credit memo, a regulatory filing.

This is the gap that makes validation officers nervous, and it’s the right thing to be nervous about. The traditional playbook (open the model, check the math, confirm the assumptions are conceptually sound) assumes you can see inside. With a frontier vendor LLM, you can’t. So the question stops being “is the model correct?” and becomes “can I prove the outputs are correct, repeatedly, without ever looking inside?”

Why weight inspection was never the only path

SR 11-7 and the frameworks that grew around it were always built on two legs, not one. Conceptual soundness (does the model’s design make sense) is the leg everyone quotes. The second leg is outcomes analysis: does the model’s output match what actually happened, or what a trusted reference says should happen? For decades, banks validated vendor scorecards and pricing models they couldn’t fully open by leaning hard on that second leg. They benchmarked the output against known results.

Vendor models have always lived under this rule. Supervisory guidance has consistently held that a model bought from a third party gets validated under the same principles as one built in-house: the process bends to the constraints, but the bar doesn’t drop. Proprietary internals don’t earn you a pass. They earn you a heavier burden on outcomes testing.

LLMs just push that logic to its limit. There is no scorecard to peek at, no regression coefficients to sanity-check. You have a prompt going in and text coming out. That is exactly the situation outcomes analysis was built for. The shift now visible across model risk practice, and increasingly what examiners will accept for a Tier-1 vendor LLM, is that output validation against a labeled benchmark set can stand in for inspecting internals you were never going to get.

What output validation actually requires

Output validation is simple to state and brutal to operationalize. You need a set of inputs, the known-correct answer for each, and a way to score the model’s response against that answer at scale. Three pieces, and the second one is where most programs collapse.

The benchmark set has to be labeled. For a finance LLM, “labeled” means every question in the set has a ground-truth numeric answer that you trust enough to grade against. What was Q3 gross margin? What’s the current ratio after the December close? If the December numbers grow our 13-week cash runway by 6%, what does that do to covenant headroom? Each of those needs a single right answer, and that answer has to be defensible: not “the analyst’s best guess,” but a figure that ties to the ledger.

That’s the catch nobody wants to say out loud. You cannot output-validate an LLM against numbers you can’t yourself defend. A benchmark built on a spreadsheet that doesn’t reconcile is a benchmark that launders errors into your validation evidence. Garbage truth in, false confidence out.

The benchmark is a reconciled model, not a quiz

Here’s the reframe that makes this tractable. The hardest part of validating a finance LLM isn’t the LLM. It’s producing a ground truth you can stand behind.

A reconciled financial model (one single source of truth that ties out to the general ledger) is precisely that ground truth. If your numbers come from a model that reconciles to the penny against QuickBooks, Xero, NetSuite, Sage, or your warehouse, then every figure in it is a pre-labeled benchmark answer. Gross margin, runway, segment contribution, the covenant calc: all of it has a known-correct value because the model that holds it is reconciled, traceable, and auditable.

This is the architecture Rexfin is built around, and it’s worth being explicit about how the pieces line up against a validation program:

  • The labeled set comes for free. Every metric in a reconciled model is a question with a defensible answer. Your benchmark isn’t a side project; it’s a query against your source of truth.
  • The math doesn’t run through the LLM. Calculations execute in a deterministic engine. The model retrieves figures and frames answers; it doesn’t do the arithmetic. That collapses an entire class of failure (the LLM confidently inventing a number) into something you can test for and catch.
  • Every output traces to source. When the model returns “runway is 11 months,” you can follow that figure back through the calculation to the reconciled balances it rests on. Traceability is what turns a validation finding into a closed item instead of an argument.

When the benchmark is a reconciled model rather than a static answer key, quarterly output validation stops being a special project. You re-run the labeled set against the current model after each close, score the LLM’s responses, and log the deltas. Same harness, fresh truth, every quarter.

What this doesn’t fix

Output validation is necessary, not sufficient, and a credible program says so out loud.

A benchmark only covers the questions in it. Real users ask things you never labeled, in phrasings you never anticipated, and a model that aces your set can still mislead on the long tail. Coverage is a permanent open question: you expand the set, you sample production traffic, you accept that some risk lives outside the benchmark.

Output validation also tells you nothing about how a wrong answer happened. A vendor can silently swap in a smaller or quantized variant of the model to cut costs, and your scores can drift before you understand why. That’s why output validation pairs with ongoing monitoring rather than replacing it. You validate quarterly and you watch continuously, because the box you can’t open can also change without telling you.

And none of this addresses retrieval. If the model pulls the wrong figure from the right source, or the right figure from a stale one, that’s a data-layer failure your output benchmark may or may not catch depending on what you labeled. Validation of the model and control of the data are two different jobs.

The takeaway

You will never validate the weights of a vendor LLM, and you should stop trying to design a program that pretends you can. The defensible path is output validation against a labeled benchmark, and the benchmark is the part that decides whether your evidence is real. Build it on numbers that reconcile to the ledger, run the math deterministically so the model can’t fabricate, keep every answer traceable to source, and quarterly validation becomes a routine you can show an examiner without flinching.

If you’re standing up a validation program for a finance LLM and the benchmark is the piece that’s missing, that reconciled ground-truth layer is exactly what we built. Book a demo and we’ll walk through how the benchmark falls out of the model.

For the bigger picture, the governing AI in finance pillar covers how this fits the wider control framework, including what the 2026 model-risk framework changes and the reference architecture for an AI control stack that holds it all together.

Part of Governing AI in Finance: Model Risk, Controls, and Validation for the LLM Era

Keep reading

Book a demo

See your numbers tie out.

Book a 30-minute demo. Bring a question you can never answer fast enough, and we will model it live against real financial data.