Skip to content
New: ask the Rexfin Analyst Agent about your model. Every figure comes back cited.
· 6 min read

Turning a Printed Label Into a Concept You Can Actually Compare

Rexfin maps every statement line to a stable IFRS taxonomy concept and canonical predicate through a hand-verified dictionary, so comparisons hold across companies, years, and languages.

By The Rexfin team

Two companies both report “Property, plant and equipment.” A third calls it “Fixed assets, net.” A fourth splits it across two lines that a fifth reports as one. Every one of those labels might mean the same underlying concept, or might not, and an AI tool that treats label text as the source of truth will happily compare things that don’t belong side by side. Rexfin doesn’t compare labels. It compares concepts.

What tagging actually does

Every line on a company’s face statements (balance sheet, income statement, cash flow) gets mapped to two things at ingestion: a taxonomy_element from the ifrs-full taxonomy, and a normalized canonical_predicate that Rexfin uses internally to identify what that line is, independent of how the filer chose to word it. “Fixed assets, net” and “Property, plant and equipment, net” land on the same predicate. A subtotal like “Total non-current assets” gets its own distinct predicate rather than silently inheriting one of its components, which matters more than it sounds like, because a subtotal sharing a predicate with a line item it contains is exactly how double-counting sneaks into a rollup nobody double-checks.

The mapping runs off a hand-verified dictionary, not a model guessing at synonyms on the fly. That’s a deliberate constraint: a canonical predicate is a promise that every number carrying it means the same thing everywhere it appears, and that promise only holds if the mapping was checked by someone rather than inferred fresh each time a new filing’s wording shows up.

Why an untagged number is a blocked number

Rexfin’s ingestion rule is that a face-statement figure without a non-null taxonomy tag and canonical predicate is not seedable into the model at all. That sounds strict, and it is meant to be. The alternative, letting an unmapped number sit in the model with a best-guess label, is exactly the kind of soft failure that looks fine until someone builds a comparison or a trend line on top of it and gets a result that’s quietly wrong. “Zero manual entry” only means something if every number that enters the model has already been resolved to a concept a human vetted, not one an algorithm improvised under time pressure.

Why this is the layer comparability actually runs on

This tagging step is what makes quick company comparables possible without someone manually reconciling label differences first. It’s also what a cited AI answer is actually referencing when it says “here is total assets”: the AI isn’t matching your question against a page’s raw text, it’s matching it against the canonical predicate the tagging step already resolved, then pointing back to the exact cited figure behind it.

The taxonomy element also carries metadata for IFRS-18 category and operating-profit-subtotal status alongside the mapping. That matters because IFRS-18 changes how operating profit and its subtotals have to be presented on the face of the income statement, and the comparative period for that transition needs to hold up under the new categorization without re-extracting the filing from scratch. Carrying that metadata now, at tagging time, is what lets a company’s historical statements stay usable once the new presentation requirements apply, instead of forcing a re-ingest the year the standard takes effect.

The honest scope today

At this stage, the dictionary covers the full set of face-statement lines needed for a clean, no-double-count mapping across income statement, balance sheet, and cash flow: verified concept by concept rather than assumed. Extending that same rigor to a scored fuzzy match across the entire IFRS vocabulary, so an unusual or highly bespoke label finds its concept without a person adding it to the dictionary first, is later-stage work. Until that lands, an unusual label a filer invents for a line that doesn’t cleanly match anything in the dictionary is exactly the kind of case Rexfin would rather flag than force-map.

Who reaches for this

Anyone building a comparison across companies or years who has been burned before by two filings using different words for the same line, or the same word for two different things. If your analysis depends on “total assets” meaning the same thing every time it appears in the model, this is the step that makes that true rather than assumed. It sits directly downstream of document intake and classification, and feeds everything else in the product tour.

Part of Rexfin Product Tour: Every Number Traceable

Keep reading

Book a demo

See your numbers tie out.

Book a 30-minute demo. Bring a question you can never answer fast enough, and we will model it live against real financial data.