How can I be sure the numbers are right?

Every figure in a Finn report is fetched from source, computed in a spreadsheet, and reviewed blind before delivery. The full architecture, and what it costs.

This is one of the most common questions we have to answer on every call with a new prospect. The concern is valid - a language model cannot tell the difference between a figure it verified and a figure it half remembers from training, and it states both with the same confidence.

Considering that risk, Finn is built so that fetching the data and the model is never trusted with a number. The model reads and judges. Code around it fetches every figure from source, a spreadsheet does the arithmetic, and a set of checks decides whether the report can go out.

Every Finn report passes four checkpoints before it is sent.

Checkpoint

What it tests

What stops delivery

Source resolution

Every figure fetched on this run, through a fixed order of sources

A field no source can resolve

Workbook

The arithmetic, plus identities like market value against price times shares

A table that disagrees with the workbook behind it

Mechanical checks

Citations resolve, tables match the workbook, links are live

A citation pointing at nothing

Blind review

Whether the captured evidence proves each written claim

Any finding still open at the end of the run

What the LLM is allowed to do

The process is run by code, not by an LLM, and the same inputs go through the same steps every time. A run starts when a message arrives, a calendar event is scheduled, or a recurring commitment is due. The code works out what is due, assembles the material the model will work from, applies the rules below, and decides what happens to the output.

The model starts each run from a briefing pack of cited sources. Anything already covered in the last comparable report is left out so it does not come back as new, and whatever the model remembers about the company from training has no place in the report. The model's output then goes through the checks before anything is delivered.

Where each figure comes from

No figure in a Finn report comes from the model's memory. Prices, share counts, market values, exchange rates, and multiples are fetched on the run that uses the data. A remembered figure is treated as a defect even when it turns out to be correct, because afterwards there is no way to tell which of the remembered figures were right by luck.

Each field is looked up in a fixed order: licensed market data providers first, then the primary documents and cited sources behind them. This happens field by field, so if the first source cannot fill a field, the system moves to the next source for that field. When a report says "not disclosed," it means every source was tried for that field, and none had it.

Figures that matter to a conclusion are also checked against the company's or the exchange's own documents. When a data provider and press coverage disagree, the company's own filing takes precedence, and the report explains which figure was used and why.

Every price shows its as-of date and basis. Every currency conversion shows the pair, the rate, the provider, and the timestamp, and is calculated by a tool rather than from a remembered rate. Each figure is labeled as one of three things: company disclosure, consensus estimate, or Finn calculation. These are never mixed into a single number.

The arithmetic happens in Excel

Language models get arithmetic wrong often enough that Finn does not let the model do any. Tables are built in an Excel workbook with live formulas, and the report shows what the workbook computed. The workbook is sent with the report, formulas intact, so you can open it, trace a number back to its inputs, or change an assumption and see what moves.

The system also checks its own output against relationships that have to hold. Market value has to equal price times shares. A stated percentage change has to match the two figures printed next to it. These checks need no judgment, so they run on every report.

The blind review

After the mechanical checks pass, the report goes to a second model, a different one from the model that wrote it. If I were evaluating a vendor, this is the part I would push hardest on.

The reviewer has no web access and no tools. The only thing it sees is the evidence that was captured for each claim while the research was being done. Its job is to check whether that evidence supports the claim as written. We keep it isolated on purpose. A reviewer that can search will go and find its own, better evidence and grade the report against that, which is an easier job and misses unsupported claims.

The review asks five things of every claim:

  1. Do the numbers, tables and charts agree with each other?

  2. Does the evidence prove the claim, or has an inference been written up as if it were a company disclosure?

  3. Is it the right company, the right period, the right event date?

  4. Does the conclusion match what was actually gathered?

  5. What did the report leave out?

Each claim comes back as verified, flagged, or unverified. If the review cannot run under these conditions, it fails, and the report is held.

Flagged findings go into a single repair pass, limited to the parts of the report they affect. If any one finding cannot be verified and tied to a specific place in the report, the whole repair is rejected, and nothing is applied, because a partly repaired report has not been checked end-to-end by anyone.

A model never edits a workbook. If the defect is in the spreadsheet, the spreadsheet is regenerated from the source data, and the report waits for it.

What happens when a check fails

If the review is unavailable or incomplete at the end of a run, or any finding is still open, the report is held. Every finding has to be closed; each closure is verified separately, and all internal review notes are removed before the report is delivered.

There is one exception. If the review provider runs out of quota or time halfway through a run, the report is delivered with a clear disclaimer and the full audit appendix attached, so you can see which items were still open. That is the only case where review material reaches a reader.

None of this can be switched off. You can change base currency, risk framing, writing style, output format, and delivery schedule. The checks and the hold rule are not settings.


The path a Finn report takes before it is sent. Four checkpoints, each of which can stop delivery.

Finn report processing flow with the four checkpoints and what stops delivery at each one.

What this design costs

All of this takes time. A typical ad-hoc request comes back in about an hour, longer when the work is heavier, and scheduled monitoring is set up to arrive one to two hours before its calendar time. Nothing in Finn is real-time. A price move gets picked up when the next monitoring run covers it.

We think that is the right trade-off for this kind of work. If you need an instant answer, use a chat assistant. We use them every day ourselves.

Could I just build this with Claude?

Part of it, yes, and it is worth trying. Give Claude a name, a filing, and a question, and you get a decent first draft in a few minutes. That is the same reading and judging that the model does inside Finn. What you do not get is anything in the table at the top of this post. For how that plays out on a real question, see The chatbot answered, Finn did the work.

To get the rest, you would build it yourself. Market data with programmatic access, which is a separate contract from a terminal seat, and code that fetches each field in a fixed order and moves to the next source when one is empty. Code that builds the Excel workbook with live formulas from that data and checks the report's tables against it. Evidence capture, so every claim the model writes is stored with the source it came from, otherwise the review has nothing to check. A second model with no tools, its own prompts, and the logic that holds a report when a finding stays open. And something that wakes up when a filing lands or a date comes due, runs the whole chain, and puts the result in your inbox.

The higher cost comes after the initial setup works. Model providers ship new versions and prompts that behaved last quarter start behaving differently. Data providers change their APIs. Checks that were tight drift loose. Someone has to own that, and in a fund of five to twenty people that someone is you, on a day you also have positions to manage.

If you have the time and like building, do it - we'll gladly provide the high-level architecture to help you get started. If you would rather spend that time investing, that is what Finn is for.

Reading a Finn report

Every number has its source and as-of date next to it. The workbook that produced the tables is attached with formulas intact. The disclosure, consensus, and calculation labels tell you what kind of number you are looking at. Checking any single figure takes a few seconds.

The judgment stays with you. Finn does what a junior analyst does in most funds: it gathers and verifies the evidence behind a view. The view itself, whether to buy or sell, how much, and when, is yours.

This post is about accuracy. Where your data sits, how it is isolated, and how long it is kept is covered in Data, privacy, and security at Finn. If you want to see this on your own portfolio, request access and send one or two names you cover. The first report comes back the next morning.