Skip to content
MonitoringReportsMethodologyPricingBlogAboutStart for free
Back to blog

Method

LLM as a Judge: Why a Pro Report Says Less, Not More

LLM as a judge, applied to stock research: a fourth model deletes every claim without a source, which is why a Full report pro runs shorter, not longer.

The Taufolio team7 min read
See a sample report

This is a method, not a recommendation. Nothing here, or anywhere else on Taufolio, is investment advice. Treat every example as a starting point for your own research.

Running several AI models over one company gives you less than it promises. Three answers that agree are still one answer if all three misread the same paragraph. The value of Three models and a judge sits elsewhere: the fourth model, an LLM as a judge, deletes every sentence you cannot point to in a document. So a Full report pro often runs shorter than the standard one, not longer.

I lived off the case for multiple models once, so I will not weaken it. A few years before Taufolio I built an engine that matched people to things, and the best single model got roughly seven cases in ten right. One person in three got a match worth throwing away. An ensemble of five models, none of them best on its own, beat every one of them. Those are my numbers from my own project, not a market statistic, and the lesson stuck: one confident answer is a hypothesis, not a verdict.

Except that agreement between models measures similarity, not truth.

One model has one set of blind spots

Ask a single model whether a company is any good and you get one confident register throughout. It will not mark where it is guessing. A sentence about China exposure, where it knows very little, reads exactly like one about revenue it has counted in a table. Why that is too thin to carry an analysis is the subject of the piece on why one chatbot is the wrong tool for company analysis.

Nobody calls the same witness twice and calls the second appearance corroboration. You want a witness trained differently, primed by habit to worry about something else.

Three models agreeing is not proof

A shared error is the norm here, not an accident. In the FinanceBench benchmark, built from questions about real company filings, GPT-4-Turbo with a retrieval system answered incorrectly or refused to answer in 81% of 150 sampled cases, while the same model with the whole document in context was wrong 21% of the time. Which means the result turned less on which model you asked than on whether it saw the right page at all.

The study is from 2023 and models have moved a long way since, so I am not carrying the number into today. I am carrying the mechanism. The error is born where the document is difficult, not where the model is weak.

That has a consequence model voting says nothing about. Three models reading the same tangled lease footnote trip on it for the same reason and in the same direction. Agreement there is not three independent measurements, it is one measurement repeated three times, now with triple the confidence.

I think agreement between models is the most overrated number in this industry.

Independence has to be engineered in, because it does not arise on its own. The models draft separately, never see each other's work, never read each other's conclusions. A model that had seen another's draft would stop being a second witness and start being an echo, and an echo sounds like confirmation.

What an LLM judge actually throws out

The judge gets three versions of one analysis and has no right to pick the best-written one. It has two jobs: show where the versions diverge, and cut every sentence with no address in a source.

For that check to mean anything, all three models read the same current pack: the latest annual and quarterly filings, the earnings call transcript, verified news from the period. A model left with its own memory guesses what the company did last quarter, and it guesses fluently. A model with the document in front of it also gets things wrong, only now the mistake can be pointed at, which is the point of the procedure.

Say the first model reports a gross margin decline of 5 percentage points, the second says 2, and the third does not mention margin at all. Which number goes into the report?

A vote has nothing to count, because two different claims do not make a majority. So the judge goes back to the text. If a specific number sits in the 10-K on EDGAR, that number stays with its citation. If no number sits there, the report carries no sentence about margin, though three models had plenty to say about it.

That rule costs length, and it is supposed to. A report where every finding leads to a document you can open is shorter than one where some sentences sound clever and lead nowhere.

The judge has a weakness too

A model acting as a judge is surprisingly good at agreeing with people. In the MT-Bench work, a strong judge model matched human preferences more than 80% of the time, roughly how often humans agreed with each other. The same paper lists its flaws, and one is awkward for us: a judging model favours answers written in its own style.

So who watches the judge?

A judge asked which version is better is a fourth opinion, not a control. So the question is put differently: which exact sentence is supported by which exact passage. Checking text is immune to taste in a way grading style never will be.

There is a case where the judge does not help at all. If all three models read the same sentence of a filing in the same wrong way, the check passes: the address matches even though the understanding does not. That is why a citation is not decoration but a route to the sentence you open yourself when the point is contested.

The second cost lands on the reader. True claims that could not be pinned to a line leave with the invented ones, and nobody tells you how many there were. How many the judge cuts from a typical report we do not publish. I do not have that number in a form I would stand behind, and I would rather say so than dress an impression up as a metric.

Disagreement is the product

The most interesting part of a pro report is not the agreement but the split the judge did not resolve, because both sides had sources. Say the company makes server cabling. The models agree on its position in AI and split completely on whether copper survives. One builds the case on active copper cables being cheap and power-efficient over short runs between chips. The other on co-packaged optics maturing faster than the market assumes, which cuts the copper growth engine out.

Both cases stand on documents. That is not noise to average away, it is exactly where your own work starts.

None of this voting ends in a trade recommendation, and it never will. You get the advantages, the management tone, the risks and a map of the contested points; the verdict stays with you. Agreement between independent rankings is also where the name Taufolio came from, since tau is a measure of rank agreement, not a mascot.

When a Full report pro earns its credits

A Full report pro costs 400 credits against 100 for a Full report. So one pro report costs what four standard ones cost. The Free plan carries 100 credits a month, which makes a free month exactly one Full report and nothing more. How the currency works is what the piece on credits is for.

The difference only makes sense where being wrong is expensive. I reach for the pro version with three kinds of company:

  • ones whose technology I do not understand well enough to catch nonsense in the product paragraph,
  • ones where a single supply-chain link or a single subsidiary decides the whole result,
  • ones inside a live narrative fight, where the price carries a lot of expectation.

A company you have followed for years, and the question "did anything change last quarter", do not need a judge. They need a faster document, and the tiers are compared in the comparison of the shorter report with the full one.

Once a pro report is in front of you, read it from the contested sections, not from page one. The section where the models split shows where the thesis is thinnest. It sets the list of things to check before the next quarter, and the rest of the document is background for that list.

The mechanism is called Three models and a judge and it sits behind the pro versions only. What survives its pass, and how that looks in a finished file, is set out in our methodology.

I would rather read work that admits to a gap than work that speaks about everything with equal certainty. The first kind can be checked.

Frequently asked questions

Three models get the same pack of company documents and write the analysis independently, without seeing the other versions. A fourth model, which has no name of its own, lines the versions up, marks where they split, and removes any claim that cannot be pinned to a specific passage. That is how a Full report pro is produced.
They can, and that is the weak point of every model vote. A tangled footnote or an unusual way of presenting a number can push all three the same way, because the difficulty sits in the document rather than in the model. Agreement on its own settles nothing; going back to the source does.
It does not average them and it does not pick the version that reads better. It goes back to the filing and checks which version the text supports. If none of them does, the claim leaves the report; if the split is substantive and both sides cite something real, it is written up as a split.
The questions are the same: business, competition, management, risks, valuation. The quality control differs, because the pro version is written three times independently and then goes through the judge. The result is often a shorter document than the standard one, with fewer sentences you cannot check.
The mechanism is fixed, the line-up is not. Frontier models from different providers write the same analysis independently, and a separate model, deliberately unnamed, checks their sentences against the documents. We do not publish which ones, because the composition changes whenever a stronger model ships and a list of brands would go stale faster than this page. Three models and a judge names the procedure, not the logos.
Because the analysis is produced three times and then passes a separate verification stage, so it consumes several times the model work. A Full report pro costs 400 credits against 100 for a Full report. You are paying for what gets cut, not for extra pages.
  • methodology
  • research
  • product
Share:XLinkedIn

Posts are produced with AI tools and go through editorial review by the Taufolio team before publishing.