LLM as a Judge: Why a Pro Report Says Less, Not More
LLM as a judge, applied to stock research: a fourth model deletes every claim without a source, which is why a Full report pro runs shorter, not longer.
See a sample reportThis is a method, not a recommendation. Nothing here, or anywhere else on Taufolio, is investment advice. Treat every example as a starting point for your own research.
Running several AI models over one company gives you less than it promises. Three answers that agree are still one answer if all three misread the same paragraph. The value of Three models and a judge sits elsewhere: the fourth model, an LLM as a judge, deletes every sentence you cannot point to in a document. So a Full report pro often runs shorter than the standard one, not longer.
I lived off the case for multiple models once, so I will not weaken it. A few years before Taufolio I built an engine that matched people to things, and the best single model got roughly seven cases in ten right. One person in three got a match worth throwing away. An ensemble of five models, none of them best on its own, beat every one of them. Those are my numbers from my own project, not a market statistic, and the lesson stuck: one confident answer is a hypothesis, not a verdict.
Except that agreement between models measures similarity, not truth.
One model has one set of blind spots
Ask a single model whether a company is any good and you get one confident register throughout. It will not mark where it is guessing. A sentence about China exposure, where it knows very little, reads exactly like one about revenue it has counted in a table. Why that is too thin to carry an analysis is the subject of the piece on why one chatbot is the wrong tool for company analysis.
Nobody calls the same witness twice and calls the second appearance corroboration. You want a witness trained differently, primed by habit to worry about something else.
Three models agreeing is not proof
A shared error is the norm here, not an accident. In the FinanceBench benchmark, built from questions about real company filings, GPT-4-Turbo with a retrieval system answered incorrectly or refused to answer in 81% of 150 sampled cases, while the same model with the whole document in context was wrong 21% of the time. Which means the result turned less on which model you asked than on whether it saw the right page at all.
The study is from 2023 and models have moved a long way since, so I am not carrying the number into today. I am carrying the mechanism. The error is born where the document is difficult, not where the model is weak.
That has a consequence model voting says nothing about. Three models reading the same tangled lease footnote trip on it for the same reason and in the same direction. Agreement there is not three independent measurements, it is one measurement repeated three times, now with triple the confidence.
I think agreement between models is the most overrated number in this industry.
Independence has to be engineered in, because it does not arise on its own. The models draft separately, never see each other's work, never read each other's conclusions. A model that had seen another's draft would stop being a second witness and start being an echo, and an echo sounds like confirmation.
What an LLM judge actually throws out
The judge gets three versions of one analysis and has no right to pick the best-written one. It has two jobs: show where the versions diverge, and cut every sentence with no address in a source.
For that check to mean anything, all three models read the same current pack: the latest annual and quarterly filings, the earnings call transcript, verified news from the period. A model left with its own memory guesses what the company did last quarter, and it guesses fluently. A model with the document in front of it also gets things wrong, only now the mistake can be pointed at, which is the point of the procedure.
Say the first model reports a gross margin decline of 5 percentage points, the second says 2, and the third does not mention margin at all. Which number goes into the report?
A vote has nothing to count, because two different claims do not make a majority. So the judge goes back to the text. If a specific number sits in the 10-K on EDGAR, that number stays with its citation. If no number sits there, the report carries no sentence about margin, though three models had plenty to say about it.
That rule costs length, and it is supposed to. A report where every finding leads to a document you can open is shorter than one where some sentences sound clever and lead nowhere.
The judge has a weakness too
A model acting as a judge is surprisingly good at agreeing with people. In the MT-Bench work, a strong judge model matched human preferences more than 80% of the time, roughly how often humans agreed with each other. The same paper lists its flaws, and one is awkward for us: a judging model favours answers written in its own style.
So who watches the judge?
A judge asked which version is better is a fourth opinion, not a control. So the question is put differently: which exact sentence is supported by which exact passage. Checking text is immune to taste in a way grading style never will be.
There is a case where the judge does not help at all. If all three models read the same sentence of a filing in the same wrong way, the check passes: the address matches even though the understanding does not. That is why a citation is not decoration but a route to the sentence you open yourself when the point is contested.
The second cost lands on the reader. True claims that could not be pinned to a line leave with the invented ones, and nobody tells you how many there were. How many the judge cuts from a typical report we do not publish. I do not have that number in a form I would stand behind, and I would rather say so than dress an impression up as a metric.
Disagreement is the product
The most interesting part of a pro report is not the agreement but the split the judge did not resolve, because both sides had sources. Say the company makes server cabling. The models agree on its position in AI and split completely on whether copper survives. One builds the case on active copper cables being cheap and power-efficient over short runs between chips. The other on co-packaged optics maturing faster than the market assumes, which cuts the copper growth engine out.
Both cases stand on documents. That is not noise to average away, it is exactly where your own work starts.
None of this voting ends in a trade recommendation, and it never will. You get the advantages, the management tone, the risks and a map of the contested points; the verdict stays with you. Agreement between independent rankings is also where the name Taufolio came from, since tau is a measure of rank agreement, not a mascot.
When a Full report pro earns its credits
A Full report pro costs 400 credits against 100 for a Full report. So one pro report costs what four standard ones cost. The Free plan carries 100 credits a month, which makes a free month exactly one Full report and nothing more. How the currency works is what the piece on credits is for.
The difference only makes sense where being wrong is expensive. I reach for the pro version with three kinds of company:
- ones whose technology I do not understand well enough to catch nonsense in the product paragraph,
- ones where a single supply-chain link or a single subsidiary decides the whole result,
- ones inside a live narrative fight, where the price carries a lot of expectation.
A company you have followed for years, and the question "did anything change last quarter", do not need a judge. They need a faster document, and the tiers are compared in the comparison of the shorter report with the full one.
Once a pro report is in front of you, read it from the contested sections, not from page one. The section where the models split shows where the thesis is thinnest. It sets the list of things to check before the next quarter, and the rest of the document is background for that list.
The mechanism is called Three models and a judge and it sits behind the pro versions only. What survives its pass, and how that looks in a finished file, is set out in our methodology.
I would rather read work that admits to a gap than work that speaks about everything with equal certainty. The first kind can be checked.
Frequently asked questions
What is the Three models and a judge mechanism in Taufolio?
Can three AI models hallucinate the same thing?
What does the judge do when models disagree?
How is a Full report pro different from a Full report?
Which AI models does Taufolio use?
Why does a Full report pro cost more credits?
- methodology
- research
- product
Posts are produced with AI tools and go through editorial review by the Taufolio team before publishing.