Skip to content
MonitoringReportsMethodologyPricingBlogAboutStart for free
Back to blog

Method

AI Equity Research Without Citations Is Just a Story

AI equity research without citations is a story, not a finding. Why a citation constrains the model, plus a three-minute test for any AI stock tool you use.

The Taufolio team11 min read
See a sample report

This is a method, not a recommendation. Nothing here, or anywhere else on Taufolio, is investment advice. Treat every example as a starting point for your own research.

An AI hallucination is a sentence with the rhythm of a fact and no document underneath it. In equity research the dangerous ones sit on numbers, because a number gets read faster than it gets checked. The fix is not a smarter model but a rule: every claim that could change your view of a company points at a page you can open. AI equity research without those pointers stays a story about a company.

What a hallucination actually is

The word is misleadingly gentle, because it suggests a malfunction. There is no malfunction. A model picks the most probable next word, and it does that just as diligently when it has read the filing as when it has never seen it. Wording like "steadily expanded margins through operational efficiencies" fits almost any company in almost any quarter, so it is cheap to produce and rarely clashes with the paragraph around it.

The scale has been measured. In a study published in the Journal of Legal Analysis, models were asked verifiable questions about randomly selected federal court cases, and the share of answers containing an invented element ran from 58% for GPT-4 to 88% for Llama 2. Which means even the best model tested was wrong more often than right, on questions whose answers sit in a public document.

Law is not the stock market, but the question has the same shape: it concerns one specific document, and the answer arrives from a memory of how such documents usually sound. A filing is the harder case, because a judgment has one version while a segment margin has as many definitions as the company cares to write into a note.

What AI equity research gets wrong on a number

An invented adjective usually passes through without a trace. If a report says management sounded "cautiously optimistic", nothing follows from it and nothing lodges in your head.

A number behaves in the opposite way. Say a report gives a gross margin decline of 80 basis points, from 58.4% to 57.6%. Eight tenths of a point sounds harmless, so the obvious question: how much is that? On a billion in revenue it is eight million a year. If that number was in no filing, you have just built a thesis on eight million that does not exist.

A number has the nasty property of impersonating its own source. It looks like the output of someone's work, so it enters your notes unchallenged, then your thesis, and a year on you remember it as a fact whose origin you can no longer reconstruct.

Then there is the aggregation problem. In company research almost no single claim settles anything by itself; conviction is assembled from a dozen small things at once. Margin down three quarters running, inventories growing faster than sales, management quietly dropping a disclosure it used to give. If three of those fifteen small things have nothing behind them, the whole still reads as coherent, because coherence is a property of the prose and not of the evidence under it. You usually notice at the next set of results, when one of the small things refuses to repeat and there is nowhere to check where it came from.

A citation is a muzzle, not a courtesy

Citations have the whiff of a dissertation about them: proof of diligence, a nod to the reader, something you add at the end. In that version a citation is a cost borne out of politeness and the first thing cut when the text has to be shorter. It sounds reasonable.

Except that inside a system generating text, a citation does its work long before any reader sees it.

To write "page 14 of the quarterly filing", the system has to fetch page 14 and hold it in view while the sentence is written. The order reverses, and the document becomes the precondition of the sentence existing at all. A model that did not find the page has nothing to build a citation from, so either it does not write the claim or it writes it naked, and a naked claim is visible from across the room.

A muzzle changes exactly one thing about a dog: the range of what it can do.

The reader benefit, all that ability to verify, is a side effect. The same constraint operates when nobody ever clicks a single link, which is why asking about sources tells you more about how a report was produced than about how it reads. It shows most sharply where nobody handed the model a document at all: ChatGPT asked to analyze a stock writes from a memory of companies rather than from their filings.

Two sentences about the same margin

The difference shows best in a pair of sentences about the same quarter. Take an invented enterprise software company and two versions of one paragraph.

Without a source: "The company faces growing competition in its core segment, and management has acknowledged margin pressure in recent communications."

With a source: "On the third-quarter earnings call the CFO described 'incremental pricing pressure from two new entrants in our enterprise segment' and put its effect on gross margin at roughly 80 basis points (transcript, prepared remarks). Segment gross margin fell quarter over quarter from 58.4% to 57.6% (quarterly filing, segment results, p. 14)."

The first sentence is shorter and cannot be checked. "Recent communications" could mean a call last week or an interview last year. "Margin pressure" holds ten basis points and three hundred equally well. "Acknowledged" covers both a full accounting of the damage and a graceful sidestep.

The second is longer because it is doing work. It names the speaker, the event and the document, ties the description to the number, and shows where the number sits. That is enough for someone who disagrees to know which page to open.

The three-minute test

None of this has to be taken on trust, because you can measure it yourself, once, on any tool.

  1. Pick one claim that would genuinely change your view of the company, ideally one with a number in it.
  2. See which document it points to: name, period, section or page.
  3. Open that document. US companies file everything with the SEC through EDGAR, and the SEC's investor guide explains how to read an annual 10-K.
  4. Find the number and read the two sentences around it.

Step two settles most cases and takes thirty seconds. If the document slot holds "market data", "recent communications" or nothing at all, you can stop there. If a document is named but the number is not in it, you have learned something better: the tool can point at sources and stretch them anyway.

Step four is the one people skip, and it catches the most. The number is often real while the sentence beside it says something the report does not: the decline turns out to be a one-off writedown, the growth an acquisition closed mid-quarter, the record margin a building that got sold. Two sentences of context cost about fifteen seconds and regularly reverse the meaning of a whole paragraph. Which document to open for which kind of claim is worked through in the piece on reading filings instead of headlines.

Run the test on a company you know well rather than one you are trying to learn about. On familiar ground you will catch a stretched sentence before you even open the document, so you are measuring the tool rather than your own ignorance.

Three citations that prove nothing

The test above usually fails in a specific way: the link is there, and it leads nowhere checkable. How do you spot a decorative one? Three varieties turn up regularly.

The first is a citation to a front door: the link points at an investor relations site or a filings search, so formally a source is named and practically you have been handed an address to search yourself. A good citation ends at the name of a document and the number of a section.

Next comes the citation to somebody else's narrative: a claim about margin points at an article, that article points at another article, and the last one cites "market reports". You can travel four links this way and never open a filing. What you check then is what somebody else wrote, and you come back knowing what you knew before.

Worst of the three is the citation to the wrong period. The number is real, the document is real, the quarter is not. Step three of the test comes back clean until you look at the date, which is why the period is the one thing worth checking even when everything else looks fine.

What a citation will not fix

It would be convenient to stop here. Except that a citation solves exactly one problem, and the research on tools that retrieve documents before writing shows how much is left.

The Journal of Empirical Legal Studies published a test of three commercial legal tools that advertised themselves as hallucination-free precisely because they consult a document database first. The share of answers with an invented or mischaracterised element ran from 17% to 33% across the three, against 43% for GPT-4 alone. Which means retrieval cut the error rate, in the best case by more than half, and made none of the three reliable.

Two failures walk straight through a citation. The link leads to a real document that says something other than the sentence next to it. Or the document says exactly that, but it is three quarters old and the company has since cut its segments differently.

So does that mean sources achieve nothing? I think they achieve about as much as anything can, because they change the kind of work you have to do. Checking one sentence without a source means reading an entire filing, so nobody checks. With the page named it takes a minute, and verification stops being an evening project.

I do not know whether models will ever stop inventing things. I do know that while they do, the only defence that does not depend on how good the model is sits outside the model: a document open next to the sentence.

When a missing source is honest

None of this means every sentence needs a link. A report with a footnote on every comma is unreadable, and at that density citations start impersonating rigour rather than supplying it.

The line runs along load-bearing weight. A sentence that connects paragraphs, describes a well-known industry mechanism or states the author's judgement stands on its own, provided it is marked as judgement. A sentence carrying a number, a management quote, a date or a comparison with a competitor points at a document, because otherwise there is no way to tell a finding from an impression. The same line runs through reports written by people: in a sell-side equity research report the forecast is the author's judgement, while the segment table points at a filing.

There is a third category nobody talks about, and it says a lot about a tool: the sentence admitting a document does not exist. The company does not break revenue out by region, management did not answer the question about churn, the latest transcript is silent on pricing. A report that says so hands the reader real information, because a missing disclosure is itself a finding. A report that fills the gap with something smooth about "a stable geographic mix" has swapped absent data for its own guess and marked the swap nowhere.

Where this lives in the report

At Taufolio the rule sits where the text is produced rather than in a policy page. The Full report is written with the documents open alongside it, and its findings point at the filing, the transcript or the press release they came from. In an Earnings brief a management quote opens the Transcript, so you see the sentence in the original and what was said immediately after it, not just a paraphrase.

Full report pro adds the Three models and a judge mechanism: three models write the same analysis independently, a fourth compares the versions and keeps what a source will carry. That catches a different error than the citation does. A citation makes sure a sentence has a document. The judge checks something else entirely, namely whether three independent readings of that document reach the same conclusion. A finding that showed up in one model and vanished in the other two is usually precisely the thing the document does not contain. The mechanism itself is taken apart in three models and a judge, and the methodology page shows where the data under each section comes from.

I now check tools in the reverse of the instinctive order: links first, prose second. When the links hold, the prose usually holds too, because it was written under the same constraint. If the first link leads nowhere, I close the report and do not read the rest.

Frequently asked questions

A hallucination is a sentence assembled out of probable words rather than out of a document; the model predicts what an answer sounds like, and it does that just as confidently when it has read nothing. On a number the cost is practical: you copy a figure into a spreadsheet in a second, and a year later tracing where it came from takes an hour or fails outright. An invented adjective usually falls apart on its own by the next quarter, because nothing followed from it; an invented number has by then propped up three other conclusions.
Because a source does its work while the text is being written, not while it is being read: a system required to name a page has to fetch that page first, which leaves it nowhere to invent. The side effect is the price of verification, which drops from "read the whole filing" to "open one page", and at that price people actually check. A report with no pointers asks you to trust an author you do not know and cannot question.
It names a primary document precisely enough to open in about a minute: the filing, the period, the section or page, and for a quote also the event and the speaker. "Market data" and "recent communications" are not citations, because you cannot open anything on the strength of them. The test is whether you can reach the number without asking the author where it came from.
No, and a report with a footnote on every comma is unreadable. A sentence that connects paragraphs or states the author's judgement stands on its own, as long as it is marked as judgement; a sentence carrying a number, a management quote or a date points at a document. There is a third case that looks like a missing source and is in fact a finding: the sentence saying the company does not disclose something, revenue by region for instance. That one cannot be backed by a document, because its content is the absence of the document.
Take one claim with a number in it and work from the claim to the document, never the other way round. The most common result is not "the number is invented" but "the number is real and the sentence beside it says something the filing does not": the margin decline turns out to be a one-off writedown, the revenue growth an acquisition closed mid-quarter. Second most common is a real number from the right document and the wrong quarter, which passes every check until you look at the date. Both show up when you read the two sentences around the figure.
It can be, and sometimes is, because the model is a well-calibrated guesser. The problem is that without a source you cannot separate the lucky guess from the confident mistake, and on the page those two look identical. A report you cannot check does not give you knowledge about the company, only somebody's opinion about it.
  • research
  • sources
  • methodology
Share:XLinkedIn

Posts are produced with AI tools and go through editorial review by the Taufolio team before publishing.