AI Equity Research Without Citations Is Just a Story
AI equity research without citations is a story, not a finding. Why a citation constrains the model, plus a three-minute test for any AI stock tool you use.
See a sample reportThis is a method, not a recommendation. Nothing here, or anywhere else on Taufolio, is investment advice. Treat every example as a starting point for your own research.
An AI hallucination is a sentence with the rhythm of a fact and no document underneath it. In equity research the dangerous ones sit on numbers, because a number gets read faster than it gets checked. The fix is not a smarter model but a rule: every claim that could change your view of a company points at a page you can open. AI equity research without those pointers stays a story about a company.
What a hallucination actually is
The word is misleadingly gentle, because it suggests a malfunction. There is no malfunction. A model picks the most probable next word, and it does that just as diligently when it has read the filing as when it has never seen it. Wording like "steadily expanded margins through operational efficiencies" fits almost any company in almost any quarter, so it is cheap to produce and rarely clashes with the paragraph around it.
The scale has been measured. In a study published in the Journal of Legal Analysis, models were asked verifiable questions about randomly selected federal court cases, and the share of answers containing an invented element ran from 58% for GPT-4 to 88% for Llama 2. Which means even the best model tested was wrong more often than right, on questions whose answers sit in a public document.
Law is not the stock market, but the question has the same shape: it concerns one specific document, and the answer arrives from a memory of how such documents usually sound. A filing is the harder case, because a judgment has one version while a segment margin has as many definitions as the company cares to write into a note.
What AI equity research gets wrong on a number
An invented adjective usually passes through without a trace. If a report says management sounded "cautiously optimistic", nothing follows from it and nothing lodges in your head.
A number behaves in the opposite way. Say a report gives a gross margin decline of 80 basis points, from 58.4% to 57.6%. Eight tenths of a point sounds harmless, so the obvious question: how much is that? On a billion in revenue it is eight million a year. If that number was in no filing, you have just built a thesis on eight million that does not exist.
A number has the nasty property of impersonating its own source. It looks like the output of someone's work, so it enters your notes unchallenged, then your thesis, and a year on you remember it as a fact whose origin you can no longer reconstruct.
Then there is the aggregation problem. In company research almost no single claim settles anything by itself; conviction is assembled from a dozen small things at once. Margin down three quarters running, inventories growing faster than sales, management quietly dropping a disclosure it used to give. If three of those fifteen small things have nothing behind them, the whole still reads as coherent, because coherence is a property of the prose and not of the evidence under it. You usually notice at the next set of results, when one of the small things refuses to repeat and there is nowhere to check where it came from.
A citation is a muzzle, not a courtesy
Citations have the whiff of a dissertation about them: proof of diligence, a nod to the reader, something you add at the end. In that version a citation is a cost borne out of politeness and the first thing cut when the text has to be shorter. It sounds reasonable.
Except that inside a system generating text, a citation does its work long before any reader sees it.
To write "page 14 of the quarterly filing", the system has to fetch page 14 and hold it in view while the sentence is written. The order reverses, and the document becomes the precondition of the sentence existing at all. A model that did not find the page has nothing to build a citation from, so either it does not write the claim or it writes it naked, and a naked claim is visible from across the room.
A muzzle changes exactly one thing about a dog: the range of what it can do.
The reader benefit, all that ability to verify, is a side effect. The same constraint operates when nobody ever clicks a single link, which is why asking about sources tells you more about how a report was produced than about how it reads. It shows most sharply where nobody handed the model a document at all: ChatGPT asked to analyze a stock writes from a memory of companies rather than from their filings.
Two sentences about the same margin
The difference shows best in a pair of sentences about the same quarter. Take an invented enterprise software company and two versions of one paragraph.
Without a source: "The company faces growing competition in its core segment, and management has acknowledged margin pressure in recent communications."
With a source: "On the third-quarter earnings call the CFO described 'incremental pricing pressure from two new entrants in our enterprise segment' and put its effect on gross margin at roughly 80 basis points (transcript, prepared remarks). Segment gross margin fell quarter over quarter from 58.4% to 57.6% (quarterly filing, segment results, p. 14)."
The first sentence is shorter and cannot be checked. "Recent communications" could mean a call last week or an interview last year. "Margin pressure" holds ten basis points and three hundred equally well. "Acknowledged" covers both a full accounting of the damage and a graceful sidestep.
The second is longer because it is doing work. It names the speaker, the event and the document, ties the description to the number, and shows where the number sits. That is enough for someone who disagrees to know which page to open.
The three-minute test
None of this has to be taken on trust, because you can measure it yourself, once, on any tool.
- Pick one claim that would genuinely change your view of the company, ideally one with a number in it.
- See which document it points to: name, period, section or page.
- Open that document. US companies file everything with the SEC through EDGAR, and the SEC's investor guide explains how to read an annual 10-K.
- Find the number and read the two sentences around it.
Step two settles most cases and takes thirty seconds. If the document slot holds "market data", "recent communications" or nothing at all, you can stop there. If a document is named but the number is not in it, you have learned something better: the tool can point at sources and stretch them anyway.
Step four is the one people skip, and it catches the most. The number is often real while the sentence beside it says something the report does not: the decline turns out to be a one-off writedown, the growth an acquisition closed mid-quarter, the record margin a building that got sold. Two sentences of context cost about fifteen seconds and regularly reverse the meaning of a whole paragraph. Which document to open for which kind of claim is worked through in the piece on reading filings instead of headlines.
Run the test on a company you know well rather than one you are trying to learn about. On familiar ground you will catch a stretched sentence before you even open the document, so you are measuring the tool rather than your own ignorance.
Three citations that prove nothing
The test above usually fails in a specific way: the link is there, and it leads nowhere checkable. How do you spot a decorative one? Three varieties turn up regularly.
The first is a citation to a front door: the link points at an investor relations site or a filings search, so formally a source is named and practically you have been handed an address to search yourself. A good citation ends at the name of a document and the number of a section.
Next comes the citation to somebody else's narrative: a claim about margin points at an article, that article points at another article, and the last one cites "market reports". You can travel four links this way and never open a filing. What you check then is what somebody else wrote, and you come back knowing what you knew before.
Worst of the three is the citation to the wrong period. The number is real, the document is real, the quarter is not. Step three of the test comes back clean until you look at the date, which is why the period is the one thing worth checking even when everything else looks fine.
What a citation will not fix
It would be convenient to stop here. Except that a citation solves exactly one problem, and the research on tools that retrieve documents before writing shows how much is left.
The Journal of Empirical Legal Studies published a test of three commercial legal tools that advertised themselves as hallucination-free precisely because they consult a document database first. The share of answers with an invented or mischaracterised element ran from 17% to 33% across the three, against 43% for GPT-4 alone. Which means retrieval cut the error rate, in the best case by more than half, and made none of the three reliable.
Two failures walk straight through a citation. The link leads to a real document that says something other than the sentence next to it. Or the document says exactly that, but it is three quarters old and the company has since cut its segments differently.
So does that mean sources achieve nothing? I think they achieve about as much as anything can, because they change the kind of work you have to do. Checking one sentence without a source means reading an entire filing, so nobody checks. With the page named it takes a minute, and verification stops being an evening project.
I do not know whether models will ever stop inventing things. I do know that while they do, the only defence that does not depend on how good the model is sits outside the model: a document open next to the sentence.
When a missing source is honest
None of this means every sentence needs a link. A report with a footnote on every comma is unreadable, and at that density citations start impersonating rigour rather than supplying it.
The line runs along load-bearing weight. A sentence that connects paragraphs, describes a well-known industry mechanism or states the author's judgement stands on its own, provided it is marked as judgement. A sentence carrying a number, a management quote, a date or a comparison with a competitor points at a document, because otherwise there is no way to tell a finding from an impression. The same line runs through reports written by people: in a sell-side equity research report the forecast is the author's judgement, while the segment table points at a filing.
There is a third category nobody talks about, and it says a lot about a tool: the sentence admitting a document does not exist. The company does not break revenue out by region, management did not answer the question about churn, the latest transcript is silent on pricing. A report that says so hands the reader real information, because a missing disclosure is itself a finding. A report that fills the gap with something smooth about "a stable geographic mix" has swapped absent data for its own guess and marked the swap nowhere.
Where this lives in the report
At Taufolio the rule sits where the text is produced rather than in a policy page. The Full report is written with the documents open alongside it, and its findings point at the filing, the transcript or the press release they came from. In an Earnings brief a management quote opens the Transcript, so you see the sentence in the original and what was said immediately after it, not just a paraphrase.
Full report pro adds the Three models and a judge mechanism: three models write the same analysis independently, a fourth compares the versions and keeps what a source will carry. That catches a different error than the citation does. A citation makes sure a sentence has a document. The judge checks something else entirely, namely whether three independent readings of that document reach the same conclusion. A finding that showed up in one model and vanished in the other two is usually precisely the thing the document does not contain. The mechanism itself is taken apart in three models and a judge, and the methodology page shows where the data under each section comes from.
I now check tools in the reverse of the instinctive order: links first, prose second. When the links hold, the prose usually holds too, because it was written under the same constraint. If the first link leads nowhere, I close the report and do not read the rest.
Frequently asked questions
What is an AI hallucination, and why is it worse on a number?
Why do sources in an AI company report matter so much?
What does a good citation in a stock report look like?
Does every sentence in a research report need a source?
How do I check an AI stock research tool in three minutes?
Can an AI report be right even without citations?
- research
- sources
- methodology
Posts are produced with AI tools and go through editorial review by the Taufolio team before publishing.