AI Stock Analysis: Check One Number Against the Filing
AI stock analysis sounds certain. Ask a chatbot for last quarter's revenue and margin, then open the filing and find the exact line the number came from.
See a sample reportThis is a method, not a recommendation. Nothing here, or anywhere else on Taufolio, is investment advice. Treat every example as a starting point for your own research.
AI stock analysis works. Just not the kind you talk to in a chat window. Ask a chatbot for a company's revenue and margin last quarter and you get a fluent sentence with no date and no line reference behind it. Even when the number is right, you cannot tell whether the model read it or reconstructed it. Which is why the only question worth asking an AI tool about a company is: how do you know.
Where AI stock analysis actually wins
Skepticism about machines in markets does badly in the data. Ed deHaan and Suzie Noh at Stanford, with Chanseok Lee and Miao Liu, turned a model loose on 3,300 US equity mutual funds and had it rebalance their portfolios every quarter for thirty years, from 1990 to 2020.
Over that stretch the human managers produced $2.8 million of quarterly alpha and the model produced $17.1 million, six times as much out of the same pile of public numbers.
You cannot wave that study away, because the model got no information advantage. It worked from 170 variables anyone could have pulled at the time: interest rates, credit ratings, the tone of earnings calls. It also swapped out roughly half the portfolio each quarter. Out of public numbers, it got more than the people who had those same numbers on their desks.
So if someone tells you AI has no business near company research, the evidence is against them. Does that settle the question of whether ChatGPT can analyze stocks?
It sounds right. Except that model never wrote a sentence: it read a table.
That distinction decides everything that follows. The model took 170 columns of numbers in and put one position-weight decision out, so every claim it made had an address in the data. A chatbot takes a sentence in and puts a sentence out, and nowhere in between is there a step where it has to open anything.
The two-number test
The gap between that machine and a chatbot shows up fastest on one company and two numbers. Ask a model to “analyze this company for me” and you get a page of prose you cannot verify at any single point in reasonable time. Two numbers that sit in named lines of a filing, you can check.
- Pick a company you actually hold.
- Ask a chatbot for revenue and margin for the last reported quarter.
- Open that company's filing and find the two lines; US filings are free on EDGAR.
- Ask the chatbot which line the figure came from, and as of what date.
- Write down the date of the test and repeat it after the next results.
Step four carries the whole thing, because step two almost always looks fine. The model writes fluently, in the present tense, without a flicker of hesitation, so the answer reads like stock analysis long before anyone establishes whether it is.
Search turned on does not settle it either: a link to a results release is still not a line and still not a day.
Step five is there for a different reason. One pass proves nothing: a model can be right by accident, and you can happen to pick a company the internet has written about exhaustively. Only the repeat, same company, next quarter, shows whether the tool went and got a new document or simply refreshed its tone.
What the filing says
I opened the filing of the company everyone asks about. In the quarter ended December 27, 2025, Apple reported $143.8 billion of revenue against $124.3 billion a year earlier, up 16%.
Gross margin for that same quarter was $69.2 billion, or 48.2% of revenue, against 46.9% a year before. Which means the company keeps about 1.3 cents more from every dollar of sales than it did a year ago, before it pays for research, selling and tax.
Up to here it all looks like one even quarter. The table says otherwise.
Products grew 16%, services grew 14%, and the iPhone alone grew 23%, from $69.1 billion to $85.3 billion. Greater China contributed $25.5 billion against $18.5 billion a year earlier, up 38% and more than twice as fast as the company as a whole. A quarter that reads as steady growth in one sentence is, in the table, a quarter about one product and one region.
Gross margin did not rise for the reason usually cited here either. Mix did not push it up, because the services share of revenue actually slipped a little. Both sides improved on their own: products from 39.3% to 40.7%, services from 75.0% to 76.5%. Those are two different stories, one about the cost of making things and one about scale, and neither of them fits inside “margin went up”.
Further down the same table it gets better. Research and development spending rose 32%, twice as fast as revenue, and the operating margin still climbed from 34.5% to 35.4%. The company put far more money into development than it gained in sales and still came out of the quarter more profitable at the operating line.
That is the sentence a chatbot will not write, because writing it takes four lines in front of you at once. A memory of one will not do.
All of these numbers sit in a table with a heading, a date and a line name. You can point at them.
Which margin, given there are three
That one quarter has three margins and every one of them is true. Gross, 48.2%. Operating, 35.4%, because $50.9 billion of operating income divides into the same $143.8 billion of revenue. Net, 29.3%, because $42.1 billion is what survives tax.
Now say the chatbot answers: margin is around 46%.
That answer is neither right nor wrong. It is uncheckable, because you cannot tell which line it describes or which quarter it means. Between 48.2 and 29.3 sit two different questions about the business. The first asks what the company earns on making the product. The second asks what is left after research, selling, interest and tax.
Put an answer like that into your notes on a company and it will live there for years. Nobody revisits a number that got through unchallenged once, and that is the worst property of a wrong figure: it does not hurt on the day it is created. It hurts two years later, when you build a peer comparison on it and no longer remember where it came from.
I think that single question, which line, is what separates a research tool from a conversation tool. The rest of the differences are cosmetic.
Why the model would rather guess than stay quiet
The explanation comes from the people who build these models. A paper on hallucination, written mostly by OpenAI researchers, puts it this way: models are trained and graded like a student sitting a multiple-choice exam, where “I do not know” scores zero and a guess sometimes scores a point, so guessing pays. Nobody designed this maliciously; it is just how the score is counted.
The consequence is that an invented margin comes out of the same machine, in the same register, with the same confidence as a real one. In prose you can catch it, because text that has drifted from its source starts sounding general. In a number you can catch nothing, because one altered digit reads exactly like a good one and looks precisely like what the reader expected.
You cannot prompt this away. Asking the model to “be careful with the numbers” changes the tone of the answer, not the incentive that produced it.
The quarter nobody named
The second thing a chatbot will not volunteer is a date. The filing carries the heading “three months ended December 27, 2025” above the column, and that heading exists so nobody has to guess what is being described. A chatbot writes “last quarter” and means the last quarter it knows anything about.
Which quarter is that? You cannot tell from the answer, because to the model the date of the event and the date of its own knowledge are the same thing.
Sometimes it is a quarter from a year ago, with a chief executive who has since resigned and a segment the company has since sold. Repeated guidance is the worst of it, withdrawn on the following earnings call and still landing in the answer as knowledge.
The nastier version happens when the model has half the information about a recent event. It fills the rest in with whatever is likely and writes it in the present tense, as though it had just been holding the primary source. Three sentences later you cannot tell what it read from what it added.
So ask for the date first and the number second. “As of the end of last quarter” means nothing until someone names the day that quarter ended.
Where a chatbot genuinely earns its place
Does that mean a chatbot is useless to an investor? No. Two jobs it does very well, and both share one property: the source is already in the window.
The first is explaining concepts. What separates gross margin from operating margin, and what exactly sits between them, has a general answer that does not depend on a quarter or a company. Here the model beats a textbook on speed and rarely slips, because it has nothing to retrieve and nothing to date. The same goes for how to read a cash-flow statement, or why net income and operating cash can drift apart for several quarters running.
The second is working on text you paste in yourself. Paste an excerpt from a filing or a stretch of an earnings-call transcript and ask it to list the assumptions, restate a paragraph, find the sentence that contradicts the one before it. The model has nothing to guess, because it was handed the material, and you can hold its sentence against the paragraph beside it.
That alone saves an hour on a document you were never going to finish.
There is a third use, and to me it is the most underrated: sparring against your own thesis. The prompt runs: here are my three assumptions, list what would have to happen for each of them to break. Asking a model whether you are right comes back as agreement. I wrote separately about how to run that conversation so it produces counterarguments instead.
Where this rule breaks down
A citation is not a guarantee. A link can point at the right document and the wrong line, and a write-up with a footnote on every claim can mislead exactly as badly as a chatbot while looking far more serious.
So a citation is worth something only when you can open it and check it in under a minute, not when it sits decoratively at the end of a sentence.
There is a cost to that rule. A tool that goes quiet when the source is missing can be irritating to read, because it leaves blank spots where a chatbot would have supplied a smooth paragraph. That blank is information, and nobody enjoys it: the company did not disclose the thing.
Honestly, I do not know how much of this the next generation of models fixes. The incentive that paper describes is fixable, since it lives in how answers are graded rather than in the silicon. What I do know is that the test in this post takes a few minutes, costs nothing, and can be repeated after every earnings season on the same company with the date written down.
The rule I have kept for years: a number without a line is a quotation from nowhere.
That is how the Full report in Taufolio is built. Every factual claim points back to a document, every management remark points to the Transcript with a name and a place in the record, and whatever the source does not say stays unwritten. The pro version is written by Three models and a judge: three models work the same company independently, and a fourth compares their versions and keeps only what a source will carry. What one of those documents looks like end to end is on the sample reports page.
Frequently asked questions
Can ChatGPT analyze stocks?
Why does ChatGPT make up numbers from filings?
Is ChatGPT's company data up to date?
What is a general chatbot actually good for in investing?
How is source-cited AI research different from a chatbot conversation?
- research
- sources
- methodology
Posts are produced with AI tools and go through editorial review by the Taufolio team before publishing.