Reading SEC filings using LLMs
beatandraise.com
beatandraise.com
A core assumption underlying this is that LLMs are already pretty good and will continue to get better at reading texts. If provided with the right thing to read, they will do very well on 'reading comprehension'.
Open ended writing is more susceptible to errors, especially in questions related to finance. For e.g google's revenues are just as likely to be 280.2 billion vs 279 billion in a probabilistic model that guesses the next part of the sentence - Google's revenues for FY 2022 are ....
So this leaves us with the main problem to solve; Serving the right texts to the LLM aka text search.
Once the right text is served, we can generate any pretty much anything in the text, Income statements, ceo comments, accounts payable on the fly, For e.g try - `can you get me Nvidia and AMD's income statement from March 2020 ?` as in here. https://imgur.com/gallery/H8Vfd5X A few more examples, Apple's sales in China, Google's revenue by quarter: https://imgur.com/a/oCCay3o
Currently, the application supports ~8k companies that are registered with the SEC. Pdfs are still work in progress, so tesla etc don't work as well.
The stack is Nextjs on Supabase. So Postgres's inbuilt text search does a lot of heavy lifting.
If one thinks of the bigger picture, we can extend/improve this to pdfs and the entire universe of stocks and more. a.k.a a big component of what CapitalIQ, Factset, Bloomberg and Reuters do can now be generated on the fly accurately for a fraction of the cost.
Generating graphs with gross margin increasing etc are just one step further and stuff like EV/Ebitda, yet another step further, as one can call a stock pricing api for each date of the report.
I would guess a number of LLM applications follow a similar process, ask a question --> LLM converts to query --> datalakes/bases --> searching and serving texts --> answer. Goes without saying, I would appreciate any feedback, especially from those who are building stuff that looks architecturally similar :) !
The LLMs are able to hone in on the details and provide interesting responses like "What questions were asked of the minister that they failed to address" and "What should the opposition leader have mentioned in their response that the minister would have found difficult to answer"
Crucially, they're well transcribed: https://api.oireachtas.ie/
For e.g, if the question is: "What should the opposition leader have mentioned in their response that the minister would have found difficult to answer."
An embedding based search will find it fairly difficult to match this against a text. Based on my experience, you have to figure out what is a nonanswer first (i don't think that's easy but gpt4 is very good at a lot of language based stuff.)
you can try Q1: question A1: Answer then prompt GPT, do you think A1 answers the question and then save it. And then, Q1: question, A1: Answer, Q2: Follow up based on the following questions, do you think A1 answers Q1 and then save it to a db.
You can then augment it in the code with our own knowledge of how politicians lie, using certain words etc :) to improve what gpt4 might miss...
https://chat.openai.com/share/c736abf4-ae3c-4fbd-9427-b7d2f9...
I was also thinking about a society of bots type application where you could have autonomous bot researchers, debaters, judges and audience. Would be interesting to feed in the topics and grab some popcorn
Leaving aside the bots that are trained to answer "as an LLM I do not have opinions..." it's going to be a very basic probabilistic yes or no based on tone and numbers of pros and cons rather than knowledge of the surrounding political context and higher order reasoning about the accuracy of the claimed pros and cons.
When you say that generating graphs of gross margin, EV/Ebitda is just one step further, are you talking about generating those based in just a single report's information or are you combining information from multiple years to for example show gross margin trends over 10 years and EV/TTM Ebitda?
You can think of it like this, you now have an army of readers that can go through tables really quickly.
FactSet, CapIQ etc use a combination of automation/manual entry and fit these tables into a homogenized schema so that they can be saved, compared etc. So if you want to get Apple's Greater China sales from 2020, you would be lucky if they decided to create an item for that. https://imgur.com/a/bp2hb7n
Here are two examples, AMD's revenues and AMD's revenue outlook using beatandraise.com https://imgur.com/a/61jqiUk I doubt you can get AMD's own outlook on CapIQ for e.g
Context sizes mean getting 100s of reports on one call is not possible, but multiple iterations will still do the trick. So in effect, you can actually create a dataset like FactSet for a lot lower cost, more comprehensive and can be customized to what the user wants, if you see my point... :)
If we run this query over an api for income statements for all 8k companies, we pretty much have all income, balance sheet, cashflow items, shares outstanding etc. Add stock price data, that can give you EV/Ebitdas, P/Es and all that stuff.
How do you parse the PDFs?
I used poppler on a digital ocean droplet, but the sheer variety of company pdfs especially european companies, some of which have to be OCRed, meant results were not really uniform. GPT still does very well, but not as well as on text documents directly. So in short, this is next on the list...
Avoiding any fears of hallucinations.
But earlier you say "LLMs will get better at comprehension". So are you using an LLM to markup the original text in some way ?
I am not making 4 calls with the entire text, I instead get pieces of each which would best match the question. There are additional challenges, for instance, GPT4 struggles to know the difference between the words guidance and outlook, which mean the same thing but somehow they don't for GPT4.
When I say they become better readers, I meant it in a general sense as in better in the case above. You basically have someone who can read through tables really well, and that can change investment research fundamentally, which is a lot of reading tables and graphs :)
What ChatGPT completey missed was stuff like omissions (I'll kind of give it a pass here, how can software analyse the absence of something without having access to supplemental documents; it still shoes hoe dangerous it is to rely only on a LLM for such analysis) and, more importantly, the connection between certain tid bits and, and here it became outright dangerous, ChatGPT didn't provide anything meaningful on risks and financials.
The tid bit it missed, one of the most important ones at the time, was a huge multi year contract given to a large investor in said company. To find it, including the honestly hilarious amount, one had to connect the disclosure of not specified contract to a named investor, the specifics of said contract (not mentioning the investor by name), the amount stated in some finacial statement from the document and, here obviously ChatGPT failed completely, knowledge of what said investor (a pretty (in)-famous company) specialized in. ChatGPT did even mention a single of those data points. Fun fact, said contract covered a significant junk of the amount the investing company had invested to begin with. And all that during a time in which the financial stability of the reporting company was at least questionable. Oh, and ChatGPT didn't even realize that risk (cash and equivilants on hand devided by burn rate per year is simple maths), or repeat the exact passage in which the SEC filling said that the survival of the reporting was in doubt.
In short, without some serious promp working, and including addditional data sources, I think ChatGPT is utterly useless in analyzing SEC filings, even worse it can be outright misleading. Not that SEC filings are increadibly hard to read, some basic financial knowledge and someone pointing out the highlights, based on a basic unserstanding of how those filongs actually work are supossed to work, and you are there.
Instead, the mental model is you have an army of people who can read texts really well, as in 'reading comprehension' as they call it in english tests... This army can get you information on the fly.
Investment research involves a lot of back and forth reading and fetching tables and making conclusions, which in turn might not have much to do with stock price performance, but there's a whole industry of financial information and news for that :)
So currently, the scope is to make widely available information beyond what FactSet and CapIQ offer and even that's a long way away :)
I'm curious how you're accurately extracting the data though. Are you prompting to respond in a JSON format, using OpenAI's functions or something else? How do you ensure you have the correct label, dates, values, etc?
In regards to extracting the correct data, is that done through the function definitions? You specify all the fields you want to extract in the function and then let GPT go ham?
E.g. operating income, interest income, interest expense, etc.
So like a lot of applications, the problem boils down to being able to serve the right text. You have something that can read and do basic inference .... You need to tell it what to read so that it can answer your question. But it can only read 16k tokens (20k words at best). So that's the basic problem. As it's universal, i.e a problem across applications, its going to get better and information will be a lot easier to get access to...
I wouldn’t yet feel comfortable with this without some automated reconciliation which to my mind defeated the point of my hobby project but I’m curious if you’ve seen different? No doubt you’d expect this to improve over time though.
This is the source: https://www.sec.gov/Archives/edgar/data/2488/000000248823000...
The sum of all of these is 12831.
Total Current Liabilities 7572 Long-term debt, net of current portion 1714 Long-term operating lease liabilities 393 Deferred tax liabilities 1365 Other long-term liabilities 1787 12831
I don't see anything in your privacy policy that gives me comfort my (even anonymized) interest in a particular firm isn't feeding your own signals.
Even if your policy promised it, I'd want to see technical controls implemented, since incentive to leverage information on searching would be so high.
People search for a company for a reason, generally to inform a thesis towards buying or selling, and regardless of their thesis, there will generally be a price move when they act.
> And these are searches on publicly available data, ie data that is not proprietary and already filed with the sec.
Of course. But the information that someone is interested is not public, so you are crowdsourcing indication of interest.
Effectively you have a leading indicator (however soft) for order flow.
> Needless to say, none of the chat traffic is used or will ever be used to feed any trading signals for anyone.
Point is, it's not "needless" to say, it's "needful". And something other than "trust me" would be in order.
As you can see, inspite of all the training in the real world, gpt4 thinks google cloud is more of an entity than google is, based on that question :)
I doubt any finance professional would want to move from Excel to ChatGPT for these use cases.
A secondary issue is that numbers in financial reports need to be standardized before they can be used for analysis - say comparing ROE of a stock with industry average. That’s not possible with numbers from raw reports.
When you login into capitaliq or factset or look at a bloomberg screen, you access data, you can then do the same here. Excel plugins go on top which can then get this data into excel to build models. The api that powers this app, can also send data into excel for instance.
Data can be copied already directly from the chat box into excel, maintaining the table format for e.g
Regarding standardization, I think data was standardized not to enable comparison but simply to fit into the same schema for every industry. Most companies within the same industry report the same way. REITs do not report like SaaS, but the existing datasets put it all into one set. Raw data from source is always better as you can convert it to standardized but you can't go back...
Retrieving numbers via ChatGPT and then feeding it into Excel via APIs seems like a very odd thing to do. Especially when most of these numbers are already available on Yahoo finance for individual investors and CapitalIQ for professionals.
Check this out: https://cofinapp.com Let me know what you think!
I guess it really depends on how your targeted user would use it... :)
1st question = got the ticker wrong. asked about company A - gave info on B 2nd question = please give info on A. it says sure 3rd question = please give info on A. pops a modal asking to pay
messed with yours and had a good experience = false alarm
We also just dropped our price to $1.99 per month from the previous $49.99.
Enjoy.
This feature stayed on the site for a couple months but ultimately we mutually decided to remove it since the traction didn't meet the costs. I'm open to revisiting this topic all in the name of making verbose financial disclosures easier to understand for all of us.
Sorry if I missed something, but skimming through the site I didn't find any information about this specific issue.
From my experience, GPT4 has been very good in following instructions and doesn't make up numbers which Bard is much more susceptible to. Bard relies heavily on snippets from the web search and completes the rest...
I am pretty sure if google wanted to train it, it would get all the answers right, but the way its designed, it gets one or two numbers and makes the rest up rather hilariously :) and even adds a breakdown which is also made up..
You can take a look here : https://imgur.com/a/vDxOV9D
You basically monitor the SEC website for when filings are made public, quickly parse the document in some automated way and make trades based on the output.
It wouldn't surprise me if some hedge fund has been trying to build something with LLMs for past few years (and possibly already deployed it).
I think using LLMs to ask a company about their company figures is missing the point. After all, XBRL already exists, and that data is widely accessible anyways. Where LLMs would add more value is analyzing changes in different sections of a company's filings, such as Risk Factors, MD&A, or governance filings.
I’ve pulled out the structured data separately, and now looking at breaking up the text into paragraphs to get embeddings.
Why are you using inverted index style text search instead of embeddings? I can see doing both perhaps…
What kind of data can be extracted from these filings, that is not available using "traditional" methods?
Nothing wrong with that :) information retrieval is hard. And CTRL+F won't render a table for queries like "what is the % increase of revenue over the last 3 quarter"-type queries.
You can also do this across time with constraints on context limits ...
What is a traditional method, would you search the string or would you find the named entity (using NLP) and look for the entity ? or do you mean a ctrl +F in the document ?
I guess the premise of LLMs, is that we have an intelligence that can read and write, so it can do this in an automated way in different applications. In this case, I am trying to automate and speed up the investment research process, along the way, we can create our own dataset of financial data that can be generated on the fly as necessary.
Also, how would you extract an income statement from the text, also in the same img using a traditional method ? I personally find it magical it can do that, it knows where it ends and so on...
IME, the most useful software tools in investing are analogous to the most useful software tools in software-building: the IDEs, libraries and frameworks that make an increasingly-expert human more effective.
Unfortunately, there aren't any long-term-valuable shortcuts to attaining expertise in any domain, but there are an increasing number dead ends that really look like shortcuts.
We found the hardest part about this is the pre-ranking of sections and finding the relevant information. Do you feed the whole report into OpenAI?