89 karma · joined August 16, 2021
Things feel off,but you have (mostly well-off) people talking about how great things are going to be and are. I suspect a lot of this correlates to the stock market.
For the median millenial atleast, whatever you were taught growing up is just the wrong guide to understanding the world today.
I think social media by its very nature, has to be understood inverted. A high number of posts about how great things are suggests otherwise. Comparisons with Europe have gone exponential that tells you more about how the posters feel than anything else...
When you login into capitaliq or factset or look at a bloomberg screen, you access data, you can then do the same here. Excel plugins go on top which can then get this data into excel to build models. The api that powers this app, can also send data into excel for instance.
Data can be copied already directly from the chat box into excel, maintaining the table format for e.g
Regarding standardization, I think data was standardized not to enable comparison but simply to fit into the same schema for every industry. Most companies within the same industry report the same way. REITs do not report like SaaS, but the existing datasets put it all into one set. Raw data from source is always better as you can convert it to standardized but you can't go back...
You can also do this across time with constraints on context limits ...
What is a traditional method, would you search the string or would you find the named entity (using NLP) and look for the entity ? or do you mean a ctrl +F in the document ?
I guess the premise of LLMs, is that we have an intelligence that can read and write, so it can do this in an automated way in different applications. In this case, I am trying to automate and speed up the investment research process, along the way, we can create our own dataset of financial data that can be generated on the fly as necessary.
Also, how would you extract an income statement from the text, also in the same img using a traditional method ? I personally find it magical it can do that, it knows where it ends and so on...
As you can see, inspite of all the training in the real world, gpt4 thinks google cloud is more of an entity than google is, based on that question :)
I guess it really depends on how your targeted user would use it... :)
If we run this query over an api for income statements for all 8k companies, we pretty much have all income, balance sheet, cashflow items, shares outstanding etc. Add stock price data, that can give you EV/Ebitdas, P/Es and all that stuff.
So like a lot of applications, the problem boils down to being able to serve the right text. You have something that can read and do basic inference .... You need to tell it what to read so that it can answer your question. But it can only read 16k tokens (20k words at best). So that's the basic problem. As it's universal, i.e a problem across applications, its going to get better and information will be a lot easier to get access to...
From my experience, GPT4 has been very good in following instructions and doesn't make up numbers which Bard is much more susceptible to. Bard relies heavily on snippets from the web search and completes the rest...
I am pretty sure if google wanted to train it, it would get all the answers right, but the way its designed, it gets one or two numbers and makes the rest up rather hilariously :) and even adds a breakdown which is also made up..
You can take a look here : https://imgur.com/a/vDxOV9D
This is the source: https://www.sec.gov/Archives/edgar/data/2488/000000248823000...
The sum of all of these is 12831.
Total Current Liabilities 7572 Long-term debt, net of current portion 1714 Long-term operating lease liabilities 393 Deferred tax liabilities 1365 Other long-term liabilities 1787 12831
Instead, the mental model is you have an army of people who can read texts really well, as in 'reading comprehension' as they call it in english tests... This army can get you information on the fly.
Investment research involves a lot of back and forth reading and fetching tables and making conclusions, which in turn might not have much to do with stock price performance, but there's a whole industry of financial information and news for that :)
So currently, the scope is to make widely available information beyond what FactSet and CapIQ offer and even that's a long way away :)
https://chat.openai.com/share/c736abf4-ae3c-4fbd-9427-b7d2f9...
I am not making 4 calls with the entire text, I instead get pieces of each which would best match the question. There are additional challenges, for instance, GPT4 struggles to know the difference between the words guidance and outlook, which mean the same thing but somehow they don't for GPT4.
When I say they become better readers, I meant it in a general sense as in better in the case above. You basically have someone who can read through tables really well, and that can change investment research fundamentally, which is a lot of reading tables and graphs :)