I'm curious how you're accurately extracting the data though. Are you prompting to respond in a JSON format, using OpenAI's functions or something else? How do you ensure you have the correct label, dates, values, etc?
I'm curious how you're accurately extracting the data though. Are you prompting to respond in a JSON format, using OpenAI's functions or something else? How do you ensure you have the correct label, dates, values, etc?
In regards to extracting the correct data, is that done through the function definitions? You specify all the fields you want to extract in the function and then let GPT go ham?
E.g. operating income, interest income, interest expense, etc.
So like a lot of applications, the problem boils down to being able to serve the right text. You have something that can read and do basic inference .... You need to tell it what to read so that it can answer your question. But it can only read 16k tokens (20k words at best). So that's the basic problem. As it's universal, i.e a problem across applications, its going to get better and information will be a lot easier to get access to...