BloombergGPT: A Large Language Model for Finance
arxiv.org
arxiv.org
I use ChatGPT to keep track of tasks and Todo lists. It works phenomenally well for me, and the natural language back-and-forth helps keep me motivated. I give it a set of tasks, with time estimates, and it organizes these tasks for me, and I tell it when I complete them, and it updates my task list.
The one funny mistake it makes is that when it groups my tasks (say I have 3 "Work" tasks and 2 "Personal" tasks") it sums up the total estimated time for each task group, but the totals are often wrong, especially when I start adding new tasks or completing tasks.
When so much of finance requires numeric accuracy, I'm curious how BloombergGPT handles numbers.
ChatGPT: Hello hn_throwaway_99! I'm happy to help you as a personal coach to address your problems with procrastination and to help you prioritize tasks. It makes perfect sense to use our conversation as a way to keep you on track and avoid distractions like mindlessly browsing the internet. Feel free to share your current tasks or goals, and I can provide suggestions, encouragement, or strategies to help you stay focused and productive.
Me: Great, thanks very much. I'm going to group my tasks by topic. Note these tasks aren't listed in priority order. For each task I'll give an estimate for how long I think it will take.
Work tasks: 1. Finish Jira ticket foo. Estimate: 2 hours. 2. Write job req for bar. Estimate: 1 hour.
Personal tasks: ...
Home tasks: ...
---
From there I just tell it when I complete tasks or add new tasks, and I ask it "Can you please output my outstanding task list?" - I only had to ask that once, after that it always output my correct task list summary after I told it I added or completed a task. It groups my tasks by the categories I gave it (Work, Personal and Home). I was using GPT-4. I also like how it uses encouraging words and advice as I go through my items.
It would list stuff like:
>Troll: 70pts Total
>Base Cost: 50pts Equipment1: 30pts Equipment2: 10pts
It did write a rather compelling backstory for each character that perfectly fit the lore of the scenario.
"Doing math" is not really a goal of bloombergGPT. Take a look at our applications in the paper, which include information extraction, sentiment, reasoning, knowledge. These are models of language meant for use on text documents.
There are some aspects of the datasets that require numerical reasoning (ConvFinQA), but that's not the same as doing math.
Say, I'd like an ontology of the current stock market focusing on the relationship between natural persons and public companies, board members, well-known analysts, investors and so on. This would be tedious for anyone to do, but should be fairly simple with a LLM.
Another task, maybe a little bit further into the future is categorizing open source intelligence. Think of oryxspioenkop.com and their famous lists of lost equipment in the Russian invasion of Ukraine. It's tedious and time-consuming but generates a valuable dataset. Here, image recognition would be necessary, but the principle is still the same, no?
Come to think of it, how does a company like TomTom generate map data nowadays?
The arxiv paper is written for researchers who are building these models. We benefited tremendously from reading papers on GPT-3, PaLM, Chinchilla, Galactica, Gopher, Bloom, OPT, most of which are closed models. We are contributing back to that community, and the collective experience of people who are training these models. We learned a lot that will help others who are making their own decisions about how to train models.
If you are looking for an API to use, this isn't for you.
Publishing models is definitely better than only a paper, but I think you're being unnecessarily harsh in that case.
We would all prefer it if OpenAI released a proper paper about GPT4, for instance, even if they did not release the complete model alongside it.
(They did benchmark it against GPT-3 in general-purpose tasks and, unsurprisingly, GPT-3 came out on top.)
Press release link: https://www.bloomberg.com/company/press/bloomberggpt-50-bill...
We didn't compare to GPT-4, or any instruction tuned model. We're comparing a causal LM pre-trained only (BloombergGPT) to other causal LM pre-trained only models. You need to compare like to like.
Would an instruction tuned model work better? Sure! I would hope so. That seems like a good next step to explore.
It's also difficult to make scientific comparisons to models that we know nothing about. I don't know what data GPT-4 or GPT-3-5-turbo used for training, how many parameters they have, etc. We can't make any scientific conclusions when comparing two systems if we don't know anything about one of the systems.
We compared to GPT-3 results when they were available. Does GPT-3 do better on general purpose datasets? Yes, as we expected it would. It's a much larger model (~3.5 times the size). No surprise there. If anything, the surprise was how well bloombergGPT did against these larger models, beating them on some datasets.
Also it looks like public filings are a large dataset that isn't currently being used by other LLMs.
1). Finance is high dynamic. BloombergGPT retrains LLM using a mixed dataset of finance and general sources is too much expensive (1.3M hours). Lightweight adaptation is highly favorable.
2). Internet-scale finance data (timely updates using an automatic data curation pipeline) is critical. BloombergGPT has privileged data access and API access. A promising alternative is "democratizing Internet-scale finance data".
3). Another key technology is "RLHF (Reinforcement learning from human feedback)", which is missing in BloombergGPT. RLHF enables learning individual preferences (risk-aversion level, investing habits, personalized robo-advisor, etc.)
LLMs in general, I mean. They seem to be the first widespread application for large, unstructured datasets. Still hype-y, but maybe even a /practical/ application.
Mostly good on words based financial tasks, and it did poorly in tabular numbers tasks.
tldr it's worse than GPT3?