747 karma · joined February 7, 2020
Interestingly, my own calculations lined up pretty well with this calculation, although they approached the problem from a different direction (a leak by Morgan Stanley about how many GPUs OpenAI used to train GPT-4 as well as an estimate of how long it was trained): https://colab.research.google.com/drive/1O99z9b1I5O66bT78r9S...
Sam Altman has also stated that GPT-4 cost more than $100 million to train, and replication can cost 2-4x less compute. https://www.wired.com/story/openai-ceo-sam-altman-the-age-of...
If you know of an organization that can replicate GPT-4 for $400k to $4m I would love to know so that I can invest in them.
"Snapshot of gpt-4 from March 14th 2023. Unlike gpt-4, this model will not receive updates, and will only be supported for a three month period ending on June 14th 2023."
Unfortunately this experiment is using the frozen snapshot model gpt-4-0314 instead of the unfrozen gpt-4 or gpt-4-32k models, so any differences are literally 100% noise. This would be a somewhat interesting experiment if someone were to use an unfrozen model, though. I do appreciate the author for captioning the images with the exact model they used for generation so that this bug could be caught quickly.
1. Someone develops a procedure for training models with distributed computing resources, including consumer CPUs and GPUs. Even this is not really guaranteed to work, since corporations will probably just buy up all the consumer GPUs, since they are cheaper on a FLOPS/$ metric.
2. One or more governments provide a lot of funding for OSS models, including being willing to pay competitive salaries for the best talent (potentially millions of dollars per year). This is unlikely for a lot of reasons. The only thing that could speed up the process enough to compete with private organizations is a major war that required AI to win. In that case, though, open source would be the least of their concern.
3. Scaling laws stop working and Moore's law catches up. In that case adding more compute won't really help and eventually even organizations with small budgets can afford to train a SOTA LLM. We can't know until we find the limit, but we haven't hit the limit of scaling laws so far, despite scaling up massively over the last few years, so I doubt we will find the limit any time soon.
4. A bunch of corporations that have no hope of reaching first place decide to combine their resources to beat OpenAI and thus prevent a monopoly. I'm not sure if there is precedent, but even if there is, that would still require a lot of coordination and resources for little direct monetary gain.
We'll see what happens, but I'm not really confident about any of these possibilities.
My best guess is that it will be a combination of 1 and 2. I have always maintained that LLMs are very unlikely to develop into AGI on their own, but are very likely to be a critical piece of an AGI. The most recent research into scaling laws (Chinchilla, Llama) finds significant improvements from scaling data size much farther than parameter size, so memorizing facts within the parameters will become less and less feasible. This is actually ideal, though, since you want your model parameters to encode language and reasoning patterns, not memorize facts. If it's memorizing facts you either need more data or better (i.e. deduplicated) data. I'm not an expert, though, and I'm too lazy for a research review, so please don't sue me for libel if my facts are out of date.
There's already a lot of research on this, but I strongly believe that eventually the best AIs will consist of LLMs stuck in a while loop that generate a stream of consciousness which will be evaluated by other tools (perhaps other specialized LLMs) that evaluate the thoughts for factual correctness, logical consistency, goal coherence, and more. There may be multiple layers as well, to emulate subconscious, conscious, and external thoughts.
For now though, in order to prompt the machine into emulating a human chess player, we will need to act as the machine's subconscious.
1. Quickly reduce costs by increasing model and computation efficiency.
2. Massively reduce prices while still maintaining some gross margin.
3. Massively increase market size and take the vast majority of market share.
4. End up with a higher gross profit due to a much larger market size despite decreasing prices and gross margins.
5. Profit.
There's also some research on converting hydrogen and CO2 to edible carbohydrates, either chemically or through hydrogenotrophic or methanotrophic bacteria. That will be a huge revolution for either increasing the carrying capacity of the planet or decreasing humanity's impact on the planet. It will also be a huge boon for countries without much arable land to be able to feed their people without relying on imports. Electricity to food is not quite ready for scaling up yet, but synthetic fuels are absolutely ready to go as soon as solar electricity prices drop just a bit more or fossil fuel prices rise a bit more.
[0]http://www.greenrhinoenergy.com/solar/technologies/pv_manufa... [1]https://www.sciencedirect.com/science/article/abs/pii/S00380....
[1]: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6611928/#:~:tex....