ML training compute has been doubling every 6 months since 2010
twitter.com
twitter.com
I think the advent of retrieval models (retrieval transformers) will continue this compute trend in a more efficient manner. They allow focusing of the compute onto indexing of the knowledge.
For datasets benchmarking new task adaptation there are Meta-Dataset (MD)[2] for Meta-Learning, and the Visual Task Adaptation Benchmark (VTAB)[3] for Representation Learning. Recently to compare the two approaches VTAB+MD[4] was created. Of course, there is also model quantization[5], mixed precision, sparse networks[6], brain floats[7].
Well I've dropped a lot there, but there is a lot more outside the mainstream. I've been building something on a tight budget for a few years so this has been important for success there. We are are very focused on being data and compute efficient.
[1]: https://proceedings.mlr.press/v70/finn17a/finn17a.pdf
[2]: https://arxiv.org/abs/1903.03096
[3]: https://arxiv.org/abs/1910.04867
[4]: https://openreview.net/pdf?id=Q0hm0_G1mpH
[5]: https://arxiv.org/abs/2105.08819
[6]: https://arxiv.org/abs/2112.13896
[7]: https://en.wikipedia.org/wiki/Bfloat16_floating-point_format
One thing I'd object to is coining (or maybe just using?) the term "Large-Scale Era". This seems like polluting the discussion with a meaningless phrase.
The basic message is essentially, deep learning continues to grow in scale exponentially, many people consider the dividends still worth the costs. No need to "sexy things up" with what no more than a marketing term.
I think the Large-Scale Era does point to a new phenomenon that emerged pretty discontinuously, which is that there are now 'two lanes' in ML scaling. Prior to 2015, academic and industry would train roughly similarly compute intensive models. Since then, a small number of industry players frequently train models with 10-100x more compute than what the typical researcher uses.
Something about environmental problems causes people to focus on vary narrow types of harm rather than the major ones like habitat destruction. Domestic cats kill ~5 orders of magnitude more birds than windmills, yet somehow bird strikes of windmills is in the public continuous. I think it comes down to the inability to really grasp the difference between large numbers combined with specific harms used as distractions.
We use multiple GPUs at 100% capacity (all cores) for ML training.
When need be, we use parallelization/concurrency to use every core in computer for our application.
Almost no computer application today buffers while processing (thanks to parallelization / concurrency)
Software can improve a lot faster than hardware can (due to low capital required)
This hypothetical GPT-4 would be big enough that it shouldn't have the context window or BPE token problems of GPT-3. It would be superhumanly good at predicting token sequences-- generating text. What that would look like, exactly, is not entirely clear. (What does it mean to be twice as good as a human at writing an essay?) If it had the same architecture as previous GPTs then it wouldn't be "conscious" or be goal directed, but it would be able to correlate more information and draw conclusions that we couldn't. (Would we understand these conclusions is another open question.)
It would also be quite expensive to run in a capital-amortized sense, at first, maybe hundreds of dollars per word. The commercial use of such a thing would be limited. Large tech companies seem to be building AI because of how self-evidently useful it will be... eventually.
The economic case for high-cost NNs is the opposite of most automation, which started from the bottom up. If a net is expensive to train and run then you have to pursue the Tesla strategy-- start from the top down. So you need to target high-price knowledge work in terms of dollars per hour, but is tolerant of small variance, since the results are still probabilistic. In the near future you won't see NNs designing jet engines, producing complete software systems from requirements documents, or even standard grunt-level software engineering work of wiring two systems together, since you can never be quite sure GPT-programmer is producing the right work.
Similarly, GPT-lawyer or GPT-CEO would be tough to do. GPT-hollywood would be interesting, if you could get it to crank out a complete Marvel movie with superhumanly good CGI. More terrifyingly, GPT-advertiser, if it can produce super-appealing ads, would be able to pay for its own runtime at the cost of a machine hijacking human minds to plug money into gacha games or the equivalent in the 2030s.
Terminator, Terminator 2
2011 1e+14 = 100 000 000 000 000
2021 1e+21 = 1 000 000 000 000 000 000 000 <- you are here
2031 1e+28 = 10 000 000 000 000 000 000 000 000 000
I've been working with MNIST with C and CUDA with dynamic parallelism for a couple months and it's been extremely enlightening but I'm kind of ready to move on.
CNNs maybe?
Why try to beat nature?
I'm being only semi sarcastic here.
The vast majority of AI research (over all time) has not been seriously intended or expected to arrive at a general intelligence.
Sure there are some true believers there in some sort of handwavy construction argument ("and then a miracle occurs") and there have been serious research attempts to understand intelligence, of limited scope an resources. But to be clear, whether this wave or previous ones, the money washes in and the army of PhD's happens when there is a sniff of practical, usable algorithmic results from machine learning. It's not the same thing.
Parallel application of even relatively stupid algorithms targeted at the right tasks will rapidly outproduce any reasonable number of organic intelligences of the type you describe :)
Or maybe your newborn is more capable than mine!
"Why try to beat nature?"
If we wouldn't try "beat nature" we would be still in caves or would be extinct. Nature is "beating" itself with evolution and we are part of that.
oscar wilde's the soul of man under socialism probably better describes the idea.
which I thought was brilliant, ha!
2. Because ai doesn't have biological scaling limits. There's nothing particularly special about human level of intelligence.
Big claim since nothing in the known universe has human level intelligence besides humans. It's like holding the original declaration of independence and saying that there is nothing particularly special about this piece of paper.
What do you propose is special about human-level intelligence? If anything, it is in constant struggle with biological and emotional needs. Yes, of course, human brains are amazing, the declaration of independence is great, but the entire human project is the push for progress.
IMHO in the community there aren't any significant objections to the Legg & Hutter definition proposed in that document "Intelligence measures an agent’s ability to achieve goals in a wide range of environments." - perhaps you can tweak the wording some more, but that's the direction implied by people talking about building general intelligence and (at some future point) human-comparable general intelligence; in essence it's about 'wide-coverage' intelligence of being efficient at different purposes and figuring out what is required to be successful for these purposes, as opposed to narrow single-task specific effectiveness; but it's essentially a metric on which one quite reasonably can imagine something being equivalent to or better than the average (or x-th percentile) human.
It's not about the capabilities of unassisted body - I can influence stuff in a volcano without sticking my bare hands in it.
- Is that 2 (as in binary 10)?
- 108 as in a typo, or
- an order of magnitude of 6 thereof?
However, I've replaced the submitted title ('ML training compute has been growing by a factor of 10B since 2010') with what the tweet says now.
edit: s/parameters/FLOPs