Deep Neural Nets: 33 years ago and 33 years from now (2022)
karpathy.github.io
karpathy.github.io
The original training took 3 days on a Sun 4/260 workstation; I can't find specifics but I believe that era of early SPARC workstations would likely pull about 200 watts in total (the CPU wasn't super high powered but the whole system, running with the disks and the monitor etc would pull about that).
So 200 watts * 72 hours = 14400 watt-hours of energy.
Karpathy trained the equivalent on a Macbook, not even fully utilized, in 90 seconds. Likely something around 20 watts * 0.025 hours = 0.5 watt-hours.
An energy efficiency improvement of nearly 30000x.
By any measure that puts energy used by the brain in the denominator, humans are probably dumber than ants. But that doesn't mean those measures are always accurate.
(For contemporary neural networks, you also have to distinguish training costs from inference costs.)
The greatest form of general intelligence at 20W.
A MacBook Air is ~30W.
https://www.jackery.com/blogs/knowledge/how-many-watts-a-lap...
Given a laptop is at 30W how much can a laptop do disconnected from the internet? Now how much can it do with the internet? Now how much information does the internet cost in terms of wattage? Now what’s the ratio?
Doubling every 2 years = compound annual growth rate (CAGR) of ~41.42%
CAGR = ((End Value / Start Value)^(1 / Number of Years)) - 1
((2 / 1)^(1 / 2)) - 1 = 0.41421356237
Therefore in 34 years since then: 1 * (1 + 0.41421356237)^34 = ~131,072
So x30k is ~4.4x less than 131k. Then again, that's equivalent to ~x1.833 every two years, compared to Moore's Law of x2 every two years, so only ~8% less growth per two years, which coming back to the fact that Moore's Law is a rough estimate concept not an exact fact, doesn't seem to far off!Modern laptops give great efficiency, when I went for solar power here the first thing to go was the desktop computer. I still have it, but it hasn't run in over a year and the elderly thinkpad that is now my daily driver uses far less power and still has enough compute to serve my modest needs. But if I would dive into something requiring much more compute I'd have to start the desktop again. Unfortunately power management is not such that computers can really throttle down to 'miser mode' when you don't need it, it's a good step but not as good as the jump between desktop and laptop.
Imagine a GPU with a 128G or even 256G slot based memory section that is sold unpopulated. 8 SODIMM slots or so.
8% less growth is not the point. The "law" has stood the test of time which says something about the guy and his vision
You mean joules (up to a constant factor)?
It's quite possible that none of these predictions will come true simply because the timeframe is large enough for many unanticipated breakthroughs and roadblocks.
Maybe someone will figure out a much, much simpler foundational architecture than "perceptrons++", maybe we'll all be training clouds of 3D gaussians, maybe quantum computers will finally take off and we don't even have the nouns for the building blocks we'll use.
On the negative side perhaps we hit a hard scaling limit (in hardware or training) that we didn't see coming. Or a civilizational setback.
All that said, though, if I were a betting man I wouldn't exactly wager against the article's conclusions; they're probably the best we can extrapolate knowing only the past and present state of affairs.
I would lean to them being even more dramatic, due to the opportunity to advance algorithms, not just resources.
On the more obvious side, most libraries are not yet taking full advantage of many known gradient optimization techniques. It’s been so much easier to just add data & processing that there is an overhangs of tools to still apply.
And large successful models are telling us important things.
For instance, it is clear that language models are learning a kind of logic of language similar to how we process thoughts, allowing highly disparate types information to be woven together sensibly.
At some point, identifying the nature of that processing could radically simplify language processing.
That is just one opportunity for radical architecture and algorithm advances, and it would be revolutionary.
Remember, 33 years ago "connectionist AI" wasn't the dominant AI paradigm, and "symbolic AI" wasn't the only other approach either - there were others, like "robotic functionalism" (the idea that you couldn't have true intelligence with interacting with the physical world). Maybe in 33 years some of these other approaches will have a resurgence, perhaps in combination with connectionist approaches. Or maybe they'll even be some entirely new approach.
My world has been very exciting in the last 18 months. I spend as much time as I can exploring self hosted LLMs, APIs from Hugging Face, OpenAI, etc.
My mind is blown even thinking about tech 33 years from now!
Little images of characters is a trivia type problem, very different from training on the linguistic and visual communication of essentially the whole human race.
Another 33 years of expanded computing resources won’t be training models to mimic the behavior and knowledge of humanity.
That problem (us!) will have been reduced to a toy problem long before then.
Model architecture doesn't matter compared to the dataset. Any model from a class can learn the same skills from the same data, but change the data and they all change their abilities - the intelligence is in the data.
The future is data engineering, not model architecturing. Human culture, by analogy, evolves faster than human biology. The data is evolving faster than the model. And we are seeing a drastic reduction in novel architectures in AI, diverse datasets applied to the same transformer models in recent years. Even among the transformers, very few variants are largely used, thousands of them abandoned.
I like to think of it as language evolution by memetics being the real engine behind intelligence. We and AI are riding the language exponential together.
You might be right in the same sense that big-O notation is 'right'. Constant factor can matter; especially once you have to take energy use into account.
My Tesla drives and navigates itself most of the time. 90-95% at least, just not 100%.
As apposed to cars 10 or more years ago which didn’t do any of that.
To me it is much like the “God of the Gaps” when tremendous progress on a big problem is dismissed negatively, due to the (continuously shrinking) gaps of what it can’t do.
But I think all these complementary takes on the problem, with significant year-to-year progress by all three firms, are fantastic.
That used to be considered a fast learning curve!
And not delivering AGI yet is a problem?
What are these broad technology schedule based criticisms founded on?
I really want to understand this viewpoint!
Hopefully not the over-optimism of anyone who uses optimistic timelines as a motivational force. That’s not real data. Or a suitable benchmark for human progress.
I read the article and I was thinking "my God, I remember I used MSE that weekend in my pet ML project and it really didn't work out that well; wrong loss function." Our current crop of LLMs, or the one next year, will be perfectly able to tell me how I can improve my code and graphs, which means that I can deploy some expert-level techniques that otherwise would be "locked" to me by 50000 hours of "mastery acquisition".
A part of me is telling me that we humans are doomed, and that in 33 years we would have created a world in which we humans are irrelevant. But another part tells me that if we avoid that fate and all the other dooms, the future might just be quite bright.
We have heard, and will continue to hear, this sort of thing rather a lot. The last 5 yards are the hardest, but without them the previous 5 miles are of limited utility.
And I think that we have a responsibility to any intelligent being we bring into the world. Some lament that there’s no test to become a parent-what about creating a million copies of a new virtual brain from scratch? And basically so they can be born into lifelong servitude.
The lambda calculus (a system we know is capable of infinite self-complexity, learning, etc. with the right program) can be described in a few hundred bits. And a neural net can be described in the lambda calculus in perhaps a few thousand bits.
Also, we have no idea how "compressed" the genome is.
Compression still must obey information theory.
Not sure all the things it captures, but a model could be trained on the combination of audio/video/spatial/iris/what have you...
Now though, I’m sure people are taking LLMs and putting them together to do forward and backward chaining.
But a Turing award is pretty neat as well.
The new stuff is better, by a lot, and with implications more to come.
But those of us paying attention then had a frame of reference where “so much better it’s crazy” still stops short of “it’s out of control”.
It’s a lot better.
This is exactly why I didn't start experimenting with this stuff back then. I read some articles and had the interest, but having no access to existing training data or "fast" computers was really a show stopper. This article really convinced me that the amazing results today are mostly due to hardware advances.
I will add my own view that 1) hardware will not be advancing anywhere near so much in the future. And 2) training and inference have to be done together like real brains do. Then the AI will learn from experience while deployed and you can clone the best ones later.
For LLMs that is true. But many other things like Whisper, Stable Diffusion etc. could in theory have been made a decade earlier.
Fine-tuning would then be used for specific environments (human hair, an apartment) and robots (robotic barber, cleaning robot).
Imagining I could take a foundation model, run it on my specialized task, and measure which regions of the neural network light up.
Then I could carve out unused regions of the network to create a more lightweight model.
Or is this a silly idea?
LeCun is not even really interested in supervised learning anymore, for example.
https://youtu.be/vyqXLJsmsrk?si=8n0ylC6qdLX06CmY
Note that the talk is not really primarily about ChatGPT even though that's in the title. The new ideas are a little bit in. The beginning of the talk is just him explaining how unimpressed he is with LLMs. Which I think is a misjudgement but that doesn't mean his plan doesn't have merit.
Evolution has optimized animal brain and bodies to survive and take care of the next generation. They have a good grasp of environment, where they are, where food is, where predators are, basic communication if they live as a group. Babies grow up and start learning.
Our current AI is extremely power hungry compared to a brain. Cruise & Waymo put large power hungry supercomputers in cars. The computing system costs 100k+. They still make silly mistakes like crashing into fire trucks, blocking roads, driving into wet cement etc.
ChatGPT and friends make silly mistakes for trivial math problems that require a few hierarchical planning steps.
All in all, brains have some form of symbolic computation and reasoning that we haven’t been able to replicate with current AI algorithms.
I’m not saying we’ll never be able to but current AI is really hyped. Kinda like crypto boom of 2019.
There are some really hard algorithmic problems to be solved.
Google, Microsoft, Meta could have 1000x more computing power and data, however in the grand space or all algorithms there exists a learning algorithm that is probably >10000X more efficient at generalized modeling and reasoning than what we have.
The proof that we (20W biological generally intelligent computers) exist validates the hypothesis that there is a lot of advancement we can still do at the algorithm part.
Deep Neural Nets: 33 years ago and 33 years from now - https://news.ycombinator.com/item?id=30673821 - March 2022 (5 comments)
will there really be 10 million times 400 million images floating around then?
Perhaps it will be trained on whole videos, or a combination of different inputs from agents that move about in the real world / or a video game.
Not yet.
As the joke goes:
People in the 60s:
I better not say that or the government will wiretap my house
People today:
Hey wiretap, do you have a recipe for pancakes?