Energy, not compute, will be the #1 bottleneck to AI progress – Mark Zuckerberg [video]
youtube.com
youtube.com
This whole movement of simply increasing compute or throwing in more data to hopefully attain AI results is just a dead end.
This is like saying if human beings just flap their wings hard enough, they will be able to fly like birds.
We need to finance original research, alternative research and think of working on an actual theory of intelligence.
The stage AI is at right now is more like the Ptolemaic model of the universe.
We need a Galileo and Kepler of AI. Maybe they already exist but are being marginalized by the current establishment?
There's no way for you or for anyone else to know that it's a dead end.
Also, it's not "simply throwing more compute" - the algorithms have evolved too and they'll continue to.
Also also, it's not "hopefully attain AI results" - we have AI results now and they keep getting better.
People are spending resources on a trajectory which has worked well - we've been cranking up the compute and better results keep coming out, why stop now?
It's ok to have deep learning. But the herding towards deep learning is just too much.
Recall that before Hinton and his group presented thier record breaking ImageNet results, deep learning was a fringe technique that the likes of Marvin Minsky regarded as a failure in "Perceptions"
Scientific progress is made by researchers exploring other alternatives.
Herding can only slow down progress.
The “other strategies” have been completely sidelined. Even back when we had no transformers, only CNNs and RNNs, the other strategies were tossed aside because they didn’t show much promise in cifar or mnist or whatever benchmarks the NN community deemed the most suitable sota. Even before that, when we only had decision trees and catboost, people had already dismissed loess/splines etc. You can get a flavor of the dismissals here -
https://news.ycombinator.com/item?id=19145706
At this point that horse has left the barn.
It’s no surprise that industry has been pushing for more compute ever since.
Are people not exploring other strategies?
They are very few. Too few.
You don't know that either. What if there are actually no alternatives.
Isn't that what we are trying to simulate?
Has Generative AI Already Peaked? - Computerphile (https://www.youtube.com/watch?v=dDUC-LqVrPU)
Then generative AI became good. Good enough for me to use it as a daily assistant. So good it scared me and led me to rethink my position.
I have missed one crucial thing: Big enough changes in quantity do lead new qualities as well.
Training models with insane amounts of data was the necessary step to finally make the qualitative leap into developing models that were practical for mainstream use.
One of the big things holding back AI research has always been trying to be too clever. This article sums it up perfectly:
> The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin. http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Let's imagine a modern day Galileo discover a new revolutionary algorithm that gives the same results for 1000x less compute.
How would he transform this competitive advantage into some money without revealing the secret sauce ? Today's he will be in competition with people spending $10^8 to train. Even with the 10^3 cost reduction it's out of range for most people. And then he has to build a business around while keeping making sure his secret techniques don't leak.
VCs and CEOs are just playing the controlling the business game, raising barrier to make sure they stay in control. Incremental advancements are great for them, because they get to stay in the lead while reducing the possibility that a revolutionary advancement comes to shuffle their game.
Renaissance Technologies could maintain an edge over the whole stock market for plenty of years. AI-controlled company will probably be able too, but you won't hear about them.
> We need a Galileo and Kepler of AI. Maybe they already exist but are being marginalized by the current establishment?
The problem with "theories of intelligence" is that they need to yield results. You can come up with whatever brilliant theory you like but at the end of the day you need to produce a tool, a model, a result, something that does something.
And that's exactly what "theories of intelligence" have failed to do. Your argument is like Chomsky whining in the NYT about machine learning. What did Chomsky's linguistics actually do? What did they enable us to build? The answer is probably not "nothing" -- it's more like "not much". Meanwhile, machine learning gave us a tool that can translate from one language to another reasonably well and didn't stop there.
Theory isn't useless but we have plenty of it. There's so much borderline useless theory about language and intelligence in academia and it has produced so little tangible value. We absolutely do not need more of it.
Hahaha, oh man.
Von Neumann, and many of his contemporaries, weren't secretive researchers toiling in obscure corporate labs -- they were famous public intellectuals with renowned accomplishments across many fields. There are no men of that caliber alive today, and the educational system absolutely excels in ensuring that such well-rounded men are not produced.
SV runs 80/20 on investors getting fleeced vs. vaguely realistic business models.
It's very difficult to get funding on projects that don't already have a track record of success.
Which is why all research facilities working in AI today that I have read about only do research on Neural Networks
Just look at what is being published in Arxiv.
Call your friends in university departments.
Ask people working in startups or corporation.
AI at this point is just Neural Networks.
Do I think it's a dead end for AGI?
I really don't know what AGI is, but for what ever new possibilities we can unlock.
We need newer techniques that are less intensive computationally.
It appears that a lot of breakthroughs came from further understanding how the brain works but we are just trying to brute-force our way to do harder things.
(And our brains also do things that AIs cannot even when running on a few salads a day.)
Hopefully? The last 10 years of astonishing AI results have largely been driven by “throwing in more compute and data”.
It’s the logical incremental step until an actual breakthrough is model on technique
All evidence is the throwing even textbook qualify data at a model of almost any size just approaches an asymptote just a tiny bit above GPT 4.
A better model of some data starts to look increasingly like then data, not like something else beyond the data.
A huge Freudian slip there.
I wonder how much of this over supply can be taken up with AI training and how much that would change the calculations.
If we want an alternative I think we'd need to try alternative training algorithms that are easier to distribute. Predictive coding for example can approximate backprop, but requires more compute. Something in that direction might work https://arxiv.org/abs/2006.04182
AI requires high bandwidth and dynamic compute. You need a whole supply chain to a data center. Repairs become expensive when you have to ship in highly trained people to fix the machines. Swapping out an OAM/UBB AIA isn't like pulling a gpu card from a PCIe slot.
AIA servers are complex 350lbs beasts that have a lot of moving components and require dedicated trained technicians to replace parts. They have multiple ways they connect to multiple networks. They require a lot of software. They are very dynamic with continuous updates.
It's like saying your employees need oxygen to breath or they can't work at scale.
- creating and executing complex plans in arbitrary scenarios
- counting to 6 with 99% reliability in arbitrary scenarios
- understanding object permanence with 99% reliability
- understanding simple, intuitive causality with 99% reliability
But rats can't answer Python questions, so according to a truly depressing number of tech folks, of course GPT-4 is smarter than a rat! AI researchers/capitalists seriously need to spend more time reading animal cognition. If Silicon Valley was in charge of animal cognition research, spiders would be considered smarter than crows.
Memory currently is a concern.
Algorithms are my only concern. Its not like these algorithms can snag some protons and electrons and analyze them, they are at the whim of human progress/data collection.
Energy, above my paygrade.