The Bitter Lesson (2019) [pdf]
cs.utexas.edu
cs.utexas.edu
http://www.incompleteideas.net/IncIdeas/BitterLesson.html
u/dang, swap links if you see this?
Big reminder that the Bitter Lesson is _not_ saying "just scale your methods and they work". What the Bitter Lesson _is_, however, is "work on methods that scale". There is a _huge_ distinction between the two, in my opinion.
Effectively, if I can add my layer of (re?)interpretation on it, it's saying that specialized boutique, hand-designed solutions don't play very well in the long-term in an arena with Moore's Law and money. But, what it's not saying, is that just making an algorithm bigger is the solution. This is how I see most people missing its interpretation.
For an algorithm to be effective, it needs 3 things in my opinion: 1. It needs to scale (measurable by some factor) 2. Because we have 1, it needs to have an implementation with extraordinarily rapid iteration time 3. If we have 2, we need an implementation that is extraordinarily lightweight (enough to run on consumer machines)
These three factors together, in my personal estimation, unlock algorithmic research progress in an area.
An ancillary fourth rule that really drives progress IMO is competition, formalized or otherwise, that is 1. open, 2. well-known, 3. incentivized and 4. accessible.
This is my personal opinion of course, and bound to be flawed in some way, but -- from personal experience, at least -- when I've seen this field (or research fields in general) align with this kind of method of research is where the speed of algorithmic research has absolutely exploded. :) <3
> They said that ``brute force" search may have won this time, but it was not a general strategy, and anyway it was not how people played chess.
But it is now clear that these people were completely right! The Deep Blue approach (apparently mostly brute force search) didn't scale. E.g. Go wasn't solved this way, and the AlphaGo approach (based on reinforcement learning) did in fact scale to a lot of other games with AlphaZero and MuZero.
Sutton in fact says that both search and learning produce the big success in AI. But now, with the dominance of ML, it seems that mainly learning is responsible the big successes in AI, not search. Machine learning means that the AI models aren't written by hand, GOFAI style, like Deep Blue, but instead by a learning algorithm, which in turn is written by hand. That seems to be the breakthrough, not brute force search. (Do GPTs use "search" in any conventional way? I don't think so.)
The purpose of SGD is to find parameter values that minimize a training loss.
Pretrained dense GPTs don't search in any conventional way. However, when these GPTs are fine-tuned using RL, an aspect of search is reintroduced. One the most interesting directions labs are pursuing to improve on today's best LLMs is to fine-tune them with search (RL) after a learning-based (next-token prediction) pre-training.
I think they're currently solving it for eight pieces.
So it's learning but then brute-force for the end game.
Regardless, the author's point is that computation is a better way of finding and exploiting patterns/strategies than our own intuitions. The distinction between search and learning is not the important one here.
Parallelizable methods win (e.g. DL), by using the computation available
Further: progress will occur on real-world problems that can be solved by those methods, making those applications dominant throughout society. (i.e. everything looks like a nail to someone with a hammer... but if it's a really good hammer, that may be the best approach)In my head this always parallelled the "premature optimization" conversation in programming. Most programmers would say that inlining etc was only justified when you've benchmarked so you know how the benchmarks are performing but the experience of whole program optimization in things like hotspot jvm and llvm seems to suggest that even benchmarked optimizations can be premature because only a vm can do optimizations on the real world use case.
One of the points, I think, is that there is very little a researcher in the 70s could have done to make progress on these problems, because the computing power just wasn't there yet; and once the computing power was here, almost all of that 70s researcher's work became obsolete.
John McCarthy described chess as "the drosophila of AI":
http://jmc.stanford.edu/articles/drosophila/drosophila.pdf
Who was it that described Go as the drosophila of AI?
See section 7:
As a fourth Drosophila I would like to mention the research on Computer Go.
Thanks for the correction.
this is even more striking in poker, where counterfactual regret minimization wasn't invented until 2007, despite being a relatively simple algorithm to describe and all the essential intellectual building blocks being known since the 1960s.
von Neumann had invented extensive form games with the express idea of modeling poker; there were researchers in the 1950s/60s (Hannan, Blackwell) working out the core ideas of regret minimization. One innovation (the sequence form representation of game strategies) was only known in the Soviet Union in the 1960s and not propagated to the West. But this was independently rediscovered in the early 1990s and it still took another 15 years for CFR to be developed.
I agree and think “ML explainability” efforts are doomed to fail as ML becomes increasingly more effective. There is no a priori reason that the human brain should be capable of intuitively grokking sufficiently advanced general learners. We can invent them and improve them, but saying that we will be able to understand what the myriad matrix multiplications are “doing” will be like saying we understand the human brain because we can model the physics of its constituent atoms. The emergent complexity is too high for us to make any sense of it.
P-:
> How are you going to improve something you dont understand?
Is just nonsense. Evolution understands nothing, yet produced a mind. Closer to us, the early people who produced all the crops that led to the shift to agriculture, and it's later improvements, absolutely did not understand how any of it worked.
Evaluation and selection are sufficient to improve things. Understanding is useful, but optional.
So how would you "improve" here an now any algorithm. Create and insert random code and evaluate? Jeepers people are losing touch with reality.
How does that opinion not make sense? There are numerous things humans have invented for which we have little understanding of how they work: medical drugs, anesthesia, certain quantum phenomena utilized in semiconductors, etc.
I would argue that for current state-of-the-art LLMs, the implementation is likewise ahead of the theory at the moment.
There is a strange emerging AI cult that is also in force here in HN that seems to believe these algorithms have evolved themselves or were some random trial and error. Ergo, they can keep evolving and the researchers dont need to understand a thing about how they work.
Serendipity in combining ideas that prove effective plays a role but within a fairly well defined conceptual sandbox. But progress with AI is more or less conditional on people having sufficient understanding to coax algorithmic structures in the desired direction.
Which is basically what we've been doing with AI so far, as the paper notes (among other fields, like medicine). Have you heard the joke "grad student descent"?
Brook's post goes over the classics (Moore's law is ending, curating a dataset requires human intervention, etc) and posits that making a huge model won't be a competitive strategy for long because it gets to expensive to train and use.
It's a bit early to tell, but so far that hasn't materialized. OpenAI got state-of-the-art results with GPT-4, AFAIK by sticking for very-super-big models together. Open source experiments with LLAMA show you can still get good results with heavy quantization. Distillation hasn't be too explored by mainstream projects, but I bet there's lots of potential there too.
Right now the winning strategy looks to be "go really big, then figure out how to go small".
If you do not have solid scaling, your link between micro methods predicting mega methods is broken. Hence, Sutton's bitter law forms the foundation for a few other lemmas that I think underpin what makes really effective research (which is iteration time, and how we effectively reduce it as much as possible and make it as accessible as possible -- which thankfully for ML algorithms seems to go hand in hand! <3 :')))) )
1. Not all nets work for all problems, those that work tend to have the right inductive biases. We discovered the architectures partially by trial an error, nevertheless they work because of encoded prior information.
2. Data and computation are bounded. GPT4 was basically trained on all text, further advancements probably need more insight not more data.
ConvNets hardcode translation equivariance, more general convolutions can hardcode more general equivariant structures.
The whole field of geometric deep learning is about constructing nets based on insights about the structure of the data manifold, and many successful nets turn out to do this.
Well, there's multi-modal training. There's tons of untapped audiovisual data.
Transformers are preferred to ConvNets these days in Computer Vision despite the latter having all sorts of vision based inductive biases.
GPT-4 has not been trained on all text lol
2. Gpt 4 was trained on 13T tokens, all books ever written would be about 6.5 T tokens by Fermi calculation.
2. Assuming the rumor is fact, GPT was trained for multiple epochs and books are a small percent of what trains LLMs lol.
So it was always going to end, and further advancement was always going to revert being driven by a neural net of some sort. In fact, it looks to me we are already at that point.
The interesting question is what neural net will end up driving it.
[0] https://www.qualcomm.com/news/onq/2023/07/generative-ai-tren...