http://www.incompleteideas.net/IncIdeas/BitterLesson.html
I would love to hear HN's take on this argument.
http://www.incompleteideas.net/IncIdeas/BitterLesson.html
I would love to hear HN's take on this argument.
It used to be almost everyone worked in agriculture, now about 1% does, so the others are free to do something else. Prior to the microprocessor making a computer required manual assembly of thousands of parts, early microprocessors contained thousands of parts manufactured by a small number of photographic and chemical steps, and the number of parts has grown into the billions without the number of steps expanding millions of times.
On the contrary, neural networks will never give you a better way to convert to the frequency domain than the Fourier Transform. At best, they might approximate it.
No. That is the use case for numerical analysis, where you can develop a high performance, accurate algorithm based on the mathematical analysis of the problem.
It is really kind of silly to want to encode the solution in some neural network.
Maybe there are some very special problems, without well developed silver where they work.
For medical imaging your solver is most likely "physics based" in any case. PINNs want to encode the physics inside a neural network, instead of developing an appropriate algorithm (which obviously has to also incorporate the physics) which solves the problem and which can be analyzed mathematically.
See, the reason for the Bitter Lesson being a lesson is that Neural Nets, that Sutton is mainly writing about, are pretty crap at representing background knowledge. The only way you can store expert knowledge in a neural net is to modify its structure and its weights. The weights you can only modify by some kind of learning procedure like backprop, in practice. Very limited forms of background knowledge, like convolutions can be encoded in a neural net's structure, but imagine trying to represent, I don't know, the last ten lines of code you wrote today as a bunch of neural net connections. Continuous functions is just not the right kind of notation for that sort of thing.
If neural nets were any better at encoding background knowledge, they would use it, but they can't so they have to rely on data. And that's why they need so much of it. Background knowledge functions as a strong inductive bias- it directs the search for a hypothesis to hypotheses that we know make sense (again, think of convolutions). Without background knowledge, or with only a little background knowledge, you need tons of examples to learn anything useful.
So the Bitter Lesson is basically making a virtue out of necessity. In any case, it's not a prescriptive thing, only descriptive.
Now, Sutton argues that some of those systems at least did not rely on domain knowledge. They all did: Monte Carlo Tree Search, used for board game-playing AI agents, is nothing else but an encoding of domain knowledge - specifically, domain knowledge about the structure of two-player, complete information games. It's the same domain knowledge that was used to create the minimax-based DeepBlue software that won against Kasparov.
Neural nets have still not managed to win against human players in board games without incorporating such a strong, knowledge-dependent component, as a game-tree search.
So sutton is fudging the details. Not on purpose. He's an RL person. In RL, as in planning and other disciplines, folks tend to forget all the knowledge they put into their systems in the form of inductive biases, or auxiliary (but can't-do-without) algorithms like MCTS. He's like the proverbial fish that don't know what water is, because they swim in it.