Gradient Descent Models Are Kernel Machines
infoproc.blogspot.com
infoproc.blogspot.com
I myself am very interested in [2]. It's fairly dense, but I've been meaning to go through it and the larger Tensor Programs framework ever since.
[1] https://www.reddit.com/r/MachineLearning/comments/k7wj5s/r_e...
[2] https://www.reddit.com/r/MachineLearning/comments/k8h01q/r_w...
[3] https://www.reddit.com/r/MachineLearning/comments/k8h01q/r_w...
With a sufficiently magical kernel function, indeed you can get great results with a kernel machine. But it's not so easy to write a kernel function for a domain like image processing, where shifts, scales, and small rotations shouldn't affect similarity much. Let alone for text processing, where it should recognize 2 sentences with similar meaning as similar.
Every tom, machine-learner, and harry, is an expert on proving to the whole world that thing-A is thing-B. The only problem is people don't hire them with a million dollars a year total compensation.
I think the key issue at hand is that gradient descent is easier to train than a model using kernel functions. Someone could absolutely devise a mechanism for back propagation of errors with kernel functions, but at that point, it is basically a neural network.
With a way to turn a gradient descent into a roughly equivalent kernel machine you can turn gradient descent into an embedding, so rather than merely seemingly learning features you can recover the actual feature space.
It's also interesting how the result heavily depends on the path taking during the gradient descent rather than merely the end result. This is somewhat counterintuitive when for the equivalent gradient descent only the end result is seemingly important.
I don’t understand. What do you mean here?
I only (quickly) skimmed the article but it says that the limit is equal to a (not the) zero of your gradient.
The result should be the same for every trajectory leading to the same zero. Put another way, it doesn’t depend on the path if it leads to the same zero.
I would say here that not all zeros are equal though. :)
After reading just the linked article, not the original paper, my interpretation is that the kernel isn't calculated during the gradient descent process, but from the final model after training by gradient descent. If that's right, it's not path-dependent.
How so? The existence of equivalent kernels doesn't imply they aren't clever or good.
Basically what they're doing is just y = y0 + \int dy/dt dt and then rephrasing this equation as a kernel machine. Which I guess should also work for other forms of learning where you slowly change the parameters.
This statement is a bit strange, because most of the theoretical work around NTK assumes that the tangent kernel doesn't change sufficiently during training. This seemed to make this closer to "extreme machines" than deep learning.
There was a paper from Francis Bach's group which seemed to be saying as much, but that criticism was no match for the meme-fication of this line of research by over-enthusiastic amateurs.
https://news.ycombinator.com/item?id=25314830
It's a really neat one for people who care about what's going on under the hood, but not immediately applicable to the more applied folks. I saw some good quotes at the time to the tune of "I can't wait to see the papers citing this one in a year or two."
But I get the sense that this is one of a number of big pieces of the puzzle of how current cutting edge models can be cut down in operational size. I have a feeling that new methods of creating models will stick around, but any deep learning model can be reduced to a less resource intensive approximation through findings such as this. And if (IF!) so a lot of AI hype is going to hit a real post-hype cycle plateau of productivity because all this stuff will become easier and cost effective to deploy to the real world.
... that also happen to actually work.
You can get point estimates for many completely different models using gradient descent. The article clarifies what is meant in the first line, but it still seems like a very poor choice of words.
Now get off my lawn.
[1] https://www.microsoft.com/en-us/research/blog/three-mysterie...
So at the end, it rephrased a statement from "Neural Tangent Kernel: Convergence and Generalization in Neural Networks" [https://arxiv.org/abs/1806.07572], besides in a way which is kind of miss-leading.
The assertion is known by the community at least since 2018, if not even well before.
I find this article and the buzz around a little awkward.
In many cases though, it is often insufficient. The next step is piecewise linearity and then convexity. It is said that convexity is a way to weaken a linearity requirement while leaving a problem tractable.
Many real world systems are nonlinear (think physics models), and often nonconvex. You can approximate them using locally linear functions to be sure, but you lose a lot of fidelity in the process. Sometimes this is ok, sometimes this is not, so it depends on the final application.
It happens that linear regression is good enough for a lot of stuff out there, but there are many places where it doesn't work.
Technically, only if you don't zoom in too far, quantum mechanics is linear.
- yes, QM is linear in terms of states, vector spaces and operators acting on states (its just functional analysis [1])
- it’s not linear in terms of how probability distributions and states behave under the action of forces and multi-body interactions (see, eg correlated phenomena like phase transitions, magnetism, superconductivity etc)
However, this only works (practically) for mildly nonlinear functions. If you have exp, log, sin, cos functions for instance and you're modeling a range where the function is very nonlinear, the approximation errors often become problematic unless you use a really dense cluster of knots. For certain nonlinear systems of equations, this can blow up the size of your problem. It's a tradeoff between number of linear basis functions you use for the approximation, and the error you are willing to accept.
Probability undergirds a lot of statistical methods.
(Absolutely just kidding around here)
Is this another way of saying that neural networks are just another statistical estimation method and not a path to general artificial intelligence? Or is it saying that problems like self-driving cars are not suitable to the current state of the art for AI since we have to ensure that reality doesn't deviate from training examples? Or both?
I'd love to understand the "real life" implications of this finding better.
I’m not sure why humans are so good at that. I feel like it might be that by spending the first couple years of life picking things up and rotating them and putting them in our mouths, we have some deeper understanding that images of an animal like an anteater are not really two dimensional but are supposed to represent volume and form. That’s my guess. A neural net trained to recognize 2d images will not have that.
If you start slightly before 2yo, you may find that after they see a cow they will call every large animal a cow for weeks/months - that's closer to one sample without all the models.