Polynomial Regression as an Alternative to Neural Nets
arxiv.org
arxiv.org
No good ML practitioner believes that tiny, shallow, fully-connected neural networks are the best algorithm for every problem (just look at Kaggle results if you don't believe me). Especially for small, easy datasets like the ones analyzed in this paper, NNs are often not the best choice. However, for large scale image classification, density modeling, sequence modeling, etc. (none of which are tested in the paper), NNs are SOTA.
Hilariously, the paper only compares polynomial regression to tiny neural networks. I bet if they had thrown in results from XGBoost or other classical ML techniques, polynomial regression would be blown out of the water.
There's also some heuristics you can use to tell when a paper like this isn't necessarily representative of the field of ML:
- Missing citations (e.g. "It is well-known that NNs are prone to overfitting [Chollet and Allaire(2018)], which has been the subject of much study, e.g. [?].").
- Inconsistent formatting (e.g. the tables on page 7)
- Sentences like "Much more empirical work is needed to explore these issues." following something that sounds like an easy experiment to try.
I don't understand what it is you don't understand. Clearly HN is not a site dedicated strictly to academics and people focused on the ceremony of academia. So if a paper contains even an idea that might be interesting, and seems likely to foster some interesting discussion, it's probably going to get upvotes. It making the front page or not presumably depends on which other stories are also contending for a front-page spot at the same time. In any case, getting a lot of upvotes doesn't mean that the paper was "good", that its conclusions were justified, or anything else. It just means it was something that people thought was worth talking about.
Missing citations (e.g. "It is well-known that NNs are prone to overfitting [Chollet and Allaire(2018)], which has been the subject of much study, e.g. [?].").
- Inconsistent formatting (e.g. the tables on page 7)
It's a pre-print, and not a published paper, no? My assumption was that the authors intended to fix those in a subsequent revision.
The ML community I think suffers from some (shall we say) poor sportsmanship, with people occasionally posting rather half-baked papers to establish priority on an idea, before it's even been fully fleshed out or explored. This paper feels like that, on both the 'science' and 'bothering to use reasonable LaTeX' fronts.
Ah, interesting. At first blush that seems like that kinda defeats the purpose of releasing a pre-print, but I can kinda see why someone might do it that way.
- Grandiose-sounding claims ("the NNAEPR principle") backed only by generic informal arguments and/or well-known facts ("...for more realistic activation functions, simply note that they themselves can usually be approximated by polynomials...").
Judge papers based on their methodology, evaluations and findings (at least).
For some fields, you'll find the most cutting edge research on arXiv.
Citation count is useful for determining the relative importance of published papers. 10-100 citations is ok, 100+ citations means it made a significant contribution, 1000+ citations means it's a landmark paper in the field and probably worth reading. In the case of this paper, it's very new so citation count is not very meaningful, but it's a useful heuristic to evaluate most papers.
The surprising resurgence of deep learning since 2012 is not the observation that neural network models can learn at all -- neural nets with a handful of hidden layers were known to work fairly well since the 90s -- but that deep nets with dozens of hidden layers (enabled by GPUs) broke records and achieved state-of-the-art results on remarkably challenging tasks (famously, ImageNet 2014) in multiple important machine learning subfields (has revolutionized NLP, also AlphaGo).
This paper fails to recognize why people care about deep learning as a tool for machine learning in our modern day, and by making meaningless comparisons they restrict themselves to making meaningless observations.
OTOH, no one cares what the nonlinearity actually is so long as the net trains, and there's a lot of effort to keep layer inputs in a neighborhood near zero (batch normalization), so polynomial explosion may not be such an issue.
Feel like I would like a more serious comparison of their model's results with best of class NNs. I'm suspicious that their NN character detector was basically failing...
In this sense I think you could consider a big neural network with RELU activation a (really big) polynomial.
This is about the extent of my knowledge on the subject though, for more info search for tropical geometry [3].
[1]: https://en.wikipedia.org/wiki/Max-plus_algebra
[2]: https://mathoverflow.net/questions/83624/why-tropical-geomet...
Isn't the ReLU a piecewise linear polynomial? Moreocer, if you convolve piecewise linear polynomials you get a higher-degree piecewise polynomial, typically referred to as a spline.
edit: here's a paper about rational functions and NN's: https://arxiv.org/abs/1706.03301 . I'm not really in a position to evaluate the claims, but author just won an NSF CAREER award so that's something.
This is not true. Deep ReLU/leaky ReLU NN are piecewise polynomial functions generated in a contrived way (a chain of compositions). Even if we assume that piecewise polynomial functions aren't polynomial functions, they are still in the same class of linear functions.
I think gsl and chebfun (matlab, unfortunately) can do regression in one dimension but I don't think it can in 2 or 3d (much less general high dimensions).
Secondly, there was no claim that polynomial regression is better than NN in all cases, or is more effective than SOTA boosting techniques. The claim seems to be that PR is more simple, scrutable, and effective, for certain datasets when compared to NN. This seems like a non-controversial finding to me, given that it's common knowledge. I appreciate seeing results like this.
I might be wrong. How are their results relevant today?
Did you somehow already know the findings that they've made?
I see "Letter Recognition Dataset" from UCI ML repository. What's the hell is that? Can you point to any paper published in a respectable conference (NIPS/ICML/CVPR) in the last 5 years that used it? They showed results from 9 irrelevant datasets, while a single one on ImageNet would instantly convinced everyone they are on to something.
Why are you pro MNIST but anti Letter Recognition Dataset? A quick search of Scholar shows hundreds of papers that have used it in the last 4 years.
Again, I'm still not sure what you want. They aren't doing a paper on NN vs PR for image classification. They aren't trying to get a higher score on a Kaggle classification challenge. Their hypothesis is that PR is a good alternative to NN, and their findings show that - for certain datasets - they are correct. Will more research arise from this? Will future results be mindblowing? Maybe, who knows. It's just research, and somebody has to do it.
No, they are not correct, because they have not compared their model to state of the art NNs. That's exactly why I insist they use standard, canonical datasets, such as ImageNet: because it's easy to find well-known state of the art, and compare the models. Google Scholar, last 4 years: "Letter Recognition Dataset": 98 results, "MNIST": 10,000 results, "Imagenet": 18,000 results.
They aren't doing a paper on NN vs PR for image classification
If they include image classification results, and claim their model compares favorably to NNs, then yes, their paper is, at least in part, on NN vs PR for image classification. I only commented on image classification results because that's the field I'm familiar with. I suspect other results they presented would also be unconvincing to people working in those areas.
Will more research arise from this?
If they want more research to arise from that, then they should have made sure their results are convincing, or at least promising. As presented in the paper, they are not.
It's worth noting they do not provide enough information to replicate most, if any, of their empirical results. They recently (today) removed their experiments/ directory (can find in commit history), leaving only records of the CrossFit data analysis still in the repository.
Actually that's very wrong. Some classes of NN (step function, ReLU, etc) do define piecewise polynomial functions, albeit in a contrived way when compared with simply generating a spline. So the functional space is exactly the same.
But anyway, should not be surprising that function approximation can be done with polynomials. Regression has big problems though, as the problem size increases, which SGD/backprop seems impervious to. That’s kind of the mystery behind neural nets.