I made a comment in that thread, but I didn't notice the 12 degree polynomial.
Perhaps I'm too optimistic, but I expect someone to read the paper and make a comment about that in HN.
I made a comment in that thread, but I didn't notice the 12 degree polynomial.
Perhaps I'm too optimistic, but I expect someone to read the paper and make a comment about that in HN.
Making any claims based on that kind of curve fitting is a huge red flag, especially if you don't discuss that.
One of the reasons we would prefer lower-order fits to higher-order fits is that there is a very real risk of overfitting in the interior of a data set but then providing completely inaccurate results anywhere but at the data points. Seeing that kind of overfitting in a scientific paper without any justification suggests that the author of the paper is making statements without any basis in science.
If you want to make it simple, think of using a 12th order curve as being similar to picking whatever 12 points you want to match and then drawing lines from point to point.
We later discovered that the tame functions that usually work well for extrapolating from a sample are things called splines.
It like saying that a rule with as many special cases as data points is not actually a general rule. It is just a collection of special cases that explains nothing.
More precisely, if you give a model enough parameters, you can always have it describe your data well. The question is whether it is likely to predict future data. And the answer is no.
See the last example here: https://xkcd.com/2048/
Highest I know is Lighthill's (aptly called) eighth power law which says that the sound power created by a turbulent flow scales with the eight power of the characteristic turbulent velocity: https://en.wikipedia.org/wiki/Lighthill%27s_eighth_power_law, anyone knows something higher?
The difference is that in my example and in your example there are only a few coefficients to tweak to fit the data. So the shape of the curve if fixed and it is very difficult to overfit the data.
In the paper they used a full polynomial of degree 12, that has 13 coefficients to tweak and it is very easy to get weird shapes.
> they use a 3rd degree poly to remove the noise it seems... not to fit very well
They are not trying to fit the noise, they are trying to fit the hidden smooth signal that if hidden by the noise. In some cases it is difficult to make a formula for the real signal, so you can approximate it locally with a polynomial.
The idea is that after you subtract the smooth part, you get only the noise. So you can calculate the expected noise level.
And if the "noise" has a big peak, you can guess there is something strange, like a big absorption line.
See also my other comment: https://news.ycombinator.com/item?id=24985680
EDIT: Also see the last panel of this XKCD: https://xkcd.com/2048/
(a) a smooth function that is not interesting
(b) noise
(c) a big unexpected peak in an interesting part
There is a nice preprint posted a few days ago: https://arxiv.org/abs/2010.09761 https://arxiv.org/abs/2010.09761.pdf
They reanalyze the data than in the phosphine paper.
---
If you look at figure 3, they have the data that is the skyscraper-like line, and they use a polynomial of degree 3 to approximate the signal, that is that curved unhappy smooth line.
The smooth line is (a), when you subtract this smooth function, you get the other part of figure 3, that is the noise (b). There are some high and low parts, but nothing too high or low that look special. So their conclusion is that there is no interesting part (c).
---
If you look at figure 2, top left, they use a slightly smaller interval, but now they use a polynomial of degree 12 instead of a polynomial of degree 3. This is the smooth function (a).
This is a reconstruction of the process in the original paper. They fit the polynomial using the data, but excluding the central part.
The problem with the polynomial of degree 12 is that is has too much freedom, so it fits the actual smooth curve, but it also fit the noise.
The polynomial of degree 3 has to go somewhat in the middle of the data, because it can't go up and down too many times. The polynomial of degree 12 can follow the local bumps and fit the noise.
When you subtract the polynomial of degree 12 you get the the graph in the third row, with the noise that is (b). It is copied in Figure 2, and it is very similar to the graph in the original paper.
Since the polynomial of degree 12 fit the noise, the noise is too small, so you underestimate the noise level.
And since the central part was skip in the fit, in some case you get a big bump like here. It is bigger than the apparent level of noise so it looks like an unexpected peak (c).
But here the problem is that you are comparing the peak with the surrounding noise level, but the noise level is underestimated because the polynomial of degree 12 overfit.
---
They repeat the same kind of analysis in other regions, and they get a few additional fake peaks. This are the other 5 graph it the top of Figure 3.
Often it's better to use a spline fit or other interpolation technique, once you find yourself needing to go beyond five- or six-degree polynomial fits.