New method identifies the root causes of statistical outliers
amazon.science
amazon.science
Without a model there is no difference between outliers and "regular" data points, because a data point will always match a distribution made from itself.
If this gets resolved in a HN thread you should publish it.
Treating them as exceptional is simply a sign you don’t have enough data.
Can't you just say those are outliers and not consider them? Ideally you'd build some bimodal model, but given that they don't matter, what's wrong with throwing away those samples?
Outliers are still useful data!
Definitely untrue. Suppose we throw 1,000 baseballs, shoot one rocket (45 degree angle), record the distance each one flew, and present the 1,001 data points to someone for analysis. Most people, statisticians or not, will recognize the rocket's distance as an outlier, even if they don't know the cause of it. And they'll be right to do so.
It's true that there is no universal definition of outlier. But outliers are very clearly defined in some contexts.
But I would argue that's not really a case of an outlier, that's a mixture distribution because the underlying identifying feature (baseball or rocket) is missing.
"There are two cultures in the use of statistical modeling to reach conclusions from data. One assumes that the data are generated by a given stochastic data model. The other uses algorithmic models and treats the data mechanism as unknown. The statistical community has been committed to the almost exclusive use of data models"
From the point of view of the data it is definitely an outlier.
> that's a mixture distribution because the underlying identifying feature (baseball or rocket) is missing.
This is the point of the paper. They literally say in the blog post summary:
To attribute the outlier event to a variable, we ask the counterfactual question “Would the event not have been an outlier had the causal mechanism of that variable been normal?”
Statisticians are known to use this kind of model. See https://en.wikipedia.org/wiki/Dirichlet_process
Just because you have a name for something doesn't mean it doesn't have other names as well.
The phrasing implies mutual exclusivity here, which thus implies probability one of a data point belonging to one set and probability zero of that point belonging to the other set.
To be fair, in natural language it tends to be difficult to unambiguously represent ambiguity.
No statisticians would ever propose that as a practical industrial method. Sometimes a bat falls in your vat of beer at the exact moment the lid is off and no sensible statistician proposes modelling the flight path of bats.
Understanding and removal of outliers has been part of the statisticians toolbox forever. Gosset's[1] (The inventor of modern statistics) whole work on Student's t-distribution was because he had to deal with small sample size and know when sample weren't representative, and how to discard outliers.
Studentized Residual is named for him and is (to quote Wikipedia)[2]: an important technique in the detection of outliers.
The statisticians I've worked with are very pragmatic, and none want to try to model the whole universe. It was a statistician who came up with "all models are wrong but some are useful" after all.
What is the innovation here, aside from a new software library? The quantification of each candidate root cause's influence on the outcome? I am surprised the authors found nothing similar throughout the entire corpus of academic research.
I must say, Amazon's "science" blog is the most unimpressive of the big tech companies. It churns out PR like the rest, but the others at least have some substance behind them.
Seems like a very unscientific way to conveniently ignore something that doesn't fit into the way you want it to be. "Look at all this rigorous math I'm doing! Oh, those points? In my expert opinion they're no good. Just delete them"
Seems they haven't had them for a while, anyone know why? They were interesting, especially for academic titles.
This is why in statistical process control these types of outcome are known as "common-cause variation" and "assignable-cause variation".
I am not a statistician, so I don't know under what circumstances outliers are usually thrown out.
As an industrial statistician, I can tell you: way too often.
Outliers are the signal among the noise. They indicate something. It is nearly always worth finding out what, instead of removing them. If they indicate a flaw with measurement or the process, then fix that flaw and re-do the measurement or re-run the process. Outlier gone! But in a much more informative way.