This challenge generalizes to all model fitting. Incorrectly assuming a distribution is Gaussian is a big one.
This challenge generalizes to all model fitting. Incorrectly assuming a distribution is Gaussian is a big one.
Amateur statistics is full of magical numbers and thresholds where everything "just works" :)
Statistics is more involved than it appears, there are so many ways something can subtly violate underlying assumptions, or appear to be fine on the surface while being actually meaningless. It's really easy to fit a model, it's more involved and difficult to actually understand the complexities.
http://www.math.nagoya-u.ac.jp/~richard/teaching/s2019/Cauch...
(Read: If X and Y are normally distributed random variables with mean 0 and standard deviation 1, then X/Y is equivalent to a Cauchy(0,1) distributed RV. This is a useful equivalence to know if working with standard normal RVs, which is often. E[X/Y] does not exist!)
If X ~ Cauchy(0,1), then X ~ normalized student-t distribution.
(student-t puts the “t” in t-SNE, where it acts as a weighting function on the Euclidean distance between points. UMAP uses a parameterized version of this weight function that is essentially a generalization of the Cauchy distribution: https://jlmelville.github.io/smallvis/umap.html)