Statistical Formulas For Programmers
evanmiller.org
evanmiller.org
The current selection on that page is somewhat limited, but I hope to grow it over time. The stuff at the beginning is pretty basic (e.g. standard deviation), but things get pretty gnarly by the time you get to the Kiefer equation. At some point I'll add some more references on how to implement things, e.g. find successive zeros of Bessel equations. For now it should be a good jumping-off point. Enjoy!
I have been thinking about collecting some stories about how people have put their knowledge of stats when programming (performance testing, user patterns examination, etc), what do you think? It would be a kind of applied stats for programmers thing I guess.
I should emphasize that this is not a nitpick or even a criticism, just a feature I would love to see. It's also what I spend a large portion of my time trying to track down, so having it in a convenient location would be nice.
No statistical technique is assumption free, unless it is purely descriptive.
Some of them are free of explicit assumptions known by the practitioner, but that's not the same thing. In much the same way, my code is all bug-free.
machine learning techniques, which tend to be assumption-free
ML should be a rigorous exercise in Bayesian and classical/frequentist stats, computational methods, dataset integrity, visualization etc, if you've been thru the texts by Murphy or Bishop. It often happens that people a couple years out of their last stats class only retain that high R-squared, p-, t- and f-values are what they're looking for, and heteroskedasticity and sphericity are just big words.My evidence that ML is a rigorous exercise: the free texts listed (Barber, Mackay and Smola's are excellent, ESL not as accessible)
http://metaoptimize.com/qa/questions/186/good-freely-availab...
http://udel.edu/~mcdonald/statintro.html
I point a lot of newbs to pages on that site so that they can develop a better intuition for the methods.
I wish I had this last semester during my statistics course.
It's written for biologists, but you don't really need to know much biology to work through the examples, and the focus is inherently practical.
A small amount of correlation can lead to very biased estimates of significance.
Correlation between observations will generally appear in any time series data or data that is arranged spatially.
Again, just a rule of thumb, but you should be very wary of the lack of independence between observations.
http://gdr.geekhood.net/gdrwpl/metnum.php
For me formulas written in pseudocode are much easier to understand than classic mathematical notation.
For example I've learned bayesian classification, chi-square, etc. from Practical Common Lisp (http://www.gigamonkeys.com/book/practical-a-spam-filter.html) after failing to understand how to apply formulas from Wikipedia. It was easier for me to learn Lisp than to decipher abstract declarative mathematical notation (admittedly I get a brain freeze whenever I see ∑, even though I know what it means. I prefer `for(…) acc += …`).
Not to be a dick, but getting the first example wrong like that doesn't inspire confidence in the rest of the post.
[1] https://en.wikipedia.org/wiki/Standard_deviation
[2] https://en.wikipedia.org/wiki/Unbiased_estimation_of_standar...
The usage of n-1 instead of n in the denominator was a question I was always asked when I TA'ed for intro statistics classes in grad school. An explanation of unbiasedness might be warranted if this is to be an introductory primer.
Thank you, Evan Miller! Your web page is great, it taught me new stuff from the very first formula. Much appreciated.
[1]: http://www.amazon.com/Statistics-4th-David-Freedman/dp/03939... [2]: http://www.amazon.com/Statistics-11th-Edition-Book-CD/dp/013... [3]: http://stats.stackexchange.com/questions/421/what-book-would...
Thanks, I'll check out your links.
Btw, what do you think of OpenIntro statistics?http://www.openintro.org/stat/down/OpenIntroStatSecond.pdf
I haven't seen this OpenIntro statistics before. I'll check it out!
http://www.stat.cmu.edu/~larry/all-of-statistics/
It was written as an introduction to statistics for people in CS and related fields.
The e-Handbook won't make you an expert statistician, but as an engineer needing to understand and apply statistical methods, I've found it to be a good starting point.
https://www.udacity.com/course/st101
Just took it earlier this year. It was informative and I enjoyed the class.
Just take care, there's a lot of problems from using statistics in wrong way. special care with small sample size and even large ones [http://scienceblogs.com/mixingmemory/2006/10/31/jeffreylindl...]
I would disagree with SagelyGuru in recommending robust regression for non-statisticians, though I can see where he or she is coming from. With robust regression, you don't have to worry as much about assumptions as with classical regression. But with robust regression, you need to be aware that the underlying analytical method is different and what that means. For example, the standard robust regression implementation in R (i.e, the rlm function in the MASS package) doesn't produce t-statistics or p-values. There're also warnings that especially at lower sample sizes, the standard errors produced by rlm may be unreliable. One recommended way to obtain those p-values would be to get bootstrapped standard error estimates, so that normal-theory approximation would apply.
[1] There are different types of robust estimators (e.g., M, S, MM, etc.) that have different robustness properties.
This could be useful- PHP stats functions: http://www.php.net/manual/en/ref.stats.php
"A confidence interval reflects the set of statistical hypotheses that won't be rejected at a given significance level. So the confidence interval around the mean reflects all possible values of the mean that can't be rejected by the data."
That seems a bit vague and perhaps confusing. Might I suggest something more like this:
"The confidence interval specifies a range (+/- a multiple of the above standard error [SE]) around our estimate of the mean (x-bar) such that: if we repeated our sampling process an infinite number of times (i.e. with the same sample size and forming a new x-bar and SE each time [and therefore, a new confidence interval]), Confidence_Level% of those intervals would contain the population (true) mean."
In addition, I think in this case, at least, there are no assumptions about the data to worry about, given a sufficiently moderate sample size due to the Central Limit Theorem (I'm confident about that in the case of the mean (x-bar), but I'll leave it up to others to correct me if I'm wrong about this applying to the standard error (SE)).
Most of these formulas are very rarely used even by quantitative analysts. The most used are for standard deviation and regression. The more complicated ones are generally used as a part of statistical routines, say, in R. It is very rare that someone has to code them.
> From a statistical point of view, 5 events is > indistinguishable from 7 events. What is this supposed to mean? There is a concept of statistical significance but if an effect is not statistically significant it does not follow that it does not exist. Btw where is the Bayes formula? :)
The first superpower is to look at the data and see if it makes sense using your eyes and your brain, not to start spewing confidence intervals.
As other posters have pointed out, it is even more irresponsible to start waving around things like the t-test without discussing the parametric assumptions that these things depend on for their validity.
Be great to see some pictures to illustrate the formulas and some mention of robust statistics as I find outliers to be a huge issue in application of statistical techniques.
To pick a couple examples from health news in the popular press:
NB: I'm making up all the numbers here for the sake of example.
(1) A study shows that people who consume more than 10g of added salt a day live shorter lives.
But how much shorter? If it's 30 minutes shorter, I don't care about the study and I'm not going to change my behavior. If it's 6 months longer, then I'm interested and might very well do something.
(2) A study shows that people who drink 2 or more cups of coffee a day have lower risk of Alzheimer's Disease.
But how much lower risk? If the average lifetime risk is 1 in 50, and drinking coffee lowers it to 1 in 49.997, then I don't want to waste time even reading the article. If it lowers it to 1 in a 1000, then yes, I might change my behavior.
So, in the above examples, is there any way to reduce the information into a single How-Much-Should-I-Care number?
Like this:
(1) A study shows that people who consume more than 10g of added salt a day have an ____x____ factor shorter life.
(2) A study shows that people who drink 2 or more cups of coffee a day have a ____y____ factor lower risk of Alzheimer's Disease.
Then, by looking at x and y, I can tell at a glance whether some result is irrelevant, trivial, useful, or groundbreaking. I understand that it'll still be subjective in the end -- like whether $1, $10, $1000, or $10,000,000 seems like a lot of money to an individual -- but at least it'll be one number.
There are many other ways to accomplish what you're talking about. The biggest problem with your made-up examples is that they are just cases of "bad reporting."
Or it doesn't show whether various correlation factors matter, or whether this is a statistical paradox. Did you know that babies of smokers are healthier than babies of non-smokers of the same weight? This is because baby born to the smoker will have decreased weight because of the smoking, whereas if the baby of the non-smoker is underweight it will be for other reasons that are often worse.
There's heaps of these kind of paradoxes and pitfalls that need to be taken in account.
What we need in newspapers and other media is a simplified abstract of the paper and an explanation or approval of a real life statistician, with no relation to the study.
I used it extensively for the development of different trading bots and algorithms. They say you need to be a real hotshot at maths to make it in that field. Little do they know that this world will soon belong to script-kiddies and hackers! Here's one of my favorites from them, this will rearrange any complex equation and make anything you like the subject:
http://www.wolframalpha.com/widgets/view.jsp?id=4be4308d0f9d...
The one thing you will have to learn is how to represent an equation in text based form, ie to use ^ to signify a power etc...
This will solve the following complex equation (Find both X and Y):
x+y=10, x-y=4
http://www.wolframalpha.com/input/?i=x%2By%3D10%2C+x-y%3D4...
... and thats just scratching the surface. Kids studying maths these days don't know how good they have it. I struggled so much when I was younger.
Maybe this solidifies the fact that I'm a programmer and not a statistician, but I got lost after the 2nd hyperbole in the beginning...but I like a challenge, I may have to read it multiple times until your presentation sinks in though :)