I should emphasize that this is not a nitpick or even a criticism, just a feature I would love to see. It's also what I spend a large portion of my time trying to track down, so having it in a convenient location would be nice.
I should emphasize that this is not a nitpick or even a criticism, just a feature I would love to see. It's also what I spend a large portion of my time trying to track down, so having it in a convenient location would be nice.
http://udel.edu/~mcdonald/statintro.html
I point a lot of newbs to pages on that site so that they can develop a better intuition for the methods.
I wish I had this last semester during my statistics course.
It's written for biologists, but you don't really need to know much biology to work through the examples, and the focus is inherently practical.
A small amount of correlation can lead to very biased estimates of significance.
Correlation between observations will generally appear in any time series data or data that is arranged spatially.
Again, just a rule of thumb, but you should be very wary of the lack of independence between observations.
No statistical technique is assumption free, unless it is purely descriptive.
Some of them are free of explicit assumptions known by the practitioner, but that's not the same thing. In much the same way, my code is all bug-free.
machine learning techniques, which tend to be assumption-free
ML should be a rigorous exercise in Bayesian and classical/frequentist stats, computational methods, dataset integrity, visualization etc, if you've been thru the texts by Murphy or Bishop. It often happens that people a couple years out of their last stats class only retain that high R-squared, p-, t- and f-values are what they're looking for, and heteroskedasticity and sphericity are just big words.My evidence that ML is a rigorous exercise: the free texts listed (Barber, Mackay and Smola's are excellent, ESL not as accessible)
http://metaoptimize.com/qa/questions/186/good-freely-availab...