Think Stats: Probability and Statistics for Programmers
greenteapress.com
greenteapress.com
http://statland.org/MAAFIXED.PDF
The first link is a BEAUTIFUL and thought-provoking discussion of what's dangerous about having a new mathematics professor choose the statistics textbooks for an introductory college class in statistics, with advice on what to look for in statistics textbooks.
http://escholarship.org/uc/item/6hb3k0nz
The second link is by a very famous statistician, with discussion of how the statistics curriculum could be revised to better emphasize the most important ideas.
Put the key ideas from these two resources in your own words, and you will have a good guide to programmers about how to think about statistics.
http://www.amazon.com/Statistics-Learning-Presence-Variation...
which I bought just more than three years ago after I read that. The textbook is quite thought-provoking, not just another Tweedledum to the usual Tweedledee of undergraduate statistics textbooks.
Statistics and linear algebra really should be required by all CS programs. It's funny that at many schools those courses are not, yet Calculus is. First, Calculus should have been handled in HS. Second, I've never had a use for Calculus professionally or for anything I've worked on in my free time.
- Machine Learning by Tom M Mitchell http://www.cs.cmu.edu/~tom/mlbook.html
For general reading and introductions I also like:
- Pattern Classification by Richard Duda
- Pattern Recognition and Machine Learning by Christopher Bishop
For a bit more emphasis on statistics and math, I usually dive in to
- Classification,Parameter Estimation and State Estimation by van der Heijden
And last, but certainly not least:
- Information Theory, Inference, and Learning Algorithms by David MacKay, available here:
I've read O'Reilly's Collective Intelligence. It's a great introductory survey, but it was very light on theory.
I also own Collective Intelligence in Action. It had more explanation of theory than O'Reilly's offering, but most of the chapters devolved into how to use Java data mining framework X.
[1] http://en.wikipedia.org/wiki/Z-test
[2] http://en.wikipedia.org/wiki/Statistical_hypothesis_testing
If you know enough probability theory, statistics is just a special case. The nice thing about using probability theory is if you do decide to use a 'test', all of your assumptions are put forth first. As E.T. Jaynes says:
In estimating a location parameter, for example, the sample median M is often cited as
a more robust estimator than the sample mean. But here it is obvious that this
‘robustness’ is bought at the price of insensitivity to much of the relevant
information in the data. Many different data sets all have the same
median; the values above or below the sample median may be moved about arbitrarily
without affecting the estimate. Yet those data values surely contain information
highly relevant to the question being asked, and all this is lost. We would have
thought that the whole purpose of data analysis is to extract all the information
we can from the data.
Thus, while we agree that robust/resistant properties may be desirable in some cases,
we think it important to emphasize their cost in performance. In the literature,
ad hoc procedures have been advocated on no more grounds than that they are ‘robust’
or ‘resistant’, with no mention of the quality of the inference they deliver, much
less any comparison of performance with alternative methods; yet alternative methods
such as Bayesian ones are criticized on grounds of lack of robustness,
without any supporting factual evidence.
A recent probability book I've started that I think is pretty good is http://uncertainty.stat.cmu.edu/:(
Right now all versions have the same content, but I will continue to revise, so you can think of the version on Green Tea Press as the draft of the second edition.
I know it can be confusing, but I hope the benefits of the free license make up for it.