A Programmer's Guide to Data Mining
guidetodatamining.com
guidetodatamining.com
onto the book - it looks promising for an intro to recommendation systems. no opinion about classification yet. doesn't appear to have anything on graphs or network effects which is somewhat disappointing. that being said i need to review bayesian stuff / teach myself some of the harder stuff and it will be nice to have a practical walkthrough.
that being said no one should be implementing these themselves (except the dumb stuff like distance metrics).. it's useful to learn but scikit-learn is amazing when it comes to fancy algorithms.
http://radimrehurek.com/gensim/
It has good implementations of various algorithms, some of which support streaming or dirstribution, and it allows loading and dumping data in various formats.
I've used it for building content based recommender using tf-idf, lsi and similarity index. After the index is built, queries to it are really fast. It can handle quite large corpuses with little memory.
The reason for that is a pretty epic list of dependencies (have fun explaining why the prod boxes need a fortran compiler), but in terms of efficiency and speed of development it's an obvious choice.
Hopefully the SciPy & BLAS dependencies will only get easier to install from now on... Continuum Analytics received shit loads of money and some of it is going towards better scientific Python packaging, I believe.
Thank you, Ron Zacharski!
(disclaimer: you do not want my opinion regarding any topic).
Porting Python code can be painful. (I checked the chapter 7 py file and it isn't filled with functional style code, though various kinds of arrays with index starting with 1 or so may still be an issue)
(And yes, I intend to go through this book re-writing for PHP...)
Thank you for doing this.
Ever heard of PEP8 for Python coding style? List comprehensions?
I'm afraid this falls in no man's land, with code too weak for practitioners and theory too weak for theoreticians.