Machine Learning for Hackers
shop.oreilly.com
shop.oreilly.com
Would-be brogrammers will find these books, read them, and graduate from cat-picture projects into more sophisticated applications that can solve real computer science problems. It matters less whether the readers of these books actually get a comprehensive understanding of the theoretical knowledge behind them, then it does that they enrich readers and spread the desire to learn and attain true understanding of the field.
Steve Yegge reference: http://www.youtube.com/watch?v=vKmQW_Nkfk8
PS. Attended Machine Learning class in hacker dojo with the author, he's a bright guy. Hopefully the book will be as good.
Hopefully that means they're getting ready for the other classes to go live as well.
"Due to popular demand, we are teaching a follow-up class: AI for Robotics at www.udacity.com . Also due to popular demand, we now have a programming environment, so you can develop and test software. Our goal is to teach you to program a self-driving car in 7 weeks. This is a topic very close to my heart, and I am eager to share it with you. (This class builds on the concepts in ai-class, but ai-class is NOT required)."
Preface
Machine Learning for Hackers: Email
How This Book is Organized
Conventions Used in This Book
Using Code Examples
How to Contact Us
Using R
R for Machine Learning
Further Reading on R
Data Exploration
Exploration vs. Confirmation
What is Data?
Inferring the Types of Columns in Your Data
Inferring Meaning
Numeric Summaries
Means, Medians, and Modes
Quantiles
Standard Deviations and Variances
Exploratory Data Visualization
Visualizing the Relationships between Columns
Classification: Spam Filtering
This or That: Binary Classification
Moving Gently into Conditional Probability
Writing Our First Bayesian Spam Classifier
Ranking: Priority Inbox
How Do You Sort Something When You Don’t Know the Order?
Ordering Email Messages by Priority
Writing a Priority Inbox
Works Cited
Books
Articles
About the Authorsunfortunately I see here, that there seems to be a large part of the book spent on a statistics introduction (including R) and only one machine learning algorithm actually gets introduced. and on top of all it's one of the most simple ones (naive bayes). I expected at least some further description of support vector machines or other advanced techniques.
http://www.amazon.com/Programming-Collective-Intelligence-Bu...
So while it's certainly an interesting field, I wonder how many hackers are really going to need these skills.
Sure, prediction APIs could arise that give detailed use cases for each algorithm, but then there's a problem with the fringe cases: you might not know that two pieces of data are so heavily correlated that they completely shatter a conditional independence assumption, for example.
As a hacker who originally subscribed to the belief that a thorough understanding of machine learning was overkill, it is without hesitation that I admit being 100% wrong. The truth of the matter is that when it's done properly, artificial intelligence and machine learning ought to be inextricably linked with your core business processes.
1. Derivation of slight variations on the basic principles
2. Scaling.
Both are very difficult.Sure we did some basic implementations in octave, it helps to have some idea of the internals. But that wasn't the goal of the course.
My prediction is that even the most black box ML, creatively applied, is and will be an incredible skill. Increasing levels of sophistication will continually kill off the current practices of black box ML, but the willingness to apply statistical pattern recognition to new and interesting areas can't help but be incredible.
But that said, the deeper message is in interpretation and discovery from data. Large data, small data, highly structured data, or just regularized DB pulls. The heart of it is statistical pattern recognition and it's really just begun to be broached (even academically) in the last 25 years.
The problem facing people who intend to work with data that does not yet exist becomes one of feature selection: what data matters and how do we use it? For NLP tasks, does stemming matter? What about part-of-speech tagging? Some classification problems are not linearly separable, which makes certain kernel methods impossible without using (and knowing to use) the kernel trick.
In the end, I think my reply here is tautological: ML is too complex to be transformed into a set of APIs a la Google Maps and Google Search.
Engineering even a basic ML solution is challenging---feature engineering especially.
The algorithms are not disclosed, but the docs hint that they are properly regularized so throwing more features at them is always good.
You still need to be able to reformulate the problems so that they fit a standard ML setting and then know how to tune things, but it looks like the API can get you pretty far.
Indeed, I worked on machine learning in NLP (fluency ranking, parse disambiguation). As a general rule, roughly 90% of the improvement of models is in clever feature engineering and exploiting the underlying system to get more interesting information that improves classification, 10% you get from using more clever machine learning techniques than, say a standard maxent learner with a gaussian prior (for linearly separable data).
For instance, the last relatively large boosts of the accuracy of the parser developed by our research group came from feature engineering: