Complete Course on Machine Learning
computervisiontalks.com
computervisiontalks.com
It sure helps to understand the integral equations, especially if you want to read the original literature. But realistically you are going to need to understand summing, normalizing, algorithms for clustering, and so on. You probably don't want to write your own numerical code anyway; someone else did it, and they handled all the edge cases that a naive implementation misses.
You can find PDFs of the James, Witten, Hastie, Tibshirani book "An Introduction to Statistical Learning" [1]. Scroll on through - there is nothing intimidating math wise. All the heavy lifting is left to R.
Jump in, the water is fine!
Some good R-specific resources:
http://www-bcf.usc.edu/~gareth/ISL/ https://cran.r-project.org/web/views/MachineLearning.html http://ocw.mit.edu/courses/sloan-school-of-management/15-097...
Another great introduction are the descriptive and inferential statistics courses on Udacity!
I personally went into a graduate-level probabilistic machine learning course with probability knowledge consisting of an undergraduate course that followed Ross http://www.amazon.com/Introduction-Probability-Models-Tenth-... - so there's certainly no need to have been a math major. But if you've never dealt with random variables whatsoever, you'll hit a wall following research from the last 20 years.
With applied machine learning it is certainly possible to quickly get a working knowledge without too much reliance on statistics or difficult theory. You can compare this a bit with using a sorting function without knowing exactly how it works (but you know how fast it is and when to use it).
If you have an engineering background, take a look at the wide array of high-quality ML code and tools. Study trendy and powerful tools like XGBoost.
The dependency would look like:
stats <- measure-theoretic prob. <- math. analysis
The problem with dumping the measure-theoretic probability is that you won't really know what a random variable is. It has a definition (a measurable function into the reals), and without that, you will have a tendency to think of it as "a box that produces something random when you look into it". This will limit your ability to understand papers, and will make you insecure in talking to people.Besides "random variable", other common notions will also be hard to understand without measure-theoretic probability, like "almost surely", convergence concepts, the difference between the SLLN and WLLN, etc.
The problem with dumping analysis is that you will not know some basic things like what a continuous function is. What is everywhere continuous? What is a C1 function? And again, you will have a hard time reading and speaking.
For what it's worth, I found analysis to be not that fun, but measure-theoretic probability to be really a fun, tight, theory. It was enjoyable to learn.
My school's PhD stats program does require real analysis before the prelims, but for most intents and purposes, 'multi' and 'linal' (as the cool kids say) should be sufficient for machine learning from a comp sci perspective.
I haven't fully worked through ESLR (Hastie and Tibsharini's advanced version of ISLR posted above) but the majority of the math there is linear algebra with some differential equations and calculus thrown in. I've heard Harvard Stat 210 and Berkeley Stat 205A/B cited as good examples of mathematical stat classes - if you're seriously interested maybe take a look at those syllabi.
http://www.computervisiontalks.com/variational-methods-for-c...
Here's a great Laboratory on Amazon ML for Human Activity Recognition (w/ Python). https://cloudacademy.com/amazon-web-services/labs/aws-machin...
Totally worth a look.
For probability, "Probability Demystified" is a good basic intro.
For statistics, I would really recommend Allen Downey's Think Stats (http://greenteapress.com/thinkstats2/index.html), especially if you're coming from a programming background. Most introductions to statistics focus heavily on the mathematics needed to enable certain analytical approximations to difficult probabilistic calculations (e.g. the t-test), whereas Think Stats just bites that bullet and focuses on simulation / brute force so you can spend more time on the actual fundamental theory behind statistics.
Brian Blais' "Statistical Inference for Everyone" (http://web.bryant.edu/~bblais/statistical-inference-for-ever...) also looks really good, but haven't had a chance to review it in depth.
If you prefer textbooks, I have heard good things about "Linear Algebra Done Right," [1] but I would not recommend it unless you are "math literate" at an undergraduate level already.
http://ocw.mit.edu/courses/mathematics/18-06-linear-algebra-...