A few (non-rigourous) examples in ML/statistics are:
- The use of the LASSO estimator [0] (minimizer of |Y - β X|₂ + λ |β|₁) for variable selection is justified by Candes & Tao's 2005 results [1, 2] that show, under fairly general conditions, that this estimate exactly recovers the true set of active variables. A key step in the proof is finding a concentration of measure result for sub-Gaussian random matrices. A corollary is that (with high probability) we can exactly solve the non-convex problem |Y - β X|₂ + λ |β|₀ by passing to the convex relation |Y - β X|₂ + λ |β|₁ and have strong guarantees that our relaxed solution is optimal.
- The Johnson-Lindenstrauss lemma for dimensionality reduction, etc [3] can be proven by essentially the same principle (random projections are "probably" nearly isometries) that Prof Gowers describes in this note.
[0]: http://en.wikipedia.org/wiki/Lasso_regression#Lasso_method
[1]: http://authors.library.caltech.edu/10092/1/CANieeespm08.pdf
[2]: http://arxiv.org/pdf/math/0410542.pdf
[3]: http://en.wikipedia.org/wiki/Johnson%E2%80%93Lindenstrauss_l...