Factorization Machines
tech.adroll.com
tech.adroll.com
EDIT: <strike>It's</strike> Kernel tricks in general are nice because they also generalise well to an infinite dimensional space (the RBF kernel) and compute the dot-products in that space without actually computing the embedding. RBFs: https://en.wikipedia.org/wiki/Radial_basis_function_kernel
What I meant to say was that you didn't need to compute the embedding explicitly. Since you embed into a space that has a nice structure, you can compute the dot product of the embedded vectors without having to compute the embedding explicitly.
The kernel trick uses an implicit mapping into a higher-dimensional feature-space. On the other hand, Factorization Machines uses an explicit mapping into the polynomial kernel space. However it learns jointly the "right" polynom, by mapping the base features into a low-dimensional dense space where the higher-order terms of the polynom are dot-products (we're not taking dot-products in the original feature-space!). FM then learns the right polynomial (a non-convex taks) jointly with learning the original supervised learning task.
Where I am not with you is in the hint that kernels have feature maps that are necessarily opaque. In fact inhomogeneous poly kernels ar great examples where the feature map is known in closed form. Although such a map is available it is not always efficient to use it. But what it lets you do is optimize the weights of each dimension on tje mapped feature space but executed in the native space. In fact, and I am greatly tickled by the coincidence, I had sent a mail to colleagues where I suggested a form where instead of the standard Euclidean dot product in the poly kernel is replaced by a convex combination of other favourite kernel of choice. The training algo remains almost the same. I was not aware of this piece of work but the mechanism is pretty much the same (little more general. Essentially replaces the std dot product that appear in the kernel expression by other kernel expressions. If you are thinking recursion now well that's intended) and can be pushed further.
http://www.ismll.uni-hildesheim.de/pub/pdfs/Rendle2010FM.pdf
In particular they're nice for recommendation algorithms.
D is awesome for what we do with it. It unlocks our small team of data scientists to have a huge impact on the company, as we don't have to rely on frameworks such as your typical hadoop/spark cluster which require engineers just to maintain it. We can also use the same code path offline and online where we have very strong latency and QPS requirements.
D makes us ship more often, with less humans. The runtime is not perfect yet, but the language and compiler are awesome. It keeps our codebase small, simple and very efficient.
To answer your question, FM is good for two main reasons. First, as a result of some mathematical magic, we can train the model in a time that scales linearly with the time needed to train a model without interaction terms. Second, each of the n features gets embedded in an inner product space with similar features ending up close to one another in some sense. This allows us to make a decent estimate of the interaction between features even if they do not appear frequently in the training data.