I'm pulling directly from Chris Olah's blog post with that example. But I will say that in practice, its always surprising how increasing the dimensionality of a neural network magically solves all sorts of problems. You could use a kernel if you don't have more computation available, but given more computation adding a dimension is strictly more flexible (and is capable of separating a much wider range of datasets)