tl;dr ordinary gradient descent is sensitive to the parametrization of the problem but natural gradient is not. This is an important fact (and one that is fairly well-known within the ML community), but it is not totally clear to me why it should be particularly relevant to neuroscience.