Information Geometry
math.ucr.edu
math.ucr.edu
E.g., in this text:
https://www.google.com/books/edition/The_Minimum_Description...
[1] https://en.wikipedia.org/wiki/Kalman_filter#Information_filt...
tangent:
One application of information geometry pops up when trying to do gradient ascent(descent) to maximise(minimise) some function over a space of probability distributions. Differential geometry can can define a more natural distance metric to use when performing gradient ascent(descent) on the parameter space of probability distributions. Incorporating the geometry of the space you are optimising over into the gradient can lead to a more effective optimisation procedure.
For an example machine learning application, see how the natural gradient is mentioned as part of the construction of a stochastic gradient descent algorithm for online variational inference for LDA [1].
To get a bit of intuition about this, blog posts by Andy Miller [2] & Nick Foti [3] are worth a read.
In variational Bayesian inference, to make calculations more tractable you replace the true posterior probability distribution with some approximating distribution that is computationally easier to work with. Then you need to search for an approximating distribution in your space of possible approximations that is the best approximation to the true posterior distribution. A standard way to measure the distance or "error" between two distributions is the symmetrised KL divergence.
One common way to setup variational inference seems to be to prove that there is a function, an "evidence lower bound", and that minimising approximation error - the KL divergence between the approximate distribution and the true distribution - is equivalent to maximising that evidence lower bound.
If your approximate distribution has a bunch of parameters then you can regard the evidence lower bound as a function of those parameters, and you're in a setting where (if you cannot compute the maxima analytically) you might want to do nonlinear gradient ascent to find a local maxima.
When performing gradient descent through this parametrised space of approximating distributions, the definition of the gradient relies on taking a small step in parameter space. If our approximating distribution is a multivariate normal, the parameters are a vector containing the the mean and variance. An easy definition of distance in parameter space between two sets of parameters could be to take the euclidean distance between them. But the euclidean distance between two sets of parameters for normal distributions gives very little control over the distance between two normal distributions in the sense of symmetrised KL divergence. So it's not a great distance metric for parameter space. This "easy" but not very accurate definition of distance leads to a standard gradient ascent algorithm using the euclidean definition of the gradient.
What could be a better definition of distance in parameter space? Now we're doing differential geometry, so I'll quote Foti's explanation:
> In Euclidean space with an orthonormal basis G(\phi) is simply the identity matrix. When \Phi is a space of parameters of probability distributions and the symmetrized KL divergence is used to measure the distance between distributions then G(\phi) turns out to be the Fisher information matrix. This arises from placing an inner product on the statistical manifold of log-probability densities.
where G(\phi) is the Riemannian metric tensor.[1] Hoffman, Blei (2010) Online Learning for Latent Dirichlet Allocation -- https://proceedings.neurips.cc/paper/2010/file/71f6278d140af...
[2] https://andymiller.github.io/2016/10/02/natural_gradient_bbv...