Gaussian distributions are monoids and why machine learning experts should care
izbicki.me
izbicki.me
if you combine this property with a bayesian analysis, and put a conjugate prior on the parameters of an exponential family distribution, then the posterior distribution, and the marginal likelihood depend upon on the data only through these sufficient statistics and everything else is easily computed. in this form, one of the sufficient statistics often has an interpretation as a "pseudo-count"; how many effective samples are encoded in your prior?
exponential family distributions include: poisson, exponential, bernoulli, multinomial, gaussian, negative binomial, etc.
This is a interesting exposition and places specific results in a more general algebraic framework; however, the title suggests this is a revolutionary discovery which it isn't
The product of Gaussian random variables is certainly not Gaussian.
Indeed, convolution, the common way of combining functions, is defined as a sum of products at each point.
Edit: I phrased this flippantly, but it's a serious question. My educational background is in Math, specifically Abstract Algebra. Would learning Haskell actually help me use those concepts for practical purposes?
There are "issues" with those structures in GHC, i.e. how the Functor/Applicative/Monoid hierarchy evolved in GHC (applicatives were discovered a little late)
http://stackoverflow.com/questions/7595023/why-is-this-decla...
http://stackoverflow.com/questions/5730270/alternative-imple...
But: (and this one of the beauties of Learn You a Haskell), learning the categories (functor, applicative, monoid through arrows, duality, groups etc), is a superb way to abstract the software analysis/development process and is one of the winning features of languages with good type systems (haskell, the *ML's incl F#, scala)
I'd encourage you to try Python. Its syntax is simple, natural, powerful, and very similar to mathematical notation (I'm thinking of list comprehensions). If you want to use Python for mathematics -- in the sense of an open-source replacement for Maple/Mathematica -- I recommend Sage [1].
http://izbicki.me/blog/wp-content/plugins/optimized-latex/im...
How does he go from (a - b)^2 to a^2 - b^2? What am I missing?
Also, not sure mathematicians wore skinny jeans in 1904.
What does group theory tell me in the context of 1+1 that knowing the answer is 2 doesn't? Generalizations are useful because they apply to more than one context.
That said; the interesting connection between two areas of mathematics is what I took from the article, but I'd agree that his title (well, subtitle) is way overblown - his argument for "Why ML experts should care" seems to boil down to speed, but the Haskell statistics package itself contains faster algorithms that are marked unsafe, and he hasn't demonstrated safety.
In summary: interesting idea but I'm not sure it applies outside a very narrow domain.
The beta distribution is uniquely determined by two parameters called alpha and beta. Typically, these represent the sum of observations we've seen of some events. (On the same website is an example with heads and tails from coin flips.) The beta distribution forms a monoid simply by adding these parameters from each distribution. This is very similar to the Categorical distribution, and also generalizes to the Dirichlet distribution.
Things get really interesting when you stop talking about distributions, and start using them in the context of actual machine learning algorithms. If you know how the Naive Bayes algorithm works, for example, it should be plausible that it forms a monoid in the same way that the Gaussian forms a monoid. The future posts will cover how to apply this to some of the "important" learning algorithms like this.
suppose we have:
x ~ N(0,1)
y|x ~ N(x, 1)
then we have: y ~ N(0, 2)
i.e., gaussians are closed under marginalization.however, i believe gaussians are not the only distribution with this property either: i think this property corresponds to the stable family of distributions: https://en.wikipedia.org/wiki/Stable_distributions
Gaussian distributions belong to the class of stable distributions, though, because of another of their properties: independent Gaussians, when added, are again Gaussian.
infinite divisibility is (yet another) property of gaussians!
Closure under marginalization is something else.
It so happens that the functional form of the gaussian satisfies both, but the two properties are not at all the same.
P1: X, Y gaussian => Z = X+Y gaussian
P2: X, Y gaussian => X | Y gaussian