surprisal: how surprised I am when I learn the value of X
Suprisal(x) = -log p(X=x)
entropy: how surprised I expect to be H(p) = 𝔼_X -log p(X)
= ∑_x p(X=x) * -log p(X=x)
cross-entropy: how surprised I expect Bob to be (if Bob's beliefs are q instead of p) H(p,q) = 𝔼_X -log q(X)
= ∑_x p(X=x) * -log q(X=x)
KL divergence: how much *more* surprised I expect Bob to be than me Dkl(p || q) = H(p,q) - H(p,p)
= ∑_x p(X=x) * log p(X=x)/q(X=x)
information gain: how much less surprised I expect Bob to be if he knew that Y=y IG(q|Y=y) = Dkl(q(X|Y=y) || q(X))
mutual information: how much information I expect to gain about X from learning the value of Y I(X;Y) = 𝔼_Y IG(q|Y=y)
𝔼_Y Dkl(q(X|Y=y) || q(X))