the explanation here left my head spinning
Then here's a simpler explanation:
We need to blend pixel values X and Y.
Normal image-scaling algorithms perform the operation (X+Y)/2, because this is fast.
However, the correct operation is ToNonLinear((ToLinear(X)+ToLinear(Y))/2), because X and Y represent a nonlinear scale, similar to how adding a 60db and 60db sound doesn't result in a 120db sound.