The DCs giving you a 1/8 downscale is well known, but the interesting part to me was this generalizes so you can get 1/4 (resp. 1/2) downscales by using only the 2x2 (resp. 4x4) block of LF coefficients.
In almost every case, downscaling-while-decoding is going to beat decode-then-scale on performance - merely keeping the uncompressed image in RAM is going to cost you, when you can scale it down in L1 or even vector registers.