Using nearest-neighbour resampling on the low-res image is an absolute joke. He didn't even look at an objective quality measurement (PSNR). The human visual system is very sensitive to edges, and the high-res image has more pronounced blocking artefacts. Downsampling a high-res image is an unnecessary load on the end user, the 8x8 block transform was chosen for good reason.
I'm not necessarily saying the low-res is superior, but I disagree this ad hoc method is the 'best' way (compared to optimising the coding).