Non-blind n=1 human test can be just as flawed. What you like in a particular scenario is not representative of codecs' overall performance.
Testing with humans requires proper setup and a large sample size (which BTW JPEG XL has done!)
The problem is that these codecs are close enough in performance to be below "noise" level of human judgement.
You will not be able to reliably distinguish q=80 vs q=81 of the same codec, even though there is objectively a difference between them. And you can't lower quality to potato level to make your job easier — that changes the nature of the benchmark.
People also just differ in opinion whether blurry without detail is better than detailed but with blocking or ringing. If you ask people to rate images, stddev of their scores will be pretty large, so wide that scores of objective metrics can fit in 1-sigma.
People also tend to pick "nicer" image, rather than the one that is closer to the original. That's a bit different task than codecs aim for.
Codecs can allocate more or less of the file to color channels, so you can get different conclusions based on e.g. amount of bright red in the image.
So testing is hard. Plenty of pitfalls. Showing a smaller file that "looks the same" is easy, but deceptive.