A lot of the problem is in the questions they're designed to answer versus the questions people use them to answer.
For example, if I'm comparing Python and C, I typically want to know "how much slower would my program be in Python?", not "how much slower is my program in Python if I spent so much time hyper-optimizing it that I might as well have written it in C?"
But the test cases usually try to answer the latter, not the former.