They don't mention what datasets were used. I've come across too many models in the past which gave amazing results because benchmarks leaked into their training data. How are we supposed to verify one of these HuggingFace datasets didn't leak the benchmarks into the training data boosting their results? Did they do any checking of their datasets for leaks? How are we supposed to know this is a legit result?
At this point, it should be standard practice to address this concern. Any model which fails provide good evidence they don't have benchmark leaks, should not be trusted until its datasets can be verified, the methodology can be replicated, or a good, independent, private benchmark can be made and can be used to evaluate the model.