One concern I immediately have is overfitting, particularly for claims about how various difficult values have been optimized to be the "best possible". It looks like the parameter space in use is truly enormous and so it would be very easy to come up with hypotheses that perform fantastically on your dataset but terribly in real life. This seems like it would be a first-order concern, while the ability to run tests in a single day seems second-order if those tests are producing garbage outputs.