My impressions of machine learning match up with yours. There's just so many parameters that there's so many opportunities to capitalize on chance, combined with [necessarily] huge datasets that preclude replication, that it's hard to avoid. I have also been surprised at how much tweaking there is of parameters; I'm used to work in traditional statistics where more is derived from theoretical principles, as opposed to trying a bunch of values to see what "works." This all lends itself to overfitting. I strongly believe that a lot of the adversarial input work is basically capitalizing on this overfitting.
To be fair, I don't think people are necessarily being nefarious, I think people across all fields just don't appreciate the dangers of overfitting.
The one upside of this Many Labs work, mentioned in the paper, is that it tends to show that the common criticism of "how can you generalize from this sample of undergrads" is not so much of a problem. Not that it's not an issue at all, but if you're studying some basic cognitive process it's probably not going to matter that much if you use undergrads versus some perfectly representative sample of the population. People have shown this in different ways before but it's useful to know more about. Obviously with some things sociogeographic variation will matter more though.
One thing that might not be totally obvious from this article and others is that although some effects clearly replicate, and others do not, there are some effects that seem to be in a grey area, where the effects are probably real but much smaller than originally reported. The distribution of effect size estimate distributions is more continuous, with shades of grey, than these news reports would have you think. Whether or not it matters that an effect is tiny versus zero might not practically matter, but at some level it is important to be mindful of.