Let's say you, a human, were given access to a ridiculously-large trove of almost-working software; do you believe you would be unable to learn to program correctly? (Related: would you even need to look at much of that software before you were able to code well?)
I am extremely confident that I am better than almost all of the code I learned to program with. If nothing else, someone out there must have written the best version of some particular function, and they didn't get to see a better version beforehand.
When I look at intro programming books now, I consider many of the examples sufficiently flawed that I tell people I am teaching who are using these books "well, don't do that... I guess the author doesn't understand why that's a problem :/".
And yet, somehow, despite learning from a bunch of bad examples, humans learn to become good. Hell: a human can then go off and work alone in the woods improving their craft and become better--even amazing!--given no further examples as training data.
To me, that is why I have such little fear of these models. People look at them and are all "omg they are so intelligent" and yet they generate an average of what they are given rather than the best of what they are given: this tech is, thereby, seemingly, a dead end for actual intelligence.
If these models were ever to become truly intelligent, they should--easily!--be able to output something much better than what they were given, and it doesn't even seem like that's on the roadmap given how much fear people have over contamination of the training data set.
If you actually believe that we'll be able to see truly intelligent AI any time in the near future, I will thereby claim it just won't matter how much of the data out there is bullshit, because an actually-intelligent being can still learn and improve under such conditions.