It's on OpenAI and Anthropic to prove that they obtained these results legitimately and credited all researchers who deserve credit. They do not get the benefit of the doubt.
There's clear benefit in a babelfish that can coordinate disparate efforts, the only problem with the current iteration is giving credit to said efforts.
Google went quite far down the road to hell, but stopped short of taking credit for websites' content since the company understood that poisoning the well only goes so far. At this point, one can safely conclude that _Chat_GPT was an intentional attempt to squeeze out more data once they mined the internet dry.
Obviously they stole them from the Proof Fairy.
It's grapes so sour they could etch metal.