How much of this shredding them isn’t just copyright but rather they don’t want anyone else having this information in their datasets?
Horrendous stewardship of humanities collective knowledge all for profit and the race to have the one god computer to rule them all.
As more time passes it becomes clearer that America’s AI strategy should’ve been a public private partnership where the public owned the datasets and the underlying models and we’d leave the productionizing of LLMs to private businesses