you massively collapsed what AI companies have been doing by comparing it to old internet-scraping. Facebook flat-out admitted that they scanned copyrighted books for their AI. The image generators most definitely trained on copyrighted images.
OpenAI et al also stole everything from everyone. But then they raised billions of dollars from that data and sell back their LLM to people (again, among other things). They are also very much NOT open in any way, aside from sharing their benchmarks of new models.
Also my understanding was they’re not storing the actual music, but the metadata and a link to the YouTube video.