Their solution basically just amounts of "Ethically sourced Styles" which still has all the red tape that a normal text2image model has because majority of the data is still unapproved for use in an AI model.
Businesses didn't want to get wrapped up in a pesudolegal model that really has no better legality than base SD.
Music conglomerates have money and their lawsuits will probably settle the issue.(unless they settle) That will be applied for all copyrighted works, regardless of the medium.
I believe going against the big guys is the reason why the big ones don't yet have music generation LLMs.
No one is stealing anything. It's not theft. There has been no crime. None of this is anywhere near criminal law.
I could make a more nuanced argument on copyright infringement. But to make that steelman, I'd need to accept a too large overton window shift, so I'll decline to do so here.
The problem is that you're putting it in the wrong legal framing, and it just won't fly. Willing to engage, but not on these terms.