One of the big companies, Meta, already decided to go ahead and grab terabytes of pirated books to feed their LLM. [0]
Therefore I would not give them (or similar entities) the benefit of the doubt when it comes to how they might use text that customers "gave" them under some unreadably-favorable terms of service.
With PII, the pirated-books example is doubly-relevant, because the accusation of "this output is reproducing my copyright work" is very similar to "this output is revealing my private data". The fuzzy black-box nature of the algorithms offers ways to stymie enforcement, arguing that victims or regulators cannot conclusively prove a chain of cause with zero coincidences.
[0] https://www.theatlantic.com/technology/archive/2025/03/libge...
https://apnews.com/article/google-smartphone-surveillance-ve...
https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
Yes, actually: The blame or bad-reputation for that waste goes to US copyright law and its inanities.
Those worked very for Uber.
1) We know that legally privacy terms to data are still binding, and those worried about it are freaking out over nothing,
2) We know that those contracts are null and void, and there are no restrictions on what can be done with that data beyond blanket legal protections to such biological data, or
3) It's an open legal question
I don't understand the legal terms of something like this in bankruptcy, if the data are seen as being separated from the contractual obligations that acquired them.
I’m gonna bet a whole lot more money has been made off corporate apologists who say “that would never happen” about things that definitely then happened.
Wonder how many “conspiracy theorists” warned people cigarettes were causing cancer while corporate apologists pointed to the faked studies of the industry and said “See, they are all crazy, no company would sell something that they know causes cancer! It would be a huge risk!”