Unfortunately they don't provide information regarding their training sets (https://help.mistral.ai/en/articles/347390-does-mistral-ai-c...) but I think it's safe to assume it includes DataComp CommonPool.
China must be laughing.
but its that a breach of GDPR???
Also people who have given their consent before need to be able to revoke it at any point.
idk but how can we do that with GDPR compliance etc???
Edit: that last bit is probably catastrophic thinking. Enforcement has always been precisely enough to cause compliance vs withdrawal from the market.
You can’t steal something and avoid punishment just because you don’t sell in the country where the theft happened.
Tit for tat.
NK isn’t really a business partner in the world.
Edit: After more reading. Clearview AI did exactly this, they ignored all the EU rulings and the UK refused to enforce them. They were fined tens of millions and paid nothing. Stability is now also a UK company that used pi images for training; it seems quite likely they will try to walk that same path given their financial situation. Meta is facing so many fines and lawsuits who knows what it will do. Everyone else will call it cost of business while fighting it every step of the way.
Also note that AI is not just generative models, and generative models don't need to be trained with personal data.
A normal industry would've figured out how to deal with this problem before going public, but AI people don't seem to be all that interested.
I'm sure they'll all cry foul if one of them get hit with a fine and an order to figure out how to fix the mess they've created, but this is what you get when you don't ethics to computer scientists.
China is already dominating AI, you are asking the few companies in the West to stop completely.
The regulation is anti-growth and anti-technology - the GDPR, DSA, Cybersecurity Act and AI Act (and future Chat Control / Online Safety Act equivalent) must be repealed if Europe is to have any hope of a future tech industry.
They have to be able to ask how much (if) data is being used, and how.
Rethinking Machine Unlearning for Large Language Models