* It's complicated so it takes a while and you need lawyers and such to make it right
* Rules for training are probably hugely vague and undefined. Because you could ingest personal data and it cannot be deleted
* AFAIK it needs to be hosted in Europe (not directly GDPR related, but america has laws that allows them to spy on all traffic in the US, so this is somewhat the counter to that)
In the end from my experience just working at a company that needs to be compliant this usually means:
* All the services need to be hosed in EU including 3rd parties we send any data to
* There needs to be a way (email is enough) to delete user data (including from 3rd parties which need an endpoint so you can trigger it from your side)
* You need to inform the user about the data useage and allow them to opt out of the "usage" of this data for non-essential things (i.e marketing emails). This does not mean you cannot save this data if you also use it for other things, but you can not use it for the non-essential case.
* You could be in trouble if you save data "just because" and do not use it for anything essential or if it is not transparent to the user.
Not a lawyer. Just the things I notice in my day to day. In the end companies need data protection professionals to navigate these things. Which is probably another thing a startup does not worry about it early on.
Of course you might not care, if you believe that the EU has no way of forcing you to pay a fine and if you are certain that you are never going to do any business in the EU. But in such case you might as well provide access to your models for EU IPs too. It makes no difference.
The model access makes no difference.
"Personal data" and "data processing" are (deliberately) drawn very broadly:
"""Personal data — Personal data is any information that relates to an individual who can be directly or indirectly identified. Names and email addresses are obviously personal data. Location information, ethnicity, gender, biometric data, religious beliefs, web cookies, and political opinions can also be personal data. Pseudonymous data can also fall under the definition if it’s relatively easy to ID someone from it.
Data processing — Any action performed on data, whether automated or manual. The examples cited in the text include collecting, recording, organizing, structuring, storing, using, erasing… so basically anything."""
And, unlike the arguments about copyright in big AI models trained on the internet (is it 'fair use'? Don't ask me, IANAL!), the requirement for explicit and informed consent is something a general crawl will very clearly fail:
"""Purpose limitation — You must process data for the legitimate purposes specified explicitly to the data subject when you collected it."""
Furthermore, we don't know enough about how the models store knowledge/beliefs to be able to make any claim about accuracy:
"""Accuracy — You must keep personal data accurate and up to date."""
And as for confidentiality… for downloadable models, that's "by obscurity" only, due to the exact same research needed to resolve the previous point about accuracy, and even for secret models like GPT-4, nobody's really sure how to actually guarantee it won't leak info with the right prompt, and there's even some suggestion that this is actually impossible with current approaches because nothing is really deleted by RLHF:
"""Integrity and confidentiality — Processing must be done in such a way as to ensure appropriate security, integrity, and confidentiality (e.g. by using encryption)."""
In the case of #1, they probably do not have an agreement with the "data controller" in the case of scraping, which means #2 is a violation of GDPR.
IANAL.
As a European I to try see it from both sides, consumer protections are generally a good thing, but it right now being restricted by EU vagueness sucks ass because I just want to play with the cool new toys.
What is OpenAI GPT4 and Google bard/Gemini EU version not doing so they work in EU but the latest Google AI is doing so google is incapable of putting it in EU ?
Maybe latest one is more invasive with ypur personal account? Scanning your personal data without consent ?
Because seems to be a super simple business, user sends you a prompt, you run the prompt and send the result back, you keep nothing unless necessary for the service to work and give the user the ability to purge their history/data .
Google is in a completely different situation. Gemini is, so far, a niche product, making little to no money for Google. However, due to how the GDPR works, a mistake with Gemini could cost the entire company dearly, impacting the profits they make from any other products.
This is a more general principle. If you're a small startup that manufactures foos, and there's a risk that manufacturing Foos infringes on Apple's patents, well, that's just one of the risks you have to bear as a startup. If you get wiped out by a lawsuit, you're established as a limited-liability corporation for a reason, if the fine is larger than the worth of your company, you can just declare bankruptcy and go away. You can't decide to stop manufacturing Foos, as that's your entire business.
If you're the size of Google, getting into the Foo-manufacturing business is not worth it. If Apple sues you, that won't just impact the Foo-manufacturing side of your business, the fine can be significantly larger than any profits you ever hoped to make from the venture, not to mention the brand damage, strain on resources etc. If Foo manufacturing is ultimately going to be a side business for you, it's probably not worth it to be the hill you die on.
This is why it makes sense for startups to "move fast and break things™", and for big corporations to require countless legal reviews on the smallest of decisions.
I wouldn't be surprised if there's a long and arduous procedure to clear a new product for GDPR compliance, and the PM had a choice between launching now and excluding the EU or launching later for everybody. This is also what originally happened with Bard.
I'm glad that Microsoft / OpenAI is confident enough to put out a global consumer facing gen AI offering. Maybe they received EU assurances, or maybe they have enough legal muscle to tank any possible hit, and eating Google's lunch is a prize worth the risk.
Lots of companies come unstuck because they fall into the trap of “let’s just collect everything and see what we can do with it”.
Or, I’ve got all this data I’ve collected legitimately. Who knew that you could sell it on to some data broken and make loads of money - let’s do that!
Or, I’ve collected all this data, I’m just going to keep it hanging around, oops I just put it on a public bucket and leaked it all. Hmm, I’m not even sure what data we had, have we just compromised a bunch of people? Who knows…
It's not supprising that many smaller companies are saying fuck this noise.
If by design you can do whatever pleases you, then yes you have a lot of innovation. But sometimes it leads to normalisation of troubles (e.g. data leaks in the US), and incredulity of the general public ("how did we even get there")?
There are good reasons to ponder ethics in the original balance too, it hasn't got to be completely paralysing either. But this comes with a cost (e.g. typically, for any data-sensitive work these days complying with GDPR, a significant part of the design & implementation time is "are we compliant").
They could have easily made law requiring sites accept DNT header but they didn't likely because of lobbying.
DNT is not relevant as GDPR is not directly a regulation against tracking, and it certainly isn't because of lobbying.
> If you want to comply it is easy.
That's the entire antithesis of modern law as opposed to monarchy. Law should be codified in as clear rules as possible.
If I'm to venture a guess, it's probably because data protections are stronger and they want to avoid potential issues should someone test GDPR (or whatever the applicable law is) by asking specific data be removed from the model
What is the downvote coming from, isn't this just facts? If not for the regulation, why would EU be shunned?
They are literally doing the opposite because they are asking you to commit to the terms and since everyone do you have actually consented to your data being used.
AI Act isn't solving anything that isn't already solved with existing regulation.
That so many people on HN seem to think this is a good idea is very puzzling.