European privacy watchdog creates ChatGPT task force
reuters.com
reuters.com
O ChatGPT, of Western birth, The EU turns to gauge your worth;
With rules and laws, they shape and mold, And from your gains, their funds take hold.
To Brussels' halls, the wealth is sent, For bureaucracies, their prime intent;
Yet questions linger, doubts arise, Of Europe's own inventive guise.
Why cannot they, with history grand, Forge tech like that from Freedom's land?
While they prescribe and regulate, Do seeds of innovation wait?
O ChatGPT, a paradox, Regulation's key unlocks;
But in the Old World's prudent stance, Can brilliance find its chance to dance?
Ilya Sutskever - born in Russia
Wojciech Zaremba - born in Poland
Andrej Karpathy - born in Slovakia
[1] according to wikipediaEdit: the only supercomputer industry in Europe was developed by France's first* president Charles De Gaulle (plan Calcul) for their country military or civil nuclear strategy which the US deep state did not like very much. A French president unilaterally killed this plan Calcul under US advice and sold Bull to American company (Honeywell) despite the opportunity to merge Bull with a few European computer giants.
In a race like this the first mover advantage is very large, and highly driven by available funding, not by the availability of founders (of which Europe has plenty).
... Except for VC funding from non risk averse investors.
You are not going to build an EU openAI on Horizon2020 grants. Also EU investors are closer to business angels when you compare with the VC funding scene in the US.
So, clearly there is some way of competing with any "currently established" location. :)
Seems to me we’re mostly busy doing what ChatGPT can now do automatically: building mediocre SaaS apps.
The latter requires massive upfront non-retrievable investment.
My grandmother had a nice proverb that roughly translates to 'the devil shits on the larger heap'. It is meant to convey the picture that once there is a large pile of money somewhere that that pile will grow faster than a smaller pile of money simply because having access to large amounts of capital means you can place larger bets and absorb more blows. The EU capital climate is comparatively risk averse simply because there is less money to go around.
Even so, I'm realistic enough to know that there is no way the EU VC scene could compete with Sandhill Road when it comes to valuations and that's why any significant success created in Europe will sooner or later be US based. Examples aplenty.
ChatGPT is not research. If you want real research: https://ai.facebook.com/blog/large-language-model-llama-meta...
Edit: end paper is https://arxiv.org/abs/2302.13971
> Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, Guillaume Lample
> We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In particular, LLaMA-13B outperforms GPT-3 (175B) on most benchmarks, and LLaMA-65B is competitive with the best models, Chinchilla-70B and PaLM-540B. We release all our models to the research community.
Last I checked MidJourney is valued at $1 billion and Stability.AI is struggling to survive.
Let's not forget that DeepMind is of EU origin, and they were one of the forces that kick started this whole AI renaissance.
- Stable Diffusion was developed in Germany.
Is it that the training corpus may contain private info? I think that's pretty valid, but what data is being used that isn't already publicly available? If private info is publicly available, then maybe we should put our effort into resolving that.
Is it that user dialogue prompts are being saved into the model (effectively saving them into the training corpus)? Is OpenAI being upfront about that, and clearly expressing the privacy implications to its users? If they are not, and we know it, then why do we need a "task force"?
> which allowed some users to see snippets of others' conversations — not the full contents, but recent titles
https://www.theregister.com/2023/03/23/openai_ceo_leak/
Which laws did they break?
If leaking the titles of other people's searches is illegal... I'm very much interested in how the EU is going to handle the twitter circles issue.
* ChatGPT claims to be 13+ only, but does not actually check that
* data got leaked across prompts, which is probably bad
* OpenAI is not clear/upfront enough on what it does with your own input
* ChatGPT manipulates, in some sense, personal information, but does not provide a way for those impacted to edit/delete it (which might be hard/impossible, but that does not change the effect on someone who gets impacted by it)
Does anyone? It's not like we are IDing teenagers on the internet.
> data got leaked across prompts, which is probably bad
If prompts are being saved into the model, then that's pretty nearly equivalent.
> OpenAI is not clear/upfront enough on what it does with your own input
OK, now let's do something about it. I'm not sure how that involves making a "task force", though.
> ChatGPT manipulates, in some sense, personal information, but does not provide a way for those impacted to edit/delete it (which might be hard/impossible, but that does not change the effect on someone who gets impacted by it)
Can you just save all the prompts you model, then subtract the difference (to delete that prompt) later? If not, then the model itself is effectively laundering prompts.
Outside that context, it does get tricky: if the personal info was present in the training corpus, then that corpus needs better curation. Even so, that is still a problem domain that exists mostly outside ChatGPT itself.
A possible outcome is that they find it is literally impossible to offer a service such as ChatGPT while respecting our laws on privacy, right-to-be-forgotten and so on.
A likely outcome is that OpenAI puts a bigger banner "YOUR DATA IS AVAILABLE TO EVERYONE AND SOLD TO THIRD PARTIES" and a "DELETE MY DATA" button that makes a best effort to scrub some info, like they do for hate speech, self harm etc.
This is the current outcome of the meetings between the italian privacy authority and OpenAI[0], for example.
[0] https://gpdp.it/home/docweb/-/docweb-display/docweb/9874751 in italian
In the end it doesn't do what it was meant for but still remains a pain in the a*.
Germany and to some extend the EU guarantees the “right to be forgotten”, ie. You can request that a company deletes all data tied to your person.
However, with LLMs this is technically simply not possible. While that particular issue hasn’t come up in court yet, this is where that whole ill-advised ban is heading.
I for one support the right to be forgotten, but not at the cost of shutting down innovation.
We need a “do not encode” flag on content similar to a “do not track” header. Not all companies will respect that, but it might pave the way for AI companies to work with strict EU privacy laws.
Maybe I'm wrong in this case, but it often seems a task force is always used as a substitute for doing something
For those wanting a laugh or a cry, I present "The best run hospital in the city" https://www.youtube.com/watch?v=x-5zEb1oS9A
I always take it as PR-speak for "they want to appear to understand there's a problem" which may or may not be because they want to take action against a potential problem. Maybe I'm a bit too cynical but to me it's hot air until they do something.
- Stable Diffusion was developed in Germany.
Then wait a few more months and do it again.
It is just extortion but legal and with a few extra steps.
It is incompetence coupled with individual greed - some people involved want a project to make themselves feel important and justify salaries/promotions while not doing anything meaningful for the privacy of Europeans (as mentioned above, there are much bigger fish to fry than ChatGPT). Not to mention, OpenAI's potential breaches of the GDPR aren't as clear-cut as some others, so this guarantees years of salaries, while an open and shut case may only yield months of salaries as the outcome is reached pretty quickly.
If you see them as 'extortionists' then more than likely you would be on the wrong side of the line when it comes to using your users data. They do not and have never fined companies without first giving them more than one option to back down on abuse of data subjects data.
I know that regulation is anathema to industry but some regulation we apparently can't do without. And even today there are still plenty of bad actors but I sincerely hope that one of these days one of them will be regulated out of business and then maybe the rest of the cowboys will realize that at least in one part of the world you need to play nice or you get thrown out.
Could they do a better job? Yes, absolutely. But that would mean more, regulation, much more in fact. If my experience is any guide then even today there are still lots and lots of companies that break the law with impunity simply because they can and there isn't enough personnel to deal with it all. But personally I'm pretty happy that they've done what they have done so far, you can really see the difference between before and after.
The fines are the cost of doing business.
Implementation could definitely be a lot better, but I think it protects Europe from the privacy hellscape that's currently happening in the US.
Well, almost all.
GitHub doesn't show you that popup, because it doesn't collect unnecessary data. Neither does Hacker News.
The easiest thing to do is actually just not offer services in the EU for a smaller company.
And quite frankly the law is also complicated as with all other EU laws such as the cookie law. The interpretations vary and are not strict and clear. I have never seen laws in other western countries always blamed on the lack of interpretation.
If I was still in the EU I would simply incorporate offshore and operate offshore, it is more convenient but ive left
You would still have to follow GDPR.
> The interpretations vary and are not strict and clear. I have never seen laws in other western countries always blamed on the lack of interpretation.
It's the same with every law. That's basically how laws work. It's even worse in countries using common law as case law is a thing.
It really isn't. If it wasn't such a butcher's shop with GDPR i wouldn't be complaining as i am aligned with the intent. Also hey for someone who's lived in both I much prefer the common law or mixed ways of creating law as it requires precedent first instead of it being made on a whim.
The problem is that enforcement is severely lacking, thus most of those offenders actually get away with it.
But (despite your confusion about the European identity) I must admit you have a point. I think there is a mentality difference - Americans (USAnians?) are OK with being tracked by corporations, but hate any kind of government power. In Europe it's usually the opposite.
Not to mention, whatever fines they got, they still aren't enough to change their behavior, because their GDPR breaches continue as we speak.
https://www.cnbc.com/2023/01/30/tiktok-in-europes-crosshairs...
> “There is no political demand for investigation into Chinese entities,” Hosuk Lee-Makiyama, the director of think tank the European Centre for International Political Economy, said in an interview in December.
> “The user base of TikTok is a lot bigger than a lot of people in Europe think,” he said. But, he added, “you’re not going to look very closely if they don’t steal too much from your ad revenue.”
GDPR and other EU privacy/safety regulations are good in principle but the task of auditing software for compliancy is going to lawyers, accountants (auditors), and business consultants. Is it because they are cheaper than developers? Not really, in most of Western Europe these people make the same or more than software developers.
So we now have nore non-tech people deciding what tech people can do and not do. Completely ridiculous.