Meta won't train AI on Euro posts after all, as watchdogs put their paws down
theregister.com
theregister.com
The only way to achieve that is to make the laws enforceable. They should probably try hard to fine some government organisations as well. The laws though littered with necessary exceptions for government are a step towards respect for privacy.
That's one of the large issues with the EU. Against foreign surveillance and espionage, but internally it's a different story.
I saw a fascinating outcome of this at a conference recently. Polish speaker made an IRS joke to a room full of Europeans at a conference in Europe. Everyone got it.
Honest question. I know they have firewalls, but these are one sided as far as I understand. No?
And how can EU consumers know which of the data they get is US data or EU data? I'm European but almost everything I write online even about my home country, including this comment you're reading, is in English on US developed platforms, not in my mother tongue. Now, is that US data or EU data?
AI is devaluing you as an employee. We are only at the beginning of the AI apocalypse. If things keep going this way, all that is left to meatbags like you or me will be hard manual labour under strict AI supervision. This is not the future we asked for.
I don't know if it is good or bad thing in total. I remember idealising the USA as a teenager, not only its technological might but imagining that everybody lives in large detached homes only later to find out that Americans who live like that can't do anything without a car, don't have groceries, restaurants, cafes, museums, cinemas in walking distance.
My bet is that it will simply stick differently and will create an image of America through which Europeans will reflect on themselves and have conversations about how things can be done differently.
I don't believe that hearing different viewpoints automatically makes you to subscribe to them and that's why I'm very ant-censorship and I think it's a mistake to block or remove content from the internet. Real people don't get influenced the way priests influence villagers in Age of Empires 2.
Actually, I think a bit of cultural decoupling from the US will have positive effects as the US itself is in cultural crisis.
In Poland we have a team of people who work on SpeakLeash which is 1TB of polish texts freely available for anyone to train their models on - http://speakleash.org/en/speakleash-a-k-a-spichlerz-english/
I can imagine models getting their knowledge and intelligence by getting trained on global - english and chinese - data, and then learning country specific languages and viewpoints from way smaller datasets.
You can see that already - gpt was trained on way less Polish language material than english, but it’s almkst just as smart and fluent in both.
As for cultural quirks - here it’s more about who does hrlf and alignment than the source material.
"Meta on Monday said it hoped to use Europeans' data to train its models. It promised to only use public posts and comments — not private chats and DMs — and to not use any content from anyone under the age of 18. Crucially, the biz said it would give Euro folks a chance to opt out; a safeguard not extended to the rest of the world."
Private messages have never been in play here. The regulators are arguing that stuff posted for the whole world to see is not OK to use for training purposes if it was posted to Facebook, but it is OK if it was posted to a blog.
All that data is already being scraped and processed into datasets by many companies: https://research.aimultiple.com/facebook-scraping/
The way I understand it is, the laws are made to cover cases of LLMs being trained on work of authors and artists who normally make a living from the content they create such as articles, books, music, images, videos, etc. whose work is normally copyrighted but LLMs don't give a shit and can just scrape and "learn" your content, style and patterns and then regurgitate it with slight alterations without the original authors getting paid.
It's no secret that some LLMs have been trained on pirated books off libgen.
In Poland we have a whole catalogue of voided agreement paragraphs, that gets updated regularly. It’s a safety mechanism against companies exploiting the legal system to push people into signing things that are considered unfavourable to them - https://uokik.gov.pl/niedozwolone-klauzule
We’ve had banks for example that had bad terms (or trading platforms), and then government step in and make such terms void.
The reasoning is that a consumer/user should expect fair treatment when agreeing to ToS, without spending money on legal advice before clicking “I agree”.
That is, each legally binging copyright assignment needs an enumeration of fields, like radio, tv, streaming and so on, and when a new field appears the copyright belongs to the original author by default. And statements like “in all fields current and future” are not legally binding.
There is no such thing as “people signed copyright to Meta so Meta can use it in whatever fashion”.
You can read an exercept of Meta’s ToS here - these ToS do not allow them to train models based on the content people publish. https://www.quora.com/When-I-post-a-photo-on-Facebook-or-Ins...
By the way, not everything is about money. Convincing people it is, is even a better trick than the devil convincing people he doesn't exist (that is assuming one believes in such things, and then considers the devil to be evil instead of the people ending up in hell, but I digress).
That bog tech gets away with, is not necessarily because it is legal, but rather that people don't care, law makers only start to care and even if people would care, it is nigh impossible to successfully litigate.
None of the above is written in stone, nor should it stop us from doing something about it.
I don't really care about the money but as a society we should care a lot about such a blatant heist of personal effort. For many their words online are their works of the day, it's not at all fair for a company to take that and use it, and even more insidious that it will be used to replace the people who created those works.
That's what the fuss is about I guess, we are working out the ethics our societies will accept on a pioneering front of technology. I don't think Facebook gets away on the letter of the law this time, they will need to adjust to new regulations as we decide on what's ethical.
for professional content creators, that is exactly the biggest problem. That's why the artist community is getting a huge brunt of the pushback. You don't want your content production tools like the Adobe suite to retroactively say "yeah, by the way we are scanning your content as its produced and using that to make more money".
Even if you as a person don't mind, that's a huge NDA breach for a companyu. You don't want AI training to produce assets similar to the in-production media that isn't even out yet. It'd be a huge leak at best and a DOA at worst. And I'm sure you've seen the year+ of debate on copyright issues.
I would be ok with that tweet being findable through search, but if a model is trained on the information and starts spitting out that I’m a cat hater period (without context), then that’s a different thing.
Also, there is an expectation of privacy on platforms like FB or Insta, because the posts require at least login to be read. Because of that people may be posting things that they expect only their friends to read, and friends of friends - not to become a part of a world knowledge canon.
For me personally - on platforms like HN and Reddit, I have no issue with models getting trained on. Fb/insta I treat as an extended group chat and out of limits for cralwers and ml datasets.
Meta: "Privacy is long dead, we always owned your data."
EU: "Hold up a moment, we need to check this before we can agree to this."
Meta: "Ok". Thinking: Eh, whatever. This will happen later anyway.
The AI will be very good at that it seems