Google's updated privacy policy states it can use public data to train its AI
engadget.com
engadget.com
Humans aren't property so of course they should be regulated differently from AI.
It's legal for me to go to a nude beach and stare at a topless woman. It's probably legal for me to draw a picture of that topless woman and distribute it. It's definitely not legal for me to take pictures of that topless woman with my phone and post them on the internet.
It's legal for me to overhear a conversation you and your friend are having on a bus. It's legal for me to transcribe what I heard and post it online. In most jurisdictions, it's not legal for me to record that conversation.
Ingesting data for use in machine learning models is still too new to have any specific legislation around it. But the argument that the technology is just doing a thing that humans do has zero relevance.
This is legal. You can take pictures of anyone, nude or not in a public setting and post them anywhere.
>It's legal for me to transcribe what I heard and post it online.
This is murky. It's legal to take notes of what you've heard but that comes with all the pitfalls of hearsay. Legally, it's not treated as the human equivalent of recording because humans have no such equivalent.
If we trace the path of how we ended up here, it's similar to how people incorrectly refer to loci of DNA as genes. We have behavior analysis where we speak of learning as the conditioning via the antecedant-behavior-consequence loop. There was the Hebbian theory of how the ABC loop manifested physically in neurons. Early neural net papers took inspiration from that that mechanism and called it learning.
Meanwhile, actual learning is far, far richer than the Hebbian theory of synaptic strengthening, and has a lot more going on than just operant conditioning.
So, please, it's time for everyone to stop pretending that the fact that ML inherited the word "learn" as a term of art for curve fitting has any philosophical weight.
You can spin up more compute with a credit card. You can't make 1000 people in the same manner, nor can you own them.
Lets be real here though. The only reason anyone is drooling over AI is because it potentially allows one to elide paying someone else, ehich means more money for them.
As usual in these scenarios, the only real injustice is that the people who tried to stand in the way will enjoy the benefits of progress in AI alongside those who worked to make it happen. So it goes, I guess.
I expect the absolute most aggressive enforcement of copyright in this case.
As to my more general assertion of AI only getting the traction it is because of an industry looking to devalue it's currently incredibly highly priced laborers; spend a bit of time around shareholders/management types and you'll soon understand why I think the way I do. Magical thinking, "as long as it makes my outlays lower" thinking is par for the course. There are, in fact, social classes who see the "hired help" as something meant to be out of sight, out of mind, and lucky they get what they are willing to give.
Besides the above hot take, I also see AI as being fundamentally disruptive to the human social fabric. I'm not convinced that as a society we're even prepared to have a real conversation eith regards to a technology that at any time could cross a threshold to sspience. The choruses of such individuals as Carmack and plenty of other HN posters on "it's just a statistical model", and "lets wait til it's at least a developmentally challenged toddler before worrying about those types of concerns" (where those types of questions are those with regard to sapience, and the matter of where the line between "just a statistical model" lay) only proves my point The reductionist viewpoint will be stretched right up to the point that there's a court case where the public finds out that training or instantiating models that communicate with one another basically involves torturing a collective mind that no one bothered to see that way because it was just so stupidly productive.
Hell, the outcome of said case would probably be shifting research in a direction whereby it's possible to make a construct that just barely toes the line. Which misses the entire moral point.
You could say I'm fairly black-pilled on the matter. Humanity can't even deal with one another, or competently raise their own children. We don't need to be committing terrible parenting on an industrial scale.
...If you've read through all of this, you're ptobably a better person than I currently, but know there was a time I shared your attitude toward the subject matter. Then I really started to pay attention to how people treat one another, and how money actually gets earmarked for different things. The learning experience is something I'd not wish on anyone, but as you, and our shared friends the Trafalmadorians say,
So it goes.
It's because a human you can be accountable and can feel consequences, and yes, also because you aren't as fast.
I feel like many people just equivalent "human brain and consciousness" with "neural network" way too quickly. You can't remove the human factor however much you try to equivelate(?) it with a program.
(? = non english speaker)
I still don't see it as an AI failure as a human failure in the use of sophisticated tools.
I might read GPL code once in a while. But I would never copy-paste it when someone asks for, say, how to do a fast inverse square root. (I don't really read AGPL code b/c most AGPL companies I've encountered strike me as salivating for a reason to force a license on a business entity.) The closest I did was looking at GPL code once upon a time for some geo-transform code, and frankly didn't use any of it and instead used a USGS book to re-implement everything in a fully legal-safe way.
If I'm the creator of the work, I get to say that. That I have no means to enforce that is precisely why I've taken all of my work off of the public web.
> There's no realistic way to differentiate web access for human purposes vs. AI purposes in the long run.
Right, which is a very serious problem.
Yes. And there have been numerous court cases confirming this (why we have paparazzi taking topless photos on people on private beaches from public vantage points).
I think this is a feature, not a bug, as without a straightforward rule I don’t know how society solves this without causing more harm. If something is public, it’s public. Restricting you from looking in your neighbors open window isn’t something that can realistically be “fixed.” Other than if I don’t want my neighbor to look in, I draw the blinds. This works for paupers and billionaires.
If I don’t want people to see things, I don’t make them public. I can’t set a limit of “only people making less than $20k are allowed to view this.”
So the crime would be on the leaker, not on Google for training on it.
Unless explicitly declassified the leaked information remains classified, and those who hold security clearances are legally required to avoid all classified information outside of their need-to-know [0].
Now if you have never held a U.S. security clearance, you're less likely to be prosecuted but the history and precedent is murky [1]. The average Joe Sixpack checking out Wikileaks is probably safe, but if I were a journalist publishing the next round of Pentagon Papers I would much rather have a small army of lawyers and a friend or two in Congress.
EDIT - Google has federal contracts, so they are probably bound by similar agreements to at least make an effort to avoid any such leaks in their training data for public-facing models.
[0] https://www.csmonitor.com/USA/Foreign-Policy/2010/1207/US-to...
[1] https://www.npr.org/sections/thetwo-way/2017/03/22/521009791...
So I didn’t view Wikileaks or any material as I don’t want to lose my clearance.
But the question was about whether it’s legal or not to view leaked material. Security clearances are a different matter and are going a step beyond what’s legal or admissible in court or what someone would be prosecuted for.
In the US. It is different in the EU. I know HN is US-centric, but tech is global and there are more people in the EU than the US.
Why would population number be the determining factor on anything? The world obviously doesn't operate based on direct democracy. India having 1.4 billion people doesn't give it a greater ability to dictate anything vs the EU, US or China. It comes down to power (always will, always has). For example, the US - for now - has the ability to dictate certain things to China on trade restrictions, given the US technology advantage. China is a legitimate superpower economically, has four times the population, and yet the US can still do that.
If the premise is consumer numbers: the US still has the EU beat even with fewer consumers, with a far larger, far more valuable economy.
Besides all of that, naturally each country (or as a union), to the extent it can, will attempt to set its own rules for tech. The US will do so, the EU will do so, China will do so, etc.
> Yes
Not in my country at least, but if you will I van suggest another exemple: does broadcasting a movie or a song through to the public make it public good? No. We have laws (highly variable between jurisdictions) that set the rules for these “published data”, and they are actually dependent on manu factors including whether or not you're making money out of it (you can lend a DVD to your friend for free, but you cannot start a DVD-renting business without a licence from the copyright owners).
Movies and songs are copyrighted and can’t be rebroadcast or copied without license. But viewing is perfectly legal and does not require a license.
And you can certainly start a DVD-renting business without a license from copyright holders (assuming you bought the DVD). You can’t start a dvd streaming business without a license.
That's exactly the point: “publishing” something only gives you some limited rights (an in certain jurisdictions like mine, most of these rights are limited to individuals only, and organizations are excluded).
> And you can certainly start a DVD-renting business without a license from copyright holders (assuming you bought the DVD)
Not in my country, again.
Publishing are rights for the publisher, not the viewer. As a viewer, I can view the material and don’t have restrictions on whether I can remember it or not (ie, run it through an algorithm to train a model as part of the viewing).
France. You can re-sell the DVD all you want, but you cannot lend them for money (because it means sharing the copyrighted material) without permission.
> As a viewer, I can view the material and don’t have restrictions on whether I can remember it or not
In m'y country you even have legal right to record it (=run it through an algorithm that “trains a model”, be it an overfitted one) for your own use, if you're an individual but not an organisation. And even for an individual, you can remember every lines of the movie, but if you make a new one with the same dialogues, then it's plagiarism.
Making something public (aka: publishing it) doesn't mean giving up all rights on it.
this is illegal almost everywhere.
e.g. in germany it is even illegal to take a picture of a person (one specific) in a public space without their consent (unless they are a person of public interest).
US [0] “ In the United States, photographs that are taken for editorial use in a public place generally enjoy Constitutional protection under the right of free speech.”
Denmark [1] “ you can almost without restrictions shoot anything as long as what you're seeing is visible from public property. You are allowed to shoot people, including police officers or other government officials.”
In Germany [2], you can photograph people from public locations but not if they are nude or vulnerable or in their home. “ You can’t take photos of people if it shows their helplessness.1 For example, you can’t take photos of accident victims, drunk people or nude people without their permission.”
[0] https://www.hg.org/legal-articles/what-are-the-laws-regardin...
[1] https://law.photography/law/street-photography-laws-in-denma...
My point isn’t that there are no limitations on photography or use. My point is that if you make something public, people can view it and use it. And that’s legal. People seem confused about this that somehow consent is required for use. Not for things publicly released. (In the US and many countries at least)
Not how I learned it. You can take pictures of public spaces as long as a specific person is not the focus of the picture. The other aspects, like the invasion of privacy when a person is nude, only come on top of that.
The law around questions like this is not definitive, so a lot depends on recent court decisions.
What I try to do is track to the actual source rather than relying on my own memory of things.
> According to German law and court decisions, it seems quite definitive.
That's impossible. Basically nothing is definitive in the german Medienrecht ;) Ok, not really true, but it's true that things can change and that it is a less defined area and you'd have to be really certain to know the relevant case law to be almost certain here.
Here is is: [2] https://allaboutberlin.com/guides/photography-laws-germany
And it actually references the same article you posted.
> Nach dem Bundesverfassungsgericht (BverfGE NJW 2000, 1021) ist bereits ab diesem Punkt [that a photo is taken] ein Kontrollverlust der ohne Einwilligung abgebildeten Person über das Bild gegeben, und dieser mögliche Kontrollverlust rechtfertigt sogar unter Umständen ein Fotografierverbot.
But that is debatable and actually the point that could be outdated.
In practice, the guide is not completely wrong: You can take pictures in public spaces and they can show people, but those people should not be the focus of the image and you might want to make them unidentifiable if you publish the image.
Taking a photo of a public square with a lot of people in it is still legal without asking everyone for consent. But the topless photo example would not be legal at all without consent or even just a photo where one specific person or small group is the focus.
Edit: taking the photo itself might be ok in the latter case, but publishing it (like posting to facebook or a blog or forum) is not.
******Warning****** NSFW pictures of the future Queen of England at this URL --> https://theoutsidersadi.wordpress.com/2012/09/14/click-here-...
The (UK) press get around the law, by reporting a story and then relying on the reader using search engines to get the information from other jurisdictions like I have just done here, located in UK, EU google servers, and where ever the wordpress server is located.
It kind of makes me think, what is the point of law?
You need a lot of money to fight these entities and fight them in multiple jurisdictions where relevant laws exist.
What also makes a mockery of the legal system at least the Royal lawyers is, whilst Kate Middleton has some injunction to block press publication in the UK, they cant stop the search engines from publishing the data to users in the UK, as I have just demonstrated with the link above and the search term "kate topless holiday photo" and then clicking the images option!
Now what if I meant "kate beckinsale topless photo" instead of "kate topless photo"? I just got someone elses topless photos without even expecting them, ie the future Queen of England.
Are the Royal lawyers from the stone age, do they not understand search engines or do they buy the targeted filter bubble narrative highlighted by Eli Parsier in his Ted talk?
I like German privacy, they even stood up to Google and the Streetview project, however Google have edited me from their Streetview images where I'm giving their car the bird as it drove past me so there is some human oversight, but they still have that data and refuse to hand it over via GDPR DSAR requests.
To many big entities including the Police, fulfil GDPR DSAR requests by relying on being able to identify the individual using todays existing systems.
If someone cant be identified, like my giving the Google Streetview car the bird in their streetview data, Google will say no data exists.
Yet GDPR law doesnt address developing and future technology which will be able to identify me giving the Google Streetview car the bird if they run facial recognition over their streetview data.
So the GDPR DSAR is useless law as is, although I havent read it, but I suspect Google's privacy policy is as well.
That ends up being a prohibition on taking photos in public at all except in areas devoid of others as it is impossible/unreasonable to get the written consent of all parties who would be photographed in the background of a photo of a person who requested thier photo be taken....
Quite a vulnerable place humans are in at present.
13 years ago, we treated cookssource.com like they were rubes, but they were just too early and too small.
"The web is considered 'public domain'"
https://illadore.livejournal.com/30674.html
https://news.ycombinator.com/item?id=1868736
-----
A Follow-Up to "The Web is Public Domain"
https://web.archive.org/web/20101112141752/http://www.cookss...
Language model -> AI model features -> product and features
It just looks like their lawyers tidying up the words but the intention is unchanged.
Put it behind a robots.txt or better yet a login.
This battle was fought and won. I’m not going back to the bad old days.
A bit like I will gladly build a bike for someone, but object to the bike being sold around the corner.
Or are you less concerned with their design and more concerned with their purpose? Providing citations is fine, but creating content is not?
How do you Wikipedia being used for products period, let alone used as data products, since that’s also repurposed unpaid labor?
Search engines help connect people to the things they want. At least they mostly do.
AI inserts itself in the middle and strips the creators of credit, recognition or income.
That's what I've done along with a growing number of others -- but I'd much prefer to be able to make it available to the general public. I mourn that's not possible without also aiding the training of AI.
Discussion from yesterday: https://news.ycombinator.com/item?id=36577626
My personal web site that is essentially a doc page on numerous home lab projects and other technical writings, is going dark.
I’m not training your new search algorithms so you can directly profit from my writings. I didn’t seek payments prior to this, but now this is just plagiarizing my work.
See ya
Now I'm making a policy that I'm allowed to (ab)use Googles services however I want. They can find the policy themselves. God I hate Google.