Web Scrapers Claim to Sell Personal Data on Facebook Users on a Hacker Forum
privacyaffairs.com
privacyaffairs.com
Name Email Location Gender Phone number User ID
So basically, everything I used to be able to get in a phone book. Honestly, at this point all of that information should just be considered public, because it obviously is.
If anything I think people are grappling with the fact that the Internet just makes data scraping and processing possible on a scale previously unimaginable, and that's really what people have an issue with, but I don't think there's a great answer to that. I mean, it's one thing to say the front of my house is public info because anyone can come by and take a picture, but it sure feels different when a high resolution photo (or heck, video feed) can be posted online that is instantly available to billions of people.
Fair, although you could opt out of the phone book. (And I don't think they had location/address, though it's been so long now that I can't remember for sure.)
> I think people are grappling with the fact that the Internet just makes data scraping and processing possible on a scale previously unimaginable
This is it right here. The scale and ease of access are terrifying. It's true that in the olden days, someone could follow me around and write down everywhere I went, everyone I talked to, what stores I went to, my hobbies, and so on. But someone would actually have to do that, and they would have to single me out, and even then the information they collected would be in a notebook, not distributed to virtually every human on the planet.
Now we are all being followed, all the time, and all of that information is available to anyone with almost no cost or effort. This is a sea change, and personally I find it horrifying. There are very, very few people I would trust with that much information. I definitely don't trust the whole world with it.
> And I don't think they had location/address,
> though it's been so long now that I can't remember for sure
Same here. initially i too couldn't remember for sure.Then i remembered the scene from Terminator 2 (1991) in which it looks up Sarah Connor's phone-number and home-address in a phone-book! :-)
(Name was ProPhone 1993, so not really pre 90's) https://youtu.be/yBupNdYe08g?t=1078
There were a few services in the early 2000s doing this, they were called Reverse Grey Pages
And funnily enough, family members of mine have started multiple businesses named A-1 something. :)
you can opt out of Facebook too... and Irish phone books had addresses...
This is not true.
Facebook builds profiles about millions of people who have never had a Facebook account. For example, people who happened to be in the background of a photograph taken by a stranger. Another example: people who installed an app on their phone without knowing that it included a Facebook SDK that was tracking them.
This is nothing new. It's been discussed in public, and even before the U.S. congress.
Personally, I'd love to opt out of Facebook. But I can't. Because I can't log in to my Facebook account, and Facebook ignores my requests for access. I even did the "send in a picture of your government ID" route, and nothing happened. So please inform me how I can opt out of Facebook's data gathering.
https://www.theverge.com/2018/4/11/17225482/facebook-shadow-...
There's a difference between information that government and big tech is scraping and storing, vs information that is publicly available to literally any random person online to scrape.
Both are problems, but those are different discussions and we started with talking about the issue of truly publicly available information. I think that's an interesting topic that merits its own discussion without falling into the surveillance discussion once again.
There is an interesting discussion to be had about what it means for information to be public, since public information is also now _searchable_ information, which never used to be the case. What are the implications for that when it comes to laws about what constitutes "public" information, and social norms on what we share publicly.
But once again instead of discussing that conversation, we're dragged back into the corporate and government surveillance, yet again.
You'd literally have to have the kind of Momma that would dig a hole and hide you in it at birth to really be "off the grid" at this point.
What is weird to me are so many of the responses to my comment are along the lines of "But the real danger now is that all of this data can be correlated with other sources of info, and it's all instantly searchable." It's weird to me because that's the point I was trying to make in my last paragraph, to explain that that really has nothing to do with Facebook. The "hackers" aren't even claiming there was a breach, just information they screen scraped. The ability to amass giant databases of information and make it available to the world to search is something fundamentally inherent to the Internet.
I lived in a lot of different places in the US in the 70s and 80s. All the phone books had residential addresses listed.
This is the format and font I remember: https://groovyhistory.com/content/50602/01af8c322a21e50d0b81...
But... you could tell the phone company not to list your address, and then they wouldn't.
If you wanted to pay a monthly fee for even more privacy, you could have your number unlisted entirely.
They most definitely did have addresses.
To find the phone number of a friend having just "John Smith, USA" would be impossible (or would annoy thousands of people and require months of calls).
Which cost money! Being listed was gratis.
What time you walk your dog each day would be neighborhood level public info. Phone numbers would be city level. For greater reach than that you usually had to put the info out there yourself or be someone of media prominence.
Nowadays the time you walk your dog is out on the internet because it was leaked from some Amazon S3 bucket collecting pings from your dog's smart collar. And what more it's been joined with your name, phone number, and other personal info to create an automated profile of you by interested groups.
That's a whole different ball game, and not one that many people expect despite living their lives (in their minds) the same way as before.
Also I guess user ID means it gives access to their fb profile page I guess? From there one can scrape pics etc (public ones).
With a phonebook (at least back in the day), you didn't risk having your account exposed and sold for a few dollars. Nor did that risk someone getting access to thousands, if not more, bank accounts or whatever through automation. In addition, a phonebook is easy to opt out of making things public. Facebook gives you the illusion that you can opt out of this data being public.
Edit: opt of -> opt out of
What makes you think that that isn't the case here? It sounds like they just scraped all the public profiles they could find. It's not a database "leak" or something like that.
I recall when I still had Facebook, many people would randomly add you and pretend to be random people, sometimes famous, sometimes they "just want to be friends". I know most people have some settings along the lines of "no one except friends can see all this data". This is an issue, because people will add these "friends" and forget about these permissions. These "friends" can easily be attackers like these that suddenly have access to all your data. Hence the "illusion".
With a phonebook on the other hand, when you opt out, you've opted out.
I’ve had FB since 2007. Don’t recall there ever being a surge.
Afterwards my friend's list hovered around 150. These were family, close friends, less close friends, and acquaintances. I also had hundreds on my friend requests lists by the time I deleted my FB around 2015.
But again, the big difference here is that Facebook gives you the illusion that your data is private. With a phone book, when you opt out, you're opted out.
Assuming you had phone books from every city/region in every country. Thats a lot of phone books so you must have had a large warehouse to store them all. Then there is the fact that phone books did not list number for every individual. Multiple persons routinely shared the same number.
The comparison sounds apt in theory but in practice it isn't. Try looking these Facebook users up in the phone books of their respective locales, via the telcos' online phone books or directory assistance. Then, using what you find, tell me their email address, gender and Facebook user ID.
Good luck.
The problem with this argument, "all email addresses are public", which I see regularly on HN, is that information does not become "public" and lose its "private" designation if it is published without consent or lawful purpose. If someone steals secrets and publishes them, they are still secrets.
Whether this information from Facebook is truly "private" I cannot say but I do think it is possible to have email addresses that are not made public.
The recent NSO iMessage story was interesting because the exploit seemed to rely on NSO getting lists of mobile phone numbers for the targets. Not email addresses. Yet iMessage will work without a phone number, with no SIM inserted. Perhaps the targets chose to use phone numbers for iMessage, not email addresses.
Consider what happens if someone creates a Gmail address but never uses it to send mail, and never shares the address with anyone, except Facebook. If this person does not make their Facebook profile public, how is this address public information. Google does not publish a list of every Gmail address. According to the logic of the parent comment, they might just as well. Email addresses are "public", right. Because some HN commenters think they are.
What happened when someone scraped Apple's servers to obtain the email addresses of Apple iPad users. Did federal prosecutors think the information was "public" or "private". The media called the incident "theft of e-mail addresses".^1
From your example, it's another thing to have a high resolution photo or video feed of everyone's houses and to, say, send them ads for painting services if the trim looks out of shape.
I think the important thing to get in the public consciousness is that scale alone is sufficient to make information processing fundamentally different than a human interacting with a single data point. Looking up one person in the phone book and calling them or sending them a letter is different than scanning the entire book, robocalling everyone in it, and sending junk mail to all of them. The fact that the former is accepted and that the later is merely the former repeated a million times does not make the latter permissible. The former was accepted because the way the world worked meant that it was simply intractable - an economic nonstarter, a physical and logical impossibility, humanly infeasible - to abuse it into spamming a million people.
For another example, license plates are public, required to be visible on your vehicle on public roads. Prior to license plate scanning technology, a cop could have tailed a suspect and radioed their vehicle description and license plate to have other detectives and officers disperse to intersections and track a vehicle through a city, and depending on the nature of the problem, they could spend a few hundred dollars to dispatch a helicopter to chase it across the freeway. They could conceivably tail a non-suspect, but that wouldn't make any sense, they were constrained by limited resources to only use this ability for a select few vehicles. That was how the world worked. Later, automated license plate readers were developed. With cameras deployed across every intersection in a city, it would be feasible to track all motions of every vehicle at all times; it would likely be cheaper and easier to do so than one year's expenses of deploying personnel to do so manually.
That information should be considered public, because it obviously is, but what a person is allowed to do with public information should not be limited only by what they're able to do with it.
I tend to think that there should be a publicly accessible, unique, and more or less immutable ID number for every citizen or resident. This ID would have pointers to our name, birth date and a few other identifiers that shouldn't really be considered secret.
My concern is that the absence of such a unique ID leads to a mess of overlapping systems in which only large organizations with the resources to track everyone will be able to uniquely identify people. So we'll have a degree of anonymity from random other individuals, not not from banks, tech corporations or the government. Computing power is becoming too cheap and ubiquitous to effectively hide information that isn't explicitly confidential. That is, as a society we need to adjust to a paradigm in which it is more expensive to keep information confidential than to allow it to be public. Especially keeping information private from those with deep pockets.
AND:
All PII field level data must be encrypted at rest.
UUIDs unlock both the record linking and privacy achievements.
The book Translucent Databases shows how. Hide sensitive data with salt + hash, just like with proper password files. Use UUIDs as opaque pointers for linking.
Agreed.
I remember way back when FB first launched the "feed". Folks on Slashdot (yes, that long ago) had great outcry about how much of a violation of privacy it was. I countered that all it was doing was collating all the posts that people were making on their "walls". Nothing new was necessarily being exposed.
People still didn't like it. Someone argued that the extra steps necessary to visit each "friends" wall was a valuable impediment. Obviously, that's a weak position to take, but it reinforces your point: data scraping is easier than anyone seems to be willing to acknowledge. Anything you write in any "semi-public" space should simply be considered entirely public.
> Name Email Location Gender Phone number User ID
It's never about one item of information released: it's about the aggregation and linking potential. Name/location/phone together form a pretty decent unique identifier. FB obviously gives you friends, interests, hangouts, and most importantly, photographs; none of which you had before.
Ater aggregating with other databases is when the harm comes.
No, we don't. PII is PII.
These large scale breaches harm everyone's privacy (anonymity). Even people who are not included. Because with enough data you can deanon people, eg thru process of elimination.
The flipside of this is that we need to make it such that simply knowing someone's Name, Address, DOB, and SSN is not adequate to fraudulently assume their financial identity and incur debts in their name.
Even if people didn't have unlisted numbers, phone books would allow listing only last name and first initial of one person in the household, without any location data beyond the phone book service area (you could provide more if you wanted to be found), and didn’t include gender.
You were able to opt out of phone books and they also didn't contain email and gender.
Well I consider SSNs public knowledge at this stage. You can reliably dox anyone in the US now and find out their SSNs. Also: I used to have a sticker on my laptop that had my SSN on it, and brought it to conferences, as a PR stunt for my consultancy.
I am honestly shocked at your proposal.
In a real paper book you had a choice not to get your number published.
Have you put any thought about people who are maybe running from abusive spouse or any other people who have reason not to have their location data to be broadcasted to entire world?
Your phone books had your login usernames and emails?
The freely available databases of pii in the wild can be used to infer anything missing from releases like this, and that stuff can be used to inform probabilistic password guesses, and so on. It's only a matter of time before deep learning models make most common password based security measures completely transparent and obsolete.
That's literally the joke.
I looked up Facebook IP address and tried to go there directly, bypassing the DNS. No response.
(I don't care about Facebook much, it was a test. WhatsApp though I have to use to communicate with relatives.)
The first place my mind went to after I read this was "dumb fucks" (Google this with Zuckerberg for context) but as good as the schadenfreude feels it doesn't change the very real and negative impacts from all of this.
I find it hard to find fault with people expressing disgust with how knowingly predatory and exploitative the company's leadership have proven themselves to be.
Exhibit A: https://www.bbc.com/news/technology-58678332
But maybe it's a failure with wider scope than just Facebook's DNS. Or an attack that targeted both FB and HN (and others)? Wild speculation at this point.
The whistle-blower situation is interesting though, I see what you mean that there could be regulation inbound... but how can they be regulated?
This vindicates the stance taken by Signal to not even collect metadata.
Edit: I mean surreptitiously scan the face of a stranger you see in public and the app will tell you about them. Don't know names of the apps.
Note that these services are also powered by scraping public data.
Is the downtime (at the time of writing) their way of blocking a known ongoing attack that can't be stopped fast and safely enough by other means?
Something like: 1) take everything down, 2) fix the bug, 3) deploy everywhere, 4) start everything up.
And, to stop clients from connecting, take down the DNS too. DNS is also a great scapegoat.
Over in WA land, we'd probably just kill all the frontend servers if needed. For all of FB infrastructure, killing all the loadbalancers would probably work, and also the outgoing proxy hosts, as appropriate. No need to mess with DNS or BGP.
It could be the only way to go in the highly unlikely scenario that attackers are able to compromise the management APIs of various infrastructure devices like routers, baremetal servers etc.
These APIs are usually on airgapped networks though, which makes this extremely unlikely
2. Somebody sees this information and records it.
3. They publish this information on another site.
4. "Hackers stole my private data!!!!"
Really?!
But I guarantee they still have my information because people I know use their app.
This is not okay.
Name, yes, but that's public.
I don't have a consistent "real" phone number -- I change it periodically. I then have virtual numbers that redirect to those, which I change from time to time as well.
I'm looking at this
https://www.shouselaw.com/ca/defense/vehicle-code/470b/
and IANAL but it seems that possession of a fake ID of itself isn't illegal and I don't think "using Facebook" falls under the intent to defraud, i.e. "in order to cause loss or damage to a legal, financial, or property right".
It would be a stretch, but perhaps an imaginative prosecutor could claim that the user is trying to diminish Facebook's finances/property by generating web responses (using CPU time and electricity) which it otherwise wouldn't.
That's discounting all the expansive interpretations of the CFAA which would deem any breach of Facebook's terms of service to be an act of illegal hacking, not to mention an interpretation of copyright law that asserted that the user was receiving unlicensed copies of Facebook's intellectual property (e.g. their HTML).
Source: https://twitter.com/WAK4S/status/1444276266362982400/photo/1
Of course it is personal information.
I think what you are trying to say is that you do not consider this data to be sensitive personal information. Emphasis on sensitive.
There seems to be no end to the chaos around the Facebook mafia.
Why do you think it is out of hand, and what exactly do you mean by that? I will have a careful, lengthy discussion with you, if you'd like.
I'll open with the claim that what is moral is not necessarily what is legal.
The evil co. in Ready Player One, the movie, literally imprisoned people and used them for forced labor.
(then, somehow, the CEO gets arrested by regular cops, lol).
Because he had a gun in hand not because of his forced labor practices.