>500m Facebook records were leaked and <10m records have an email address — far more records are phone number but no email.
>500m Facebook records were leaked and <10m records have an email address — far more records are phone number but no email.
It's also an interesting choice for a search tool for a quasi-public dataset.
Discussion here: https://news.ycombinator.com/item?id=26702473
Facebook then allows the person collecting phone numbers to search for the name of the user associated with a phone number (or email address).
So I'm very interested about any insight you may have.
--
https://boards.4channel.org/g/thread/80982465#p80984158
> Anonymous 04/04/21(Sun)19:14:02 No.80984158▶>>80985592
> >>80984122
> That's because it's the austrian dump file. There's an update with the australian dump in the original raidforums thread, it's just missing from the pastebin index
--
https://boards.4channel.org/g/thread/80973470#p80985270
> Anonymous 04/04/21(Sun)21:15:35 No.80985270▶
> >>80985097
> they were mixed as one file. leaker was australian
--
I obviously have no citations of my own for this anecdata.
Regarding the first post, I went through approximately ~80 pages of "thanks", and one pair of mildly sore eyes later, was unable to find any "update". A torrent *was* posted somewhere around page 80 but this had the wrongly-named file in it. I have also seen a few TG groups that have renamed the file, but it's still the wrong one (like the GP described).
I can confirm that Austraiaia has an eclectic mix of Australia and Austria though - and also that grepping for 'Australia.*Austria\|Austria.*Australia' returns 170 results, grepping for 'Austria' returns 446375 results, and grepping for 'Australia' returns 408 results.
As far as seeing what was leaked, you could find the data yourself (but I'm uncomfortable giving instruction on how to get it). It would be nice if it was possible to extend the tool to be able to send the information for a number to the number (because that's the only way I can think of that demonstrates ownership of the data) but that can't be done for free and it therefore seems too complicated to set up.
Eagerly awaiting the news that Facebook takes security seriously.
No interest in seeing other people's info, but I want to know if I'm affected.
After that, you'll find that the data is poorly and inconsistently encoded (lots of ugly BOMs), a bunch of files are split in weird ways (that I had to concatenate then give sensible names to). Figuring out what order they go in (they're not all split on a record boundary!) may take some translation, too. After that, a bunch of them have bad, broken CVS headers and you have to find something that can manage huge files to edit the first couple of lines.
Then you find that all of them use different delimiters (colon or comma), some of them quote each entry, and they each have a different set of data that doesn't match up at all. I'm still trying to figure out what the data in each file is and see if it can be normalized into one schema.
So it's about like downloading a phonebook for half the world with some extra mystery data that has no description where you have to guess what it even means.
Interestingly, I checked for a family member who deleted their FB a few years ago and they're not in there. It'd be interesting to see if there are any accounts that deleted before 2019 in the data or not. I suspect not, but this isn't enough of a test to prove it.
This is not particularly high quality data, I'm sure a lot of people didn't look at the data very much and just ignored a few lines of errors on import or did simple greps as you say because it's a pain in the rear to find editors that even can handle multi-GB files, etc. Even good converters like iconv can bomb out on, I think it was Qatar (which was also little endian UTF-16 with a BOM), etc.
But yes, it's all over the place. I love how one of them decided to number things 1, 2, III, 4, though maybe that was the translator, I dunno, I never learned to read Arabic letters...
Because yeah, let's just store a lot of data with commas in a CSV file, that won't make it obnoxious to parse or anything...
Uh, what?
There were links to a ufile folder with all the zips on here yesterday.
I took a peek, relieved I wasn't in there.
The pastebin with the links got taken down, although there's an archive.is mirror of it. The forum where it's from also has torrent magnet links.
>I took a peek, relieved I wasn't in there.
AFAIK it was publicly scraped data combined with the exploit to get phone numbers[1]. If you didn't have a public profile you're probably safe.
Unless you hire an attorney that knows these laws... because not even attorneys can know all laws.
I would recommend people check for themselves.