DeleteFB: Selenium script to delete all of your Facebook wall posts
github.com
github.com
I'd be slightly more confident (slightly) that editing the post might cause the core data in the db to be updated, however. In which case, I think the more effective script would be one that goes through all of your FB posts and scrambles them, or replaces the text with gibberish.
- Facebook publicly claims it does [1]
- Mark Zuckerberg testified in-front of congress and stated they do [2]
- Multiple government regulators have specifically checked that it does in their privacy audits. [3]
They might be lying and actively conspiring to not delete it, but they'd have to have a very good reason to take on that much legal risk.
Now, consider:
- A infinitesimal fraction of facebook users try to delete anything.
- Facebook makes money from your data by showing you ads. If you stop using facebook, you stop generating any revenue and your data becomes a liability, not an asset.
What reason would there be to lie about it?
[1] https://facebook.com/help/250563911970368
[2] https://www.washingtonpost.com/news/the-switch/wp/2018/04/10...
I just want to point out that "testified in-front of Congress" seems to have no relation to the truth of any statement. I don't recall the last time (if ever) there have been legal ramifications for lying in front of Congress.
Zuckerberg also testified that shadow profiles don't exist and that users can always remove their personal information from Facebook (they do exist, and you can't). James Clapper testified that the NSA doesn't "wittingly" surveil hundreds of millions of Americans (they do, very "wittingly").
https://techcrunch.com/2018/04/11/is-zuckerberg-under-oath-c...
John M Poindexter was one who was convicted of lying to Congress: https://en.wikipedia.org/wiki/John_Poindexter#U.S._executive...
[1] https://www.economist.com/united-states/2018/12/13/michael-c...
It's one of the reasons Cohen is in prison, isn't it?
It is illegal to lie to Congress (unrelated to whether you're under oath or not), but I can think of very few examples where there were real ramifications for it.
Sure you're not thinking of Manafort? I don't think Cohen had a plea deal, and I don't think he was charged with lying to the FBI. Manafort and Flynn did, and Manafort I think was the one that then broke the deal. There's a lot of crimes to keep track of.
I guess it depends on what you mean by real ramifications as it's being served concurrently but Cohen is serving two months for lying to Congress, a charge he plead guilty to.
They technically wouldn't be lying if they only deleted the record from application systems, but retained it in historical data by simply setting a bit flag (eg "IsDeleted") to 1.
The question about what motive they have comes down to whether or not the types of posts a person deletes indicate something meaningful about their personality. Ad analytics are ultimately seeking to understand the kind of person you are, after all.
"In determining appropriate retention periods for personal information, data controllers can have due regard to any statutory obligations to retain data. However, if the purpose for which the information was obtained has ceased and the personal information is no longer required for that purpose, the data must be deleted or disposed of in a secure manner. Full and irrevocable anonymisation would achieve the same objective." (Page 69, under 3.4 Retention)
So basically, as long as it's anonymised, the data can be retained.
Under this reading, is there even an obligation to anonymize? And how anonymous does anonymous need to be? Does a simple base64 encoding count as “anonymized”?
And the matter of anonymisation is a great question - "irrevocable" anonymisation certainly has a different meaning back in 2011, when swapping a name for a guid would do the job. Nowadays, it would require at least deleting all relationships as well, since social network analysis is much more advanced these days (especially at FB, of all places.) It wouldn't be impossible to derive the identity of an anonymous account/record based on the undeleted data associated with it. And since FB's "ghost profiles" are something we know exist, I think it's safe to assume those relationships are being maintained somehow.
Also, I'm now looking forward to Zuck's next apology for "a breach of trust" and his explanation of how they "failed to live up to their own standards" when it becomes publicly undeniable that their "Full and irrevocable anonymisation" of data they've claimed is deleted is as flawed as all the other attempts we've seen of doing that.
FTFY
#1 You can train the model continously
#2 the model stops learning
#3 it has a submodel which detects you are feeding it false data.
Given that people change over time, #2 seems unlikely. The question to me is "how advanced is #3?". Another thing I wonder about is how relative are these to eaxh other. At odds?
#2 makes no sense from a business perspective.
#3 is also well-established. This is how Google's Captchas work, for example.
If my friends or acquaintances continue to use facebook, then information about me continues to be useful to facebook as that information can be used to complete their model of what sort of people my friends are, that they would associate with somebody like me.
I wouldn't trust Zuckerberg to pour a glass of water on a burning orphan, let alone delete data for real. Nor would I trust any politician to know anything about the subject.
It wouldn't be the first time Facebook (nor any other business that profits from people's data) have been deceptive while treading a careful line between honesty and lies.
edit: I should add that if they were keeping archived records in a separate business unit, there is the possibility that they keep the original or even all edits. So scrambling your data might not do much aside confusing those who have you in their feed.
Given FB's size and the small amount of people deleting data, I doubt Facebook would take the risk to keep a few more datasets.
Granted GDPR is supposed to catch businesses that pull those kind of stunts, you have to remember that Facebook do already break GDPR in number of public ways too. So it's pretty clear they have a relatively open interpretation of the regulations (and an army of lawyers who are confident they can proceed in such a way). Or it might just be the case that even the worst fine issued by the GDPR is worth the risk given the financial benefits awarded to Facebook for retaining data.
This part at least is incorrect. Your data will always be an asset, and FB makes money not just by showing you ads but by selling your data.
And they obviously give away some data for free, especially pertaining to users who click through OAuth consent screens as in the Cambridge Analytica case: https://developers.facebook.com/docs/graph-api
What are some examples of transactions where they exchanged data for money?
That's not a great analogy because I would be violating other people's IP, whereas FB generally has rights to the data they collect.
I am not a fan of Facebook but this entire thread is just filled with misinformation.
It would’ve surprising if they didn’t capitalize on them.
I work at a company you've heard of, and you've probably (P > .5) used at least one of our products. We're international (including EU). I personally wrote our user data deletion logic for GDPR compliance.
We delete all of your PII. It doesn't matter if you're from the EU or not, because it's too expensive to figure it out and too risky because you're going to miss some weird edge case - Legal doesn't have much of a sense of humor when it comes to wiggle room.
My experience has been that larger organizations consistently spend more effort (relative to size) to genuinely comply with privacy regulations than smaller ones. The risk:reward ratio for deliberately ignoring or subverting privacy regulations is insanely bad. There are too many surfaces along which that would leak out, and the gains would be pretty marginal. Pretty much anyone who works at a large tech company can and will confirm that this is the case (see Jeff Kaufman's posts on this thread). There is no conspiracy between the six-digit number of engineers who work at these companies to keep quiet about it.
I wouldn't say "conspiracy" but they've been mums-the-word about things like shadow-profiles. Perhaps, it comes down to too much kool-aid but to say that engineers would initially speak-up in such cases has been proven wrong, time and again.
Take the Snowden revelations, as an example: How many years were the programs in service before Snowden went public? How long have we known about shadow profiles and no one from Facebook has come forward to say, "yes, this is what they're doing and it's wrong"?
Relying on people to do the good that should be done by the organisations doesn't take into consideration that those engineers face severe penalties for "going pubic" about such things - namely because whistle-blower laws do not supercede such thing as NDAs.
And then, what about all the data copied into Hive? Copied to 3rd parties? Does Facebook go and delete data that 3rd parties took?
>>> They "trust me"
>>> Dumb fucks
https://en.wikiquote.org/wiki/Mark_ZuckerbergHe's the same old dirt bag he ever was. Here are some fresh stories to back that up:
https://www.theguardian.com/technology/2018/nov/21/facebook-...
https://www.vox.com/2018/11/14/18096008/facebook-zuckerberg-...
All of Zuckerberg's and Facebook's actions, however, continue to suggest that his approach hasn't changed.
Honestly the vibe's still there if you pay him/fb enough attention. I'm working on extracting myself completely from them (Still got a few group messages on upcoming events to clear)
This is the most profound insight we have into the mind of Mark Zuckerberg and what he thinks of his user base.
Stop listening to the PR and examine his actions since then. You can see how that comment was not merely a youthful indiscretion, but his entire business.
Zuckerberg continues to display this exact same mentality, only now he has billions of dollars in resources.
I doubt this
> Mark Zuckerberg testified in-front of congress and stated they do He would tell us everything that helps him as he is an opportunist
> Multiple government regulators have specifically checked that it does in their privacy audits You're talking about the governments secretly spying on us and lying about that since decades? Ouch.
There is no legal risk for those who observe if you abide the laws.
> A infinitesimal fraction of facebook users try to delete anything What does that tell you about the majority of ppl?
> Facebook makes money from your data by showing you ads. If you stop using facebook, you stop generating any revenue and your data becomes a liability, not an asset. That's just what you think you know. What about psychological profiling, law enforcement etc.?
Also did you hear about the "shadow profiles" about users who don't have an account? The moment you're surfing into their net (as we know even through embedded like buttons etc.) you'll be milked like a dairy cow.
> What reason would there be to lie about it? What reason is there not to be honest about it? The recent events have shown very clearly that they don't have to fear much.
Or they can just claim, again, that they accidentally forgot to delete copies of these posts from their backups. It just completely slipped their mind.
I'd be much less confident, considering that they explicitly save and make public post's edit history.
That sounds hilarious. You don't happen to remember or have a source for that, do you? Would love to read more.
However at Facebook's size, and given they're known for 0 privacy, they likely track all changes anyway.
I don't think editing a Facebook post first would have the desired effect.
The README states the intent is to clean up publicly facing content. It's meant to tidy up internet presence to the general viewing public, not escape the grip of FB's data vacume.
I'm a big fan a Selenium, and this is great usage of it! Scripting a boring, repetitive browser task, that would take a large amount of time & effort to do manually.
All the big companies are doing immutable, append-only event logging and probably have no mechanism to expunge this data. All because storage is cheap and they need to hold on to everything for testing or whatever future need that might arise.
What a pitiful dystopia we're living in.
This is not true at all. Here's a recent HN discussion that talks about how Google handles this: https://news.ycombinator.com/item?id=19809259
(Disclosure: I work for Google, but I'm speaking for myself)
(I would love to see someone subpoena something deleted, say, 1y ago and write up whether it was produced.)
Not sure how many actually did that, but to be compliant append only does not fly. I was on a product that did exactly that for data lakes.
[1] http://www.digitalpreservation.gov/meetings/documents/storag...
Don't be confident. They have [1], or used to, post edit history. Even if the feature isn't there in the UI, I bet they didn't remove it internally.
[1]: DDG: facebook post edit history.
https://datacenterfrontier.com/inside-facebooks-blu-ray-cold...
It's text you wrote, why would there be any problem "covering up" your own text?
Note that, I'm not arguing you should have a right to use Selenium to scrape their site and replace the text, but if you went through every post in a non-automated fashion and changed the text, from a perspective of "covering up" the text, the purpose would be the same. Facebook almost certainly has an interest in keeping 100M users from scraping with Selenium, that's a separate thing entirely.
It’s true that editing could serve to hide it from other users, but privacy settings could do that also.
This is information that you originally put out, wrote down on Facebook, so it's not that you're hiding something you didn't want anyone to know originally. If it was libelous or something so you wanted to delete it after the fact, it's still likely possible to get it back -- even if Facebook deleted the actual message in the database, there's likely a number of methods -- server logs, db logs, cached resources, db backups, or even a simple screenshot -- that would likely be sufficient proof.
http://www.wassom.com/delete-at-your-own-risk-spoliation-of-...
As far as the rest of your post, i believe that is exactly what I was asserting originally.
If FB has this data squirrelled away somewhere in cold storage then in order for anyone to harm you with it, FB has to admit to keeping it (which would be a scandal) and it has to get out into the wild. That is a reduction in risk compared to someone just looking at your profile.
Don't let the perfect be the enemy of the good. Delete what you can, then shut down your account.
UPDATE users SET is_deleted = true WHERE user_id = ?The eerie part? It used to show faded thumbnails of my photos that I had deleted more than a year ago in Photos section when prompting me to add my pics.
[0] https://github.com/x89/Shreddit
[1] https://www.reddit.com/r/blog/comments/1dhw2j/reddits_privac...
FWIW, I believe this is still true of reddit. Deleting a post doesn't actually delete it. It just no longer shows on the webpage. To really remove it from the database, you need to edit the post/comment, replace the contents with something simple (e.g., the character "a"), save it, then delete it.
It’s really good.
Seed with: "Why does my dog speak poor English?"
> It's possible your dog has a rare genetic condition called dyslexia, which puts him or her at risk for learning difficulties. If the problem has progressed beyond a certain point—say, if your dog has been in the hospital for over a year—and when you go to bring him out of the hospital—not for exercise or exercise training—he may not speak English in the proper way. Your veterinarian will work on the problem while you are in the hospital.
What should I do with a dog with dyslexia?
It's important to be vigilant and familiar with what is going on with your dog in hospital. If you find that your dog may be learning, take him home and have him trained as soon as possible. It may take a while for your dog to learn to associate the letters of English with words.
There are times when we need to ask ourselves, are our interests really served by having this ability, when so many of us are now doing so much online anyway?
Facebook has many years of pretty smart work, and probably safe to say tens of millions of dollars just in meetings about specifically this (re costs associated with effort of user data retention).
... similar to the outcome of a ransomware attack (gets backed up, archived, overwrites old legit data): brilliant!
This script would have saved me some manual labour back then.
(Edit: we also used to endorse each other on LinkedIn for the most bizarre skills we could find. Toilets, divorce and animal husbandry to recall a few)
My friend had a profile picture in which he was doing something silly with his eyebrows. I commented "Nice eyebrows" and he replied "You're just jealous." I've since deleted my account and now there is only one comment on this profile picture, by the person the picture is of, reading "You're just jealous."
Yikes, sounds like it.
* need to retain data to fulfill government requests
* internal auditing
* it's all backed up in some "data lake" somewhere to do internal ml or analytics on
* hundreds of copies in database backups from different times
* internal logs that contain the data
* it's already been analyzed and aggregated into learning products and models that aren't going to be recomputed
It's not being "paranoid". As someone who has worked on large scale saas, I say: there is zero, 0, ZERO, 0.00 chance of that data every actually being deleted
Unless that is built by design. I happen to also work on a large scale SaaS where we take this stuff very seriously and I can say it is possible to protect this data. However I will agree that this adds considerable complexity, but for some organizations, that is totally worth it.
> need to retain data to fulfill government requests
That's a choice, not a requirement. If you encrypt the data and purposely don't store the keys yourself but instead have the customer store them, then you don't have anything of value for the government.
> internal auditing
Personally Identifiable Info is not something we want to peruse. In fact we purposely don't want to see it because that eliminates a potential for mishandling.
> it's all backed up in some "data lake" somewhere to do internal ml or analytics on
That kind of application shouldn't give carte blanche to disregard retention policies. You can run those applications against replicated shards of the original data; and when the original gets reclaimed, so does the replica.
> hundreds of copies in database backups from different times
Storing useless data forever is not cheap, especially at scale. Better store what needs to be stored and free up what can be freed when retention policies kick in (or user requests it).
> internal logs that contain the data
That's ground for failing certain compliance audits. Logs should never contain PII in the first place, that's an operational failure.
> it's already been analyzed and aggregated into learning products and models that aren't going to be recomputed
That's a tricky one, but if those are actual models instead of giant lookup tables, one could assume the data is not reconstructible. However, that needs to be a design consideration of the models themselves, to prevent user data from persisting.
See also: Myspace accidentally dumping everything pre-2015 https://www.engadget.com/2019/03/18/myspace-lost-12-years-mu...
FWIW I’ve seen comments from a privacy engineer at google who posted here and said they actually do work hard to delete your data.
I’d expect Facebook and google to successfully delete the data (after giving 3-6 months for backups to age out) but wouldn’t trust most smaller operations to do so. And yeah, that doesn’t mean ML models or whatever but just the retrievable copies of your photos and text posts.
For the first few hours, I see new default youtube suggestions. After a few days, I see a lot of old search/view videos pop back to the youtube home page. YT still seems "suggested" videos for me to watch base on the past viewing info.
E.g. a sexual photo retrieved from a private social media album is no threat to your reputation when it’s been assumed into some machine learning model for detecting sexual photos, but it’s certainly a threat if a future data leak allows your enemies to get the actual .PNG or .JPG and send it to the news media or your loved ones. Knowing that the photo can actually be deleted is valuable in this case and I’m sure there are many other similar ones that could be listed.
Users want to delete watch history, which they can. Classifier models predicting what you want to watch are not the video watch history, nor are the models capable of producing it.
But that probably is not the only model generated from your data, is it? They probably have many other models generated from your data, everything from ad-displaying models to profiling models for Hydra.
That's a far, far cry from the reality of "possibly prosecutable in EU in some scenarios." We aren't even using a website right now that's in GDPR jurisdiction. But there's also a weird fetishization of GDPR I've noticed where people invoke it like "heh, my dad works for Xbox. just you wait buddy, he'll have you banned."
In my experience I’ve definitely seen GDPR result in a large company having developers looking at how data can actually be deleted and not just set to deleted=true. I don’t think my company’s lawyers were alone in thinking this suddenly became more important than before.
Close to 0%. There is no way they're actually deleting things on the backend, but this could help prevent your content from indexed by search engines inside and outside of facebook.
https://github.com/spieglt/fb-delete
And one to download all "Photos of you", which Facebook doesn't include when downloading your data:
I hate the pic attrition on Facebook. As people unfriend me or delete albums, some of the best moments captured of my life are erased. They really should be included in the archive download and more strongly link to your account.
https://github.com/Jaruzel/DeleteFacebookActivity
It works very well, but needs Windows.
Edit: had to disable 2FA to get it to work.
However, I would like to remove the hundreds of friends and pages that I don't talk to/don't represent me/are mining my data even more than the platform is.
Specific procedure I’d recommend is to make this a regular process: open fb feed, read until you find a thing which does not improves your life, click on owner (author / group) -click “friend” -> unfollow. Iterate until sanity is restored.
And one moral of this story (hi Facebook, please understand this) is that people are multidimensional. Just because you know someone once, in one context in real life, does not, actually, imply positive relationship on informational ecologies for all time in the future. “Authentic self” as defined by Facebook as a single coherent identity is a lie.
People change. I had friends who would post tireless political rants that just post memes now.
But slowly started getting requests from many of the previous friends. Some thought I had unfriended them, so had to explain to them that it was a new account. Felt obligated to add back most of the people. Now my feed is useless again.
Perhaps this should be yearly ritual.
This way, you can add people you want to, and you'll not get requests from people you're hiding from.
There was just nothing, or so little, that 'improves my life' as you put it, nor anything close. Just shit and drivel.
And it's not made for it. Quality doesn't equal traffic, nor even engagement in a monetisable sense. If it were, it would very quickly have learned that I didn't want to see memes, full stop. Nope, feed continued to be full of valueless memes.
For a brief period I considered I might be 'missing out', but it doesn't take long to realise that if you're missing out through not being on the right social media platform... who or what exactly are you missing out on.
Then I added back the few that were actually important. Huge quality of life improvement.
You can explicitly mute people as well or even just certain websites. My aunt has been posting a ton of weird alt-right news lately and I've had pretty good success by just blocking her news sources so I can still see posts from her that I care about (family stuff.)
I also have a few friends who I've muted entirely because I still want to be able to contact them, but I really don't care about any of the stuff they're posting to facebook.
This might be a little more manual than you want, but it's pretty simple to just casually block posts you don't like as you scroll through your feed.
I also don't really scroll the feed anymore, so I'd rather just keep it to people I want to keep tabs on (close family, professional contacts)
Not sure if that's still the interface today, but it worked for me back then.
* Edit: I checked, and you can still do this. It's under the main preferences dropdown -> News Feed Preferences -> Unfollow people and groups to hide their posts
It's no longer a checbox, but you can simulate a click in JS. Haven't tested so obviously YMMV
No doubt Facebook would delete these posts permanently and DMCA them.
I'd advise against it though, at least in the USA, due to the commercial value of the Harry Potter novels, and the penalties for infringement.
brew cask install chromedriver
The script wouldn't work. Works great now!
Traceback (most recent call last): File "./deletefb.py", line 3, in <module> from seleniumrequests import Chrome ImportError: No module named seleniumrequests
Any idea what's to be done?
If so the steps would be:
1) virtualenv -p $(which python3) ./venv
2) source ./venv/bin/activate
3) pip install -r requirements.txt
Then you should be able to run it, assuming you have google chrome and the chromedriver installed. It's called "chromium-chromedriver" on Ubuntu
One of the nice things is you have more granularity in what you remove - You can leave things that mention certain words, or in a date range, etc.
Because it runs in your browser, it should work with MFA logins, etc, without submitting credentials to an external service.
[1] - https://chrome.google.com/webstore/detail/social-book-post-m...
Back in the day there was a similar js bookmarklet that could be used to export all your friends' contacts (name + emails) so you can leave without worrying that you'll be missing out on all your "connections." It would be nice to extend this with a "backup my posts + contacts" info before deleting everything.
Seems like it would be easy to pop the text (and even comments) into some sort of structured dataset for archiving's sake.
So every day she just took a minute to the cringiest of that 1/365th of her content and in just one year she was free.
Would love to scrub like 5 years or so but keep some more relevant memories.
Thank you.
My flagged reply was: "Also, anything criticizing Islam or LGBTTQQIAAP rights"
Message: no such element: Unable to locate element: {"method":"class name","selector":"layerConfirm"}
Or... is there a Python equivalent of "nodemon" I could run it behind?Thanks for the great work! I was thinking Jumbo.app would implement something similar by now but I realize the problem was hard to solve because facebook doesn't let apps delete posts.
I believe it is a race condition which wasn't handled in the most elegant way.
CEO, Jumbo Privacy.
I recently tried to script clearing out my twitter linkes. (Some of which date back to before <3s, when a star could mean "yikes, I don't want to RT this but I want to note it". Others were just from when my twitter was more personal, and now that I use it for professional purposes I don't need random likes from the 2012 election lying around.
But Twitter now says I have about 4k likes, but only displays the ones since I tried to run my script.
Occasionally, a weeks or months later older ones resurface, but if I try to unlike them in bulk they disappear again to return at a random interval.
However, I am really interested in learning more about Selenium so I'm curious to look at your project.
IIRC twitter's API doesn't allow you to bulk delete tweets from the outside. I've found that exporting all tweets, getting the IDs, and then deleting them in the browser using the twitter.com only browser API will allow you to delete tens of thousands of tweets with no rate limits. I did something similar here: https://github.com/kylehotchkiss/fakeblock but it's not documented well and Twitter's export file format actually tends to change on a somewhat regular basis.
See "How do I download a copy of my information on Facebook?"
https://www.facebook.com/help/1701730696756992?helpref=hc_gl...
Great work fellow programmer.
Any data is probably never deleted. What makes more sense if a long running script (multiple years??) that updates / inserts / deletes your profile with random information.
I want the ability to delete everything after some period of time. Every message I send, every message I receive, I want it all to work like conversation -- not contracts.
I don't want someone digging up an email from the past where I said something stupid in a moment of anger, or frustration. If all tech worked like Snap Chat, or Signal, I'd be really happy.
I have no trust that Facebook actually deletes anything. Guessing anyone who tries to delete things actually flags the user and then their team of Zucks go in and look at the juicy content you were trying to delete... seems like 5 years from now, Facebook will charge for the "not have anything dumb you said in your 20s show up in public search" feature. Cool. Cool, cool, cool.
So I can have my old thoughts