Facebook Scans What You Send Other People on Messenger App
bloomberg.com
bloomberg.com
They also crawl links you share in private messages to grab the title, intro, and favicon to generate that clickable widget link thing.
[0] https://torrentfreak.com/facebook-blocks-all-pirate-bay-link...
Any links to more reading on this?
[1] https://www.zdnet.com/article/google-facebook-is-blocking-go...
> Some speculate this could be a Google plot to spur distrust for Facebook while bringing more attention to Google+. I'm not sure the search giant would go that far, but I think it's telling the company's employees chose to post about the issue on their own social network rather than contacting Facebook directly and asking for an explanation.
[1] https://www.androidpolice.com/2016/09/09/whatsapp-is-blockin...
This is one of those situations where the mass populace just didn't want to understand, couldn't understand or combination of both, but the mass populace is finally starting to see just how much of this tracking stuff has been going on. I still think it's not enough, though.
It is perfectly possible to crawl pages to get titles and FavIcons client side or using a web service that don't necessarily keep a log of these requests.
Same thing for censoring links. You can have a static list of disallowed domains and do all the filtering client side or have a web service that given a link returns true or false if it should be censored. It doesn't necessarily means it needs to be logged, kept, associated with a user or manually analyzed.
What we're seeing now is that facebook is actively looking at whole messages, that's a step up from the previous instances. It's still unclear if this is all automated or if some are manually reviewed. It's also unclear if these are associated with a user, or anonymized when analyzing them. Are these logged? How long are they kept for? Facebook should be more clear on how this all works. Otherwise we're just left guessing.
The program probably started a year earlier than that, since in 2011, Facebook announced they were using a tool called PhotoDNA to scan every image shared for underage content. https://www.facebook.com/notes/facebook-safety/meet-the-safe...
Regarding how Facebook crawls links shared in Messenger, there's this medium story I saw on HN a while back https://news.ycombinator.com/item?id=11875419
In that medium story, the Facebook devs claim accessing privately shared messenger links is a publicly documented behavior, citing the Facebook Crawler Docs. They quote: "The first time someone shares a link, the Facebook crawler will scrape the HTML at that URL to gather, cache and display info about the content on Facebook like a title, description, and thumbnail image." https://developers.facebook.com/docs/sharing/webmasters/craw...
I would update my original comment if I could, however the edit period has passed.
Ah yes, the "Think of the children" [1] argument. This is the favorite argument of censorship [2] lawmakers, dictators and everyone who wants to destroy personal liberties and privacy. They always invoke "think of the children" arguments because you look like a monster if you oppose it.
[1]: https://en.wikipedia.org/wiki/Think_of_the_children
[2]: http://www.abc.net.au/news/2014-01-31/wolf-internet-censorsh...
For example, let's say I create an eye catching flyer image detailing locations for a peaceful protest against the firing of a Mueller. Such a system could be used to block it.
Right now no one opposes building it, but the capacity once created can be easily abused.
"Scans" should be "reads and stores"
"What you send to other people" should be "private messages, images, and videos"
"What you send to other people" implies that there was no expectation of privacy in the first place, which (while true) I think does not match the 'normal' person's expectations or understanding.
News organizations need to be more candid with the public about how their information is being inspected and stored instead of using slick language to downplay the distasteful practices of many organizations.
Accusing Facebook of selling data makes it easy for them to rebut: no data changed hands.
Similar to accusing copyright infringers of “stealing” movies. It muddies the waters.
Maybe it's better to accuse companies of "selling access to the person".
Such as if you were eating at a restaurant and touts came directly in and started trying to sell you various things. Where the touts had paid off the business for access to your person.
1) Advertiser tells tech company "please show this ad to people you think are interested in X"
2) Tech company uses its private data to figure out who to show the ad to.
3) If you don't click the ad, end of story. No data leaves the tech companies servers to tell anyone anything about you.
4) If you click the ad, the place that ad directs to will know the ad you came from. Åt most, they can use referral source to infer some data about you.
To me, data is sold means that large amounts of personal data are shared about me without my consent. What am I misunderstanding?
The explanation is in your point (4). Let's expand on it:
I click an ad targeted at African-American, Christian homosexuals, over 40 years old, living in Boston, earning >$100,0000 with a custom audience set consisting of 100 email addresses the advertiser scraped from a forum.
By the way, I never told Facebook my sexual orientation, salary or religion, it was inferred based on other sites I visit.
When I click the ad, Facebook reveals to the advertiser that I match this description: and this is exactly what the advertiser pays them for.
Short of turning off ads or ad targeting, do you have any ideas about how to make this better?
Most users don't care about federation or decentralisation. They want low-cost and convenience but they're starting to realise they don't want it at the expense of their mental health or privacy.
Nothing can be simpler for the end user than a centrally managed service: I just go to wikipedia.org and start reading/writing. Censorship issues are avoided because moderation is decentralised: the site has a governance framework.
So I believe there's a space for a similar concept in social media. There's no question social media fulfils a genuine human need. But the profit motive inevitably forces commercial social medias to make decisions that detriment the user. For example, 'engagement': a santised term for addiction. Is it good for the user that social media should be continually engineered to increase 'engagement'? A non-profit has no such conflict of interest. It may even take steps to reduce engagement if it believes users are at risk of depression, addiction, anxiety etc.
I would love to talk to anyone who has ideas about bootstrapping something like that (email address in profile).
Don't get me wrong, I think all the flak Facebook is getting is deserved but there's little in the revelations coming lately after Cambridge Analytica that is really new. However the media backlash is a lot, a lot bigger and more sustained that I thought it'd be, even here in HN. I'm not one for conspiracy theories but could it be partially orchestrated by some political powers that be to kill his political aspirations? Or even if it didn't start that way, I guess it could have been helped by this.
They are probably beyond pissed because of the ad revenue loses...wonder how many of them will fold soon?
Believe it or not, he was, at one time a fairly eloquent speaker: https://www.c-span.org/video/?c4544001/donald-trump-1991-hou...
All this talk started because he hired a top Obama campaign manager as a lobbyist for his foundation, and somehow people got "he obviously wants to be President" out of it.
https://www.theverge.com/2016/12/8/13889308/mark-zuckerberg-...
https://www.vanityfair.com/news/2016/12/mark-zuckerbergs-pol...
Is there any indication that FB doesn't scan the contents of these messages before encrypting them with your own key and sending them across the wire?
[0]: https://www.facebook.com/help/messenger-app/1084673321594605... [1]: https://www.wired.com/2016/10/facebook-completely-encrypted-...
Steve Weis was involved in its development (previously PrivateCore, Google Security Engineer where he developed 2FA and the keyczar library) and jumped on the defense after it was initially announced. Earlier versions were reviewed externally by some pretty well-known cryptographers.
That being said, meta-data around use of E2E encryption in Messenger is still an issue since it's not enabled by default.
After that, I assume anything I'm doing on the internet is being data mined for advertising or some other source of revenue.
Name one thing that people remember for more than 2 news cycles.
Here's a better game to prove my point: what happened 5 news cycles ago? Do you remember? What were the headlines in January/February this year?
As far what was happening in January/February:
Pretty sure the government shut down at some point, I think that was around then.
I also remember a lot of news about the stock market breaking records and some post analysis on the tax bill which I think had already passed by then. I could have that timeline wrong on that though.
I think the Alabama Senate election was around then. That might have been earlier though, but I think there were still headlines about it in January.
That's all I remember off the top of my head. I follow political news more closely than other news lately.
It's safe to assume that much of the data we all leak is being mined for revenue.
The reason for this question is that even though you don't have a cat lots of people do! It's absolutely not a false dichotomy. Either people who have no cats must see cat food advertisements (bad choice), or cat food advertisements must be shown to people who probably have a cat (better choice). There's really nothing in between.
This is the world we live in, and you should prefer to see baby diaper ads vs cat food ads, when you do have a baby but don't have a cat. The statistical number of cats or babies is irrelevant. As you may know, Google is pretty good (not perfect) about not allowing keyword targeting that gets down to individual people so the privacy implications really are pretty limited.
-
That said, I have a funny story to share. (About adapting to this world.) I am learning a foreign language and I decided to watch baby cartoons in that language. But before I did, I thought to myself, "Okay if I start searching YouTube for cartoons for 1 year olds, pretty soon Google is going to decide that I'm a new mother and I'll see nothing but baby cartoons in my feed for the next 5 years."
I was sure enough in my reasoning that I went ahead and created a brand new Google account for the express purpose of being able to pollute its YouTube feed. I only watch stuff related to that language learning on that account.
This had the exact effect that I wanted. That youtube became absolutely awesome for spending focused time on my language learning, using all sorts of related videos. It includes people documenting what life in that country is like for tourists and foreigners, it includes foreign-language teachers' channels, it includes related cartoons and films at a good level for me, it includes political speeches from that country subtitled in English. I couldn't be happier with the result.
You know those acknowledgments we've been clicking through for the past few years by Google saying "Hey!! We're doing this. READ THIS"? I think it makes what they do pretty above-board.
As a consumer we're able to adapt to this, but it's not something I have any problem with.
(Disclaimer: I indirectly contributed to Google in the past but not now, I would definitely list it as a disclaimer if it were happening now but I remember that it changed how I wrote about Google especially when I was the most critical of them, so I think it's worth mentioning still. I am a bit nicer when I'm really pissed off at them as a consumer - but this is not the case in this instance.)
---
EDIT: I carefully edited this as it is falling to -1. I stand by the sentiments in this comment: they are correct. Downvoters are wrong.
>You’re ok with someone reading your mail then spamming you after learning something from it?
Does not match this (OP):
>When our daughter was born, and I sent an announcement via GMail, I started seeing ads for diapers.
I can see how the language "I started seeing ads" might sound like it applies to spam (unsolicited mail or messages) but to me it's clear that the poster just meant that the types of textual gmail ads they see changed to include baby ads. They're talking about this:
https://www.google.com/search?q=gmail+text+ads&source=lnms&t...
Now there's something actually interesting here. We don't have enough information to decide, but I bet this is the reason that poster paid so much attention to those text ads: they used to be highly relevant for them!
I bet they used to be filled only with stuff like industry conferences, IoT dev kids they were really interested in, security whitepapers, all sorts of stuff. (An indication of this is that they did not learn to mentally filter those parts of their inbox out as irrelevant.)
On one of my computers I have a hosts file that blocks Google's ads - so I often have to copy the ad URL and open it myself: because I clicked on it, and still want to see it after it doesn't open, and after knowing very clearly that it's an advertisement. Advertising doesn't get much better than this.
Yes. I do not own a cat but would prefer to see cat food ads.
I would suggest however, that cat food brands should advertise on pet channels on Youtube, or on pet related blogs perhaps. This allows targeting, but without collecting data about the users, many of whom do not know enough about online advertising to provide informed consent.
The best adverts I've seen were from Carbon[1] and Fusion (which seem to have merged now?). They were well targeted, unobtrusive, respected my privacy, and generally fit well with the content I was reading around the web.
Unfortunately keybase.io and keybase as search terms don't show me any ads. Let me try "public key". Still no ad. (On Google's homepage.)
Let me try keybase public key. Still no ad.
Let me try public key server. No ad. I clicked around on some related queries, until I finally saw an ad. (The query I ended up under was: sks keyserver setup).
Now here is the ad that I saw:
OpenPGP Library for .NET
Adwww.didisoft.com/
Pure .NET OpenPGP Library Easy API
Online examplesPurchaseProduct page
(How it appeared to me, at the bottom of the search page: https://imgur.com/a/n76d3)Now let me ask you. If you are a .NET developer and you're trying to work with PGP, would you rather see what I just showed you - or would you rather see an ad for cat food (even though you don't own a cat)?
Please be honest here.
Yes.
My issues with Adwords are the following:
1) they're abused by spam and garbage, so seeing a (potentially good) software product in there immediately turns me off and is actually a red flag.
2) it shows that the developer has no idea of the target market for their product - anyone that has the brains to use a software library would be blocking the Adwords anyway.
3) the above could mean they're instead targeting managers/CTOs as opposed to the actual people who will be using the library (developers), which is also a huge red flag.
In fact we're arguing/discussing better targeting now! As a developer you don't want to see an open source product advertised.
On the other hand, if you're a manager who can't deliver a feature that your customer is begging for because there is no out of the box solution/library for it, and you don't have the personpower - but lo, you can license one for $100/month that's worth tens of thousands to your customer, you'll be thrilled if you can learn about the existence of it. (Just a hypothetical, use a different one if you want: again, we're talking about better targeting.)
You didn't answer my question about if you want to see ads for a conference (I took two long sentences to ask you super explicitly) - is it because that's a resounding yes? I get that you don't want to answer the question because it's counter to your philosophy that targeted ads are bad. That philosophy, as we're seeing, is wrong.
I'm happy to see computer-related ads on Read The Docs, but on standard Google I'd rather have cat food advertised (or other generic product).
Oh sorry, I totally missed that you were talking about ads for conferences as opposed to conferences themselves - but again the above still applies; more than happy to see conference ads on developer related websites. On Google, especially when I search for a generic keyword? No, give me cat food any day.
I've never been bothered by ads I see in gmail.
they're also super explicit about the way in which they collect and collate data and I understand that maybe I will have to make different Google accounts to be targeted to different things.
I guess your and my perspective is just very different. but enjoy your cat food or Amazing skin lightening! (for darker skinned people) ads - they're both very popular products. (the latter is a scam product category.)
One of my issues with advertising on the web is the tracking/stalking. I am okay with ads that are targeted based on the content of the current page as this does not require stalking me everywhere.
My other issue is that highly targeted ads (to the individual) create an echo chamber. The thing is, after a day of work, I no longer want to see development-related ads. Give me cat food (or other generic products) instead. This is what I actually like about print & billboard ads in the real world - they're generic and actually make me discover new brands & products I haven't heard of before.
Finally the problem with Adwords and similar online advertising is that there just isn't enough vetting nor review for them - there is a lot of spam, garbage products, etc. Print ads are much better as the publisher more or less takes responsibility for them, so they're doing their due diligence when choosing which ads to run. You don't have that on the web.
Maybe we have a difference in preferences but I just never felt stalked or the like. I suppose it might change in the future. Thanks for your thoughts.
Ads where? Do you just mean display ads on the web? If so, your browser likely ended up on a list for diaper ads by some other means. Top of mind: - Did ever add a baby-related product to a shopping cart on a retail site? - Did you visit baby-related websites?
Less likely, but possible,if your browser cookie was linked to other personal information by 3rd party data brokers: - Do you use a loyalty card at a physical store? Did you suddenly buy baby stuff for the first time? - Did you return a product registration card for a carseat or something like that? The company my resell their customer lists.
http://adage.com/article/digital/google-stop-reading-emails-...
I know the implication is negative in terms of privacy, but it has its benefits if it could actually solve problems and provide value to the majority.
The problem is, though, that any form of bargaining (and sales is bargaining) is an information game. The more they know about you, the worse a deal you will get. One recent real example was the advertisers targeting women on their periods. Taking it further, imagine a liquor company finding out that you've just lost a family member and flooding you with ads for alcohol. The more they know about you, the worse you get screwed.
https://www.nytimes.com/2012/02/19/magazine/shopping-habits....
Given how Target shut that guy up, you have to figure they're doing even more along those lines nowadays.
Your comment is making me feel really, really old. Have people forgotten Gmail's history?
Gmail is not that old. When it was new, what you are pointing out was in the news. It was all over the news. To the point of members of Congress commenting on it. It was heavily debated. Google was very open about the fact that they were doing it. Google was the first (major) email provider to offer 1 GB of mail (well over the usual paltry 50MB that was the norm). Everyone asked "How can they afford it?" And it was very much in the open that it was being paid for ads, and that Google will mine your emails and provide you with targeted ads.
I remember while reading an email on Gmail back in 2004/2005 there was a very obvious targeted ad based on the content of that email.
I honestly do not mean this as a criticism, but I am really, really surprised that a HN poster was surprised by this. That Gmail scans (or scanned) emails and uses them for ads is almost part of their identity. It's like being surprised that there are ads in a newspaper.
"A Facebook Messenger spokeswoman" who wouldn't put her name to the statement? Ugh. Child porn is terrible, but very few people produce it or want to look at it. On the other hand, opaque and unaccountable algorithmic censorship hurts everyone.
No there wouldn’t be. Do you hear that about iMessage, SMS, email or the numerous other services?
This is done with PhotoDNA, a system used by many large tech companies for child pornography detection: https://en.wikipedia.org/wiki/PhotoDNA.
Yes.
The reason why you don't hear about it is because of all the work already being done to combat it.
How much work have you done in this field?
[0]: https://www.onmsft.com/news/microsoft-updates-photodna-softw...
There are many jurisdictions were people who are legally underage engage in sexting etc. There's hardly any "universal standard". Never mind areas where "gay" sex illegal etc.
My landlord has access to my apartment, and I certainly don't expect them to just pop in and take things out that they don't like -- I at least expect some kind of notice. You can apply this to basically anything in the physical world, like mail. Having the capability to access does not equate to having permission to access.
So yes I agree with your statement but in this example it’s already a law.
So it would be ok if there was a 30 day delay on the messages Facebook was accessing?
I don't understand your analogy.
Expect more of that due to the recently passed sex trafficking legislation that got Craigslist personals, reddit escorts and all sorts of other places shut down.
That is because there are laws preventing them from doing it (both the theft and the entry without notice). And I did live in a state that allows them entry without prior notice. And they did do it. And it didn't bother me because they clearly have the right to do so.
>You can apply this to basically anything in the physical world, like mail.
Again, very clear legislation on this. I believe it is an explicit felony to open other people's (physical) mail.
>Having the capability to access does not equate to having permission to access.
Yes, but sans any legislation, doing stuff on their platform does equate to having permission to access - especially if there is no legal contract (e.g. terms of service, privacy policy, etc) stating otherwise.
I honestly don't get this. In the old days people (including me) ran message boards on this web site. There was no shock when the owner of the message board deleted posts or put filters, etc.
"I know a guy that works there, and he says they take privacy very seriously!"
"Facebook is too big to make such a stupid decision like looking into your personal communications!"
^ All of these are arguments that I've heard here on HN. I can't even imagine what people on non-tech oriented sites say.
In The US there seems the trend is if you transport it, you get to data mine it (as long as "it" is digital, and "you" isn't a post service or phone company - not sure about isps). While in Europe the GDPR states that we live in a digital world, detecting that someone made a thousand copies of your data is really hard; but we'll make sure everyone is responsible for helping keep your data safe. Like the mailman and the telephone company.
But yeah, I think a lot of people still assume that a company facilitating private conversations won't have as primary business model to spy on those conversations.
I may not be 100% correct, but the history of this is related to liability. Telcos don't want to be held liable for crimes committed using their services (e.g. planning a bank robbery over the phone). They do not want the burden of monitoring calls to catch these people. So in exchange for that kind of immunity, they had to give up the right to listen in on calls.
I don't know the legal status for FB and the like, but I imagine it would be similar. As Facebook needs ads for money, they would rather not get that kind of a deal. As they decline immunity, they likely can be held liable for crimes planned on their services. Hence, they need to monitor.
Obviously, IANAL.
The dark patterns, ugly UI and unreliability, all with zero privacy.
Messenger has been rock solid for me.
[0]: https://www.apple.com/business/docs/iOS_Security_Guide.pdf
I just read the section on iMessage (from around page 49) and I can’t see where this is written. Can you point to the part where they say this?
> The private keys for both key pairs are saved in the device’s Keychain and the public keys are sent to Apple’s directory service (IDS), where they are associated with the user’s phone number or email address, along with the device’s APNs address.
For Apple as a company, not having access to iMessages is the safest thing to do, and I believe them when they say they can't access them in the current setup and are not willing to change that. It's because this would change their status from hardware/software vendor to telecommunications provider, with all related problems and costs - and they don't need any of these, so the best option is just to shield themselves from any user-to-user communication.
If you send a group message, Apple provides your messaging client with all of the recipients' public keys that are used to encrypt the symmetric key that actually protects the message. They could slip their key into that list and I don't think you would be able to easily tell if they did that.
If you send a message to a single person, then that's just a group of one.
The interesting question to me is if Apple can be compelled to write code to do this if they haven't already done so (and I don't think they have). I wouldn't think they could be forced, but like Microsoft did with Skype, they might do it anyway.
But in this case I think it’s 100% fine, even expected in order to stop bad content (porn, abuse etc) from going through
If they told you what was being blocked and why and also gave the recipient the option to override the block, then I think the slippery slope problem is minimized.
Still waiting on that...
With messaging, I've already explicitly told Facebook who I want to contact me - my "friends". I've never received a message from a friend claiming to be Nigerian prince or trying to sell me "generic Viagra".
But to actually use the phished credentials, they would have to log on to FB. FB could simply send a text.
I was pretty shocked when FB filtered out a pornhub link I tried to send to my then-girlfriend on Messenger. I thought it was very inappropriate of them to police our relationship like that. We were both adults...
It seems that she's been sending some rather saucy texts, plus a topless photo or two and a few panty shots.
She went straight back to Facebook, though, and really doesn't seem to care that she's feeding the beast..
Gmail has been scanning the email it's servers receive since it's inception.
Initially this was to show ad relevancy. Once your email content became more valuable than showing ads, Google removed ads.
https://www.nytimes.com/2017/06/23/technology/gmail-ads.html
Google has enough data on you that it no longer needs to scan email to show you relevant ads.
EDIT: Ok, from the responses I get I was confused. Maybe Allo? Skype? I'm sure someone else other than Signal and WhatsApp were using Signal's protocol. Just ignore this post.
[0]: https://www.wired.com/2016/10/facebook-completely-encrypted-...
For example, kids at school sending these "dank memes" to eachother and Facebook slaying bans to them.
PS: this does not only apply to companies with social products.
[1]: https://www.facebook.com/help/messenger-app/1084673321594605...
...
Schumer - who in 2016 railed that "a person's cellphone should not become a James Bond-like personal tracking device for a corporation to gather information" - has stayed relatively silent since Facebook's user data scandal with Cambridge Analytica broke last month."
Source:
https://nypost.com/2018/04/03/street-artist-taunts-schumer-o...
> "You can't watch your kids 24/7," reads one poster, which has a picture of Schumer, Zuckerberg, and a shirtless Anthony Wiener outside Facebook's New York offices. "BUT WE CAN."
Are they that desperate to find their daily DeBlasio/Schumer/Cumo 2 min hate? Can they not just call random Democrats Communist to satisfy their readership? Could Sabo possibly find a more tenuous link between Schumer and Facebook?
The fact that Sabo admits to having an unnamed financial backer coupled with the fact that this non-news is reported on in the Murdoch press makes me think this is some sort of guerilla marketing by an underhanded conservative firm similar to Cambridge Analytica.
Anyone who lives in New York doesn't need the Post for that.
I think E2E encrypted messaging is the only social solution that makes sense in today's ad tech pervasive tracking world.
On one hand, we have moxie and a team of people with no evidence against their integrity. On the other hand, a culture that responds to criticism with aggression [1] and then responds to criticism of that by deleting their communications [2].
TL; DR It's reasonable to trust Signal's iOS closed-source app while distrusting Facebook Messenger's also closed-source secret mode app.
[1] https://www.nytimes.com/2018/03/30/technology/facebook-leake...
[2] https://fortune.com/2018/03/31/facebook-employees-are-report...
Same applies with photos.