Facebook Really Is Spying on You, Just Not Through Your Phone's Mic
wsj.com
wsj.com
"Facebook does not use your phone's microphone to inform ads or to change what you see in News Feed."
OK... so they are saying they don't use your microphone to target ads. But how about precisely enumerating how FB uses your microphone?
Do they use it for any purpose other than helping you communicate during a call?
Do they try to infer any persona information about you, which can then be used indirectly to make money from your data?
I too have had odd coincidences where eerily relevant ads show up after I have had a conversation. If only FB was more transparent about what they do, I might not be so paranoid about it.
I too have witnessed the uncanny ads, but not even from my own phone (I don't have Facebook on my phone). A friend mentioned a particular restaurant I have never been to, been near, or searched for.
I can only assume my friend's conversation was geo-tagged either by my phone (android, no non-system mic access), or the data was combined on the server end to place both of us at the same place / same time, and used his recording to market to me.
I'd also like to note, the 'amount of bandwidth needed' is almost nothing by today's standards. Not to mention it can wait to transmit that data until on wifi. 8khz audio (telephone quality) is just kilobytes per second. A reasonably unsophisticated algorithm could trim the audio for an highs and lows (IE, statistically, sound below or above some dB threshold is trimmed because it won't be useful) and uploaded. We're talking about just a few kb per conversation.
Of course, the Facebook app is a memory, storage, and cpu hog (IMO), so I don't think it's unreasonable that given today's modern phone hardware some word recognition software may be present on devices themselves.
So i was walking with co-workers for coffee and somehow we talked about a college. Now i can assure you that i have never ever searched for that college and i didn't even know it existed before that conversation took place. And 30 mins later, i'm browsing instagram at work and there is an ad of the said college. If this isn't creepy af, i dunno what is.
Also, let's say what you are saying is true. How come ad decides to show up 30mins - 60 mins after conversation takes place?
These kind of events happen by chance and have been happening well before almost everyone started carrying mobile recording devices in their pockets and even pre-internet.
I have no doubt that advertisers would love to have that kind of insight but I can't help but feel that the anecdote of the form, we talked about A and A showed up X minutes later, is too easily explain by coincidence. How many topics were discussed, how many ads were shown? When I read something like this I imagine the anecdote should read more along the lines of how we talked about A, B, C, D, E and F and among the dozen or so ads I saw afterwords one of them was topic A! Can you believe it? And yes, yes I can, that is a pretty neat coincidence.
On the other hand if out of all 6 of the topics discussed all of the ads that were shown afterward were related to them, maybe not all, but more than one or two. That sounds like there is something fishy going on.
I don't mean to single out this particular example, it was to be a simple comment that had more to say. So, thank you for the inspiration!
The most blatant I've experienced was on the Wii U. The controller with the screen powers up and shows ads for new games every now and then. We actually had a bit of fun with it, casually talking about new games and guess what happened. An ad for that particular game was shown. I'm 100% certain now, that it happens with smartphones as well.
How many Wii U games are there anyway?
Would you expect the Wii U to do some sort of correlation (other people that own the same games A & B also own C, so you're probably interested in C)?
"We were talking about new games and I got and ad for a new game!" sounds like pretty standard, non-targeted advertising. Easily chalked up to coincidence. Steam advertises new games to me, some I'm interested in and some I'm not, but I don't think Steam is reading my brainwaves.
I thought that was some great scheme made by the adults.
Turns out, it isn't. Just everyday life coincidences. Reinforced by the global rythm of social life.
College ads are more likely to appear in the periods when people talk about college.
But to really answer this question, we could build a better experiment. Write out a set of topics on index cards. You have to be careful that topics aren't new product rollouts (otherwise you really have to think about how you decided to write that topic down in the first place). Draw a card, don't talk about or search the topic for some amount of time, then inject the topic where it might be observed (talk about it and/or search for it somewhere), then for some amount of time, see if it comes up.
And have a control group that you simply don't bring up at all. Make sure to have more than one, so that you still have more remaining if someone around you brings up one of your control topics.
* the "somehow we talked about a college" is because someone else at coffee was looking at something college related and it was recently on their mind so they brought it up. You're associated with those people, so ad agency it decides to show you an ad. It may do this a lot, but the times it actually works really stick out to you.
* After the coffee, one or more people search or perform actions associated with that college. Ad agency knows you were recently meeting with them so decides to show you some associated ad.
Keep in mind that the ad agencies, whether Facebook or Google or some lesser known but still large one, have their own profiles of you and who you associate with.
I think Occam's razor holds that since we know there are multiple agencies tracking what you do and who you do it with online and that can conceivably be used to explain most of this, it's a much more likely explanation than a large company outright lying in a way that would have a horrendous backlash if proven (and it's not really that hard to prove if people got serious).
Edit: whoops, not that hard to prove...
Wedding rings have an EXTREMELY high margin, and almost infinite budget for advertising. If you google rings once, you will continue seeings ads for YEARS, regardless of any conversations you have.
Honestly if Facebook knows you are in a relationship for 3+ years and you are under 35, you are going to start seeing ads for engagement rings even without searching for anything.
If you really think they are getting you through the microphone, start talking to your phone without anyone around, while browsing Facebook about Baby Diapers, Baby Formula, and Baby Toys. Don't search for anything baby related, and obviously use something else if you are a Father already.
My favorite is when you finally buy something and they still show you ads. Like a coworker mentioned, if I buy a refrigerator I don't need another one... I have one house! Come on! Maybe they do it in case you change your mind? Who knows how these ad companies think.
The only way for them to know you have purchased a product is if you buy it online and get hit with the confirmation tracking pixel on the checkout page.
Most people see ads for a fridge, but then go into a physical store to make and finance the large purchase.
The ad-placement folks have technology and patents and millions tied up in all that. And what do they come up with? "Show ads for what they just bought"
The "fridge store" is doing a pixel based retargeting list, which is easy to take you off when the purchase is made.
But they also might have a list of 'customers likely to purchase a fridge' from Google searches that you are being served ads to as well. You might be on a 'look-a-like' list from Facebook, because people in your demographic tend to buy new fridges. You could be removed from that list if Google shared the IP addresses of everyone on their list, but at a serious cost to your privacy. These walled gardens are ultimately good, but result in a ton of advertising inefficacy.
I know its annoying but the advertisers are smart enough... we don't actually want them getting any better than they are now.
Facebook is not a trustworthy company, so I have a hard time believing them about this. But it's why I don't install ANY Facebook applications on my phone anymore and that includes Messenger and Instagram, or any other company they purchase.
Remember when Target got into hot water for outing a pregnant teenager based on an assortment of items she bought [1]? That was based on items like scent-free soap, cotton balls, and vitamin supplements.
That was 6 years ago and (no offense to Target) done by a company that isn't nearly as adept at doing that kind of targeted advertising or inference.
For example, Facebook might know (hypothetically) that you're a young adult in the tech field who lives in Brooklyn and spends a lot of time looking at photography or you're friends with a lot of amateur photographers. So they show you ads for high-end cameras from the most popular vendors, knowing that you're likely to develop an interest in photography if you haven't already.
That's just an extremely simple example, you can imagine that there are far more subtle indicators about someone's interests or likely interests (like the cotton balls for pregnancy).
The timing is almost certainly a coincidence. It's far more likely they show you those ads all the time and you're just noticing them when the coincidence occurs because you're on high alert now for things like this.
1: https://www.forbes.com/sites/kashmirhill/2012/02/16/how-targ...
> I've had discussions with people about things I've never searched for or purchased
That doesn't mean his network hasn't been searching for or purchasing those things or leaving behind other breadcrumbs... which is probably enough for Facebook to guess you're a decent candidate for the ads.
I just wouldn't be shocked if it was a coincidence as well, or something else driving it.
For example, the company mentioned in the OP could be something mentioned in the news recently or in a hot market position right now.
In which case, it's not bizarre that you might see ads for them on FB as well.
I think the example another person gave is almost just as creepy -- your friends influencing your own advertisements. That means that what your friends are searching for, private (embarrassing) things, could potentially leak over to you.
I'm pretty aware of the fact that if I go looking for socks, I'm going to start getting sock advertisements on almost every page I go with advertisements, on both Google and Facebook. I don't find that creepy, just dumb and ineffective because it usually starts well past the time I already bought the socks.
1) Aware of construction happening close to your place of employment (in any number of ways that doesn't require any super advanced knowledge)
2) Knows that people usually start to get fed up with construction noises after X days
3) Started showing you popular co-working spaces as a result
Or, as you noted, your co-workers started searching for co-working spaces and Facebook picks up on that and assumes something is happening in your office such that other people might also be interested in co-working spaces. Creepy, sure. But doesn't require clandestine recording and parsing conversations.
Or, even simpler than that: I see WeWork advertisements all the time despite never discussing them. It's not really that insane to think that places like WeWork might just be targeting your demographic and that's why you saw the advertisement.
Anyway point being, none of these explanations require Facebook to record you.
That being said, that's totally based on just a hunch and I have no real data to back up that assertion so I could easily be wrong.
With the amount of data that's floating out there about you and your friends, these coincidences are way, way more likely, and are very reasonable if you think about it.
Your friend told you about some restaurant, and then an ad shows up on Facebook.
Why is your friend telling you about it? Maybe he went there recently enough? Maybe he gave the restaurant a review? Maybe the restaurant has a list of people who've been to the restaurant before and is using Facebook ad manager to target those people and their friends? It's on his mind, so maybe that's causing him to leave enough breadcrumbs behind to make Facebook have an inkling of his interest, and you're his friend so Facebook knows to perhaps nudge you as well. And maybe Facebook actually has some idea that you and that person are meeting (yay geotracking!)
Now add a prediction system driven by deep learning models that are terrifyingly good at finding signal with a decent dose of probability (think of all the conversations you've had that didn't result in eery advertising), and you've got yourself a frightening reality where a company doesn't need your thoughts to make a decent guess about what you're thinking.
Frankly it might not even be this complicated. Maybe your network searching for or purchasing things is enough for Facebook to guess you're a decent candidate for the ad too, and you guys just happen to meet up on the same day.
So, I did this sort of thing years ago when I wrote a tweak for the InPulse smartwatch (later became Pebble) https://github.com/brandontreb/inPulseNotifier .I was able to hook into the system messaging, forward it to a custom bluetooth stack (sending it to the watch) and forward the message up the stack to be displayed by the system.
It would stand to reason that the same sort of process would be effective for catching Facebook invoking audio recording. Once you hook into the AVAudioRecorder's interface, you could theoretically observe the following:
1. Open the Audio Recorder app and hit Record - An alert should show to prove your tweak is working.
2. Open the Facebook app. If you receive a similar alert at some point, you could at least prove that FB is invoking the audio recorder at some point without the user's expressed permission.
Am I crazy or could this test actually work?
https://thehackernews.com/2017/10/uber-screen-record-iphone....
> information security research. ceo @ sudo security group (https://verify.ly).
> previously: founder of "Chronic Dev Team" responsible for many years of iOS jailbreaking solutions (24kPwn, absinthe, corona, greenpois0n, etc).
Certainly a fair question.
Do you have the hashes to prove that what you tested matches what is actually installed elsewhere?
No, I'm not actually claiming there actually are different versions in the wild. I just find it strange that anybody can make broad claims about what widespread software may or may not be doing. Widespread use of "A/B testing" and forced remote updates should make everyone question the nature of every binary, even when they have the same name (including version number).
The article in question starts out breathlessly accusing Uber of spying on users, only to completely walk back the claim by the end. Just by reading the article alone we see that the permission was granted to overcome a capability lapse in the Apple Watch.
That's a terrible reason. Taxi drivers also have an unreliable source of income with the burden of medallion rent in some of the larger cities.
Do you also boycott all construction since that is also unreliable for basic laborers?
If Taskrabbit started sabotaging income for highway construction workers I might avoid it (although to be fair I’ve never used it).
In the other hand that is a considerable effort for someone who does not usually work with this part of the stack... would you be able to introduce this changes in an android OS?
But you are right, this would def be a considerable effort for someone not in the jailbreaking space. I would love to hack it up, but unfortunately haven't dabbled in JB dev since 2011.
That's why I posted the comment to HN. In hopes it might inspire someone in that space to build it. Might also be worth jumping in the theos IRC channel. For someone with the toolchain already set up and a jailbroken iOS device, the code is actually pretty trivial.
On the other hand, facebook can check (at least on IOS) easily if the device is jailbroken and behave differently.
You can also patch binary and inject some code, (probably swizzle AVAudioRecorder methods) for the same effect.
In this case, Facebook can check binary integrity, and change behavior accordingly.
So this is kind a cat and mouse game.
But if they decide to do, I think best way of action will be some defensive programming around it, with plausible deniability.
I am guessing they are already checking binary integrity etc, also they can probably push code updates from server. So when you put this pieces together, they have everything they need technically.
I'm not 100% sure what's involved in jailbreaking iOS, but I'm pretty sure on a rooted Android you could put measures in place to "fake" results for any root checks the Facebook app would run. You could patch any APIs Facebook could use to make such checks.
It is a bit of a cat and mouse game, but as the long history of software cracking shows, as long as they still own the machine, the crackers always have the upper hand.
Disclaimer: I don't really know a) if there is some other way to interface with the mic or b) what I'm talking about in general.
The tricky bit is when users give microphone access to the app (i.e. for video recording functionality), but want to verify it's only being used then.
Which explain the race to voice assistant, that are used maybe once or twice a year, yet everyone invest millions. from cortana to Echo.
I am not network guy so just asking and seeking for valid explanation.
- The audio packets aren't recorded from the Fb App?
What if...
- The audio packets aren't then sent to Fb?
Fb only buys this data + integrates = problem solved
Which is why we don't want random apps having permanent mic access. Or permanent anything access. This is why data mining is bad, not just because the party doing the mining can get the data, but because they can sell it to third parties who combine it in unexpected ways to leak data that you really don't want to be public.
It's when an App goes behind your back and does it (And then sells it on) then it's bad.
In light of that restriction, what might be interesting is looking at the amount of data transferred by the Facebook app with/without the microphone/location services enabled. (this is a data project I have in the pipeline)
Certificate pinning is a hurdle to reverse engineering, but a surmountable one, at least on Android. Since the app is running on a phone where you may potentially have root, you can pick it apart with a debugger and see the traffic before it leaves the phone. This is technically challenging, but it is something that people do sometimes.
http://www.cnn.com/TECH/computing/9908/20/aolbug.idg/index.h...
If you're the reverse engineer and I'm the app author, you find/replace my CA file. Then I respond (or anticipate!) by checksumming the file to detect tampering. Then you respond by find/replace on the checksum.
Then I obfuscate the checksum string. Then you respond by faking out the platform's checksum API so that it always returns true. Then I respond by computing a checksum that I know should fail and verifying that the checksum API isn't just always returning true. Then you respond by faking out the platform checksum API with a whitelist of blobs whose checksum it should lie about.
Then I respond by statically linking my own checksum verification code into my binary instead of calling the platform's. Then you respond by patching my binary to jump around the code. Then I respond by using code obfuscation techniques.
And on and on. Given enough time and resources, any implementation I create can be subverted. But if I'm a huge tech company, I can afford a lot of time and resources too, if I want to. I can't eliminate it, but maybe I can make it something that rarely happens.
The whole cracking scene can afford far more time and resources than any one tech company. All adding protections does is make a more valuable target, because crackers love a good challenge.
"There's always a crack in everything. It's how the light gets in."
Can? Sure! Would?
Security is one of those things that most people say they care about, but they are really not willing to pay for. Big companies have resources, but they are also in the business of making money, so they will put most of those resources to work on features that produce a ROI. Security is a huge cost center, and even when taken seriously it will be pursued only to the degree that it addresses/mitigates risks enough to conduct business.
From the point of view of software, it's impossible in principle to tell whether or not the code being executed does what the user wants it to, and only what the user wants it to. Half of that is the halting problem, the other half is that "what user wants" is an General-AI-complete problem. Moreover, the software can't even tell the difference between "the user" and "a malicious third party".
In meatspace we solve this problem with rules and laws. Software, for better or worse, moves around too fast.
I gave it a thought, and decided I don't need apps to be able to obtain highly elevated privileges, root or similar. This is, indeed, dangerous, esp. regarding ADB root access (a rogue "charger" + an accidental wrong tap[1] = totally compromised device that can be only fixed by full re-flashing). I needed my own firmware that does things my way, signed with the keys I control.
So I did. Now all the "secure" apps are happy, and I still have the control over my device's behavior.
( Okay, I've cheated - I had to sanitize androidboot.verifiedbootstate when kernel initializes, because I can't control the bootloader :( )
[1] Hm, maybe password-authenticated root access is okay, though... But not a typical "tap to allow" dialog.
The idea that Facebook is doing this is just ridiculous.
Facebook invading our privacy by recording persons of interest or the people en masse would be horrible and in many ways unprecedented, but it wouldn’t be beyond the levels of abuse we have seen from powerful people in the past. I can easily imagine that the app supports hot mic capabilities and that they do turn it on sometimes at least at the request of law enforcement. And then the question is... when else would they turn it on? And would that program ever grow? Would they ever fork the program so each team involved thinks there working on a small project? This is all speculation but I can imagine a situation where it starts small and then grows until it seems like an insane program but everyone involved is accustomed to it.
Some companies do "highly illegal" things all of the time. It all comes down to the fact that whoever is in charge:
1) doesn't hold to a moral system that restrains them from doing said illegal things (or at least doesn't hold to one consistently)
and
2) thinks they can do said illegal things without getting caught, or if they are caught, thinks they'll be able to recover reasonably well from any punishment (if there is any) that is handed down.
Personally, I do not know whether Facebook is recording/transmitting data like this, but I guess if I found out they were, I would not be surprised -- given the things the company has done in the past, and the things their leadership has said and done in the past.
When you read their responses to this controversy, pay more attention to what they don’t say. Whomever aggregates your viewing habits by listening to audio from your TV may be listening, for example, Facebook just Hoovers up the data.
Even if it's an issue, a security researcher could recompile Chromium or Firefox with certificate pinning turned off and test with that.
[1] https://groups.google.com/a/chromium.org/d/msg/blink-dev/he9...
I suspect I'm overlooking something, as surely some security researcher would have done some of this already.
...and if the signing code is also signed, then patch that.
It's patches all the way down. ;-)
There are couple ways out, from revers engineering the binaries, through jailbroken/rooted phones, running apps in simulators, etc.
https://serializethoughts.com/2016/08/18/bypassing-ssl-pinni...
http://blog.dewhurstsecurity.com/2015/11/10/mobile-security-...
is uber still the most hated company or has the magnifying glass moved onto somewhere else?
I think you wildly overestimate how angry users get about privacy violations. For examples, Target, Yahoo, Home Depot, and Equifax have not been screamed into rubble.
It would likely be a blip in the news, and then people would move on, as usual.
Remember the Sony rootkit fiasco? People still buy Sony, and most people probably either never heard of that incident, don't remember, or don't care. Buying whatever the new Sony gizmo of the day is is more important to them.
Microsoft has had endless spyware fiascos, and people still routinely buy Windows, as long as they can play their games or run Office, that's all that matters to most of them.
Then there have been scandals like Enron, where the execs knew that they were doing something that was clearly illegal, and that their company really would be devastated if what they did was ever revealed. These "smartest people in the room" did it anyway.
Corporate history is full of just such deceptive and destructive practices. I'm not sure I'd put Facebook above that sort of thing, a priori.
I can see that happening again with voice recordings.
Your relative probably installed something and pressed "Next" through all the dialogs including the ones asking if they want to install super helpful bundled software.
But it's an interesting question: if someone credibly proved that FB was "wiretapping" on such a massive scale, would they get prosecuted? How much could they do in their own defense? Are they so enmeshed that prosecutors wouldn't bother?
Feels like a case of "unstoppable force meets immovable object".
So not literally everyone... but still many.
It was the second to go, just after Facebook.
For what it's worth, this public opinion backlash has not appeared with other companies: "Of course we're not working on leaked project [x]". "Look at project [x] we're working on!"
Even Facebook's own under-disclosed psychological experiments have been largely forgotten; Facebook has suffered few if any long-term ill effects from it.
There are some shady companies out there [1] that use the mic to listen to what shows are being played real time. [TVs in the US are on all the time] (Check out their customer list). These companies need this pinning.
In other words, any reason that the user should not be able to "disable" it? (Use own root CA.)
Consider that in this case, the data being transferred to Facebook belongs to the user.
Is it unreasonable for a user to require that they be able to see what data is being transferred before they agree to transfer it?
As for transparency, look to GDPR and friends to see what rights are being declared.
Also, storing the data until it's plugged in would require an unusually large amount of storage, and that would be detectable.
That aside, speech recognition isn't that heavy of a process these days if all you're looking to do is extract keywords. We used to do industry-leading large vocabulary continuous speech recognition on a Pentium 133... phones these days are way beyond that without breaking a sweat. Detectable? Sure. But remember, I was talking about this person's plan to look at network data.
Furthermore you don't need to store all data. You can store only when the phone is hearing stuff, as determined by a super lightweight measure of magnitude that does no speech recognition whatsoever. Is the storage detectable? Sure. But again, what was I responding to? Network traffic monitoring.
The thing I thought was interesting is just how adamant people are about FB spying via mic.
They also had an update in their year-end episode: https://www.gimletmedia.com/reply-all/113-reply-alls-year-en...
Both worth a listen.
If you're in the US, you can also be a privacy activist, which we need more of.
I'm of the opinion that it's easier for your average Jane/Joe to believe (and maybe even preferable to believe) that someone is listening and responding to your words than a computer piecing together a picture of you from unrelated clues via some nebulous "machine learning algorithm".
Anybody can listen to your words and advertise to you based on them. It is, on the other hand, not feasible for a human to look at a stream of unrelated posts and figure out that you're pregnant.
That being said, perhaps the ads are doing their job and planting the idea for that particular product. You then bring it up in conversation or mention it out loud. Then, when you return to facebook later, you see the same ad again and due to it's recent mention, it jumps out at you.
Both seem more plausible.
Telling people that doing A and B leads to C which leads to D which makes E more likely to happen is just a bunch of gibberish that can't be right because who can you blame?
I think I got distracted, or maybe since my show wasn't on, I just decided to go shower, and left it on for a bit. The next week, I start getting notifications about golf on my phone.
There are two possibilities here- Verizon (my cable provider) is making data available on what I watched to google, or google is using my microphone to pick up what I am watching on TV. I don't know which is more likely, but VZ and google having a partnership like that and keeping it secret seems unlikely.
Is there _any_ evidence that Facebook is spying using the mic? Surely we have to start there, right?
If you only use direct evidence to come to conclusions and toss out theories and deductive and inductive reasoning you won't be able to function in this world.
But it wouldn’t be fair to say that they rape girls in the alley, or throw up your hands and say “look, we don’t have much proof either way!”. It’s a completely baseless accusation that makes it harder to talk about real problems.
They buried a lede — the story laid out a case that Facebook doesn’t need your audio. Isn’t that a bigger story?
We showed you our friends, our relationships, our interests, our intimate and disarmed states, our rants, and probably half of the websites we visited. (in retrospect, that was dumb)
Oh, your tin can and strings might show? Competitors might get ideas? Please.
Until then, not a fan of you, not clicking on your ads, and generally avoiding your site. In fact I think I'll start deconstructing my profile as soon as I can muster the courage to choke back my gag reflex.
Sincerely, A growing group of mugged social network burnouts.
Honest question: why would it be unreasonable for us to expect server-side code to be open-source? Facebook's value lies in its brand and its infrastructure, not in its code, so there's no risk of upstarts taking Facebook's code and standing up a clone (which, even with the code, is way easier said than done).
Facebook must offer enough to the users that the network is still worth coming back to while still giving advertisers a chance at having their eyes. A major breach could cause user and partner abandonment because of security concerns. Once the genie is out, there is no putting it back in. Their stock will fall faster than they can rewrite the product.
It is unreasonable for us to expect open-source for server-side code because it exposes Facebook (and potentially it's users) to a lot of risk for only a small upside. 1) While open-source software has myriad benefits, those benefits require the public at large to audit their code as it is being continuously changed and deployed. Can we keep ahead of the criminals exploiting freshly merged and deployed commits? 2) Knowing the source code is one half the battle, the other half is knowing what is actually executing at runtime. How would users verify this to get the value of open-source? 3) Open-sourcing server side code of Facebook could have serious negative consequences for users or Facebook in the event of a breach due to intimate knowledge of the system only afforded by being privy to the source code.
Not a point, but a philosophical question: *) Where does this stop being virtuous? Should Microsoft open-source SMB tomorrow? Would you feel comfortable with that?
Edit: grammatical fixes
Isn't code infrastructure?
Yeah, at some point they blocked the mobile browser from working with messages. I think you can circumvent it by changing your user agent string to a desktop browser.
I also hate the fact to even look at my messages I have to "refresh as desktop site" on iOS
Furthermore, I believe that the Facebook app codebase is massive in scope and highly illegible due to most of it being auto-generated from other codebases. It has over 18,000 classes on iOS. The odds of anybody being able to meaningfully audit that are pretty low.
Facebook has a lot of compute resources, but they wouldn't have to use it. Your smartphone is more than fast enough to do simple speech recognition. The accuracy rate wouldn't have to be that high - you won't get mad if you see an ad for a misheard keyword.
[1] http://21stdigitalhome.blogspot.ca/2013/06/vcp200-voice-reco...
the downside was power usage. Motorola made one that was power efficient, used by nokia in the 90s and its pretty much the same chip in google's phone line today (just even more power efficient).
division is under lenovo now
Your phone reads sensor data as a base state.
It has a battery impact but much less than sending all the voice data continuously to a server somewhere. The biggest battery killer would be the wifi or 3G transmitting non-stop in that case.
Only if by "everyone" you mean people foolish and vapid enough to use Facebook and give them access to your phone.
Ad is easy. You don't have to understand context. Just listen for a thousand or so keywords related to products that are paying you. Then if detection happens apply some rudimentary sentiment analysis on the surrounding phrase and that's all you will ever need.
if you have a couple millions for me to start a small team we can offer this as a service next month or two.
This would have the upside of not requiring any reverse-engineering.
* Own a cat, dog and other animal
* Have between $100k- $999k liquid investible assets
* Have a net worth between $1 and $1m
* Am highly affluent
* Am a high spender
* Am a frugal spender
* Own a house
* Have multiple families
Yet none of these are true (well, I guess apart from the 'has income' demographic I'm in).I know Twitter isn't known for being an advertising powerhouse (esp. compared to Facebook), but I wouldn't take too much stock in Facebook serving me up irrelevant ads.
I moved to the other side of the world 2 months, updated my "living" location on Facebook and have been tagged at multiple locations in my new city, yet Facebook still serves me ads for buying an iPhone or Car back in my home town.
For example one guy had a buddy who had recently purchased a certain motorcycle, and all of a sudden he started seeing ads for that motorcycle.
But... really there's a simpler explanation than the microphone. Although of course it doesn't by itself rule out the microphone being used.
Facebook can just see when you are in the same location as some other people, and see what things those other people are into, and then signal whatever ad networks that you might also be a prospect for those things. Visit your buddy and see his new bike? Start getting ads for the same bike. No audio needed, just location services and some posts on FB from your buddy about his motorcycle. And there are other sensors beyond that. A lot of things are possible once the user has granted permission for use of various inputs.
Also if it is the microphone as the story suggests it could be in some cases, the evidence for it being any one particular app is thin. There are other apps that get granted microphone access by users all the time, and some of them should be looked at, not just Facebook. Not to defend Facebook here, but the net should be cast wider than just one app, even if the ads are appearing on Facebook, which itself is perfectly capable of gleaning interest information from multiple sources including other ad networks fed by other apps.
Less than an hour later, Instagram showed a Meow Mix advertisement on my wife's account, on her phone.
We have no animals. We have never had animals. She probably (though maybe not) has never logged into Instagram or Facebook on my browser.
It was too much of a coincidence.
I've shared this story on HN before this was what I experienced from 2016ish:
I saw an ad buried in my facebook feed to "buy Gallium and Bismuth metal in Australia" I thought it was an oddly specific ad so I made a joking post about it - turned out several of my friends were seeing the same ad. A common friend we all shared who is a high school science teacher spoke up. He explained his class was studying the periodic table and he had purchased samples of Bismuth and Gallium online to show to the class.
I'm absolutely convinced only reason my friends and I saw that ad was because we all shared a friend who was searching for this stuff online.
That level of surveillance is really creepy to me...
What’s funnier is when google decides that it needs to localize my search results for that country. So there is a lot of tracking that assumes all traffic from a single residential IPv4 address is somehow correlated.
The only way the data could have gone over is through a malicious browser extension.
Also consider this in the light of recent EU rulings on Facebooks tracking of non-users via the Like button on websites being an illegal violation of privacy[2]. As usual the law lags the technology by many years - was the EU even aware of the Facebook mobile SDK being wisely installed in many 3rd party apps when they made this ruling? (edit: reading the report from the University of Leuven it seems they were at least aware of the implications of things like Facebooks Mobile Advertiser network)
[1] https://www.google.ch/amp/s/techcrunch.com/2015/10/22/facebo...
[2] https://www.google.ch/amp/s/amp.theguardian.com/technology/2...
Here's Apple's privacy policy: https://www.apple.com/legal/privacy/en-ww/
And yes, it's followed relentlessly.
Ad blocking is another example; allowing it in iOS was probably a strong blow against Google.
Apple can and does collect a lot of data from your phone. Their business might not rely as much on individual targeting, but they still want to understand users as much as possible and have the means to do so.
In fact, they still collect tonnes of data from phones, but now they're more careful about the data not being user-accessible. A few quotes from the link you posted to their privacy policy:
> "We also collect data in a form that does not, on its own, permit direct association with any specific individual."
The "on its own" sounds a little scapegoat-y tbh.
> "We may collect information such as occupation, language, zip code, area code, unique device identifier, referrer URL, location, and the time zone where an Apple product is used"
You can learn and infer a lot from those vectors. Towards the end of the paragraph they also mention that they use this data, amongst other things, to deliver "better advertising".
> "We may collect information regarding customer activities on our website, iCloud services, our iTunes Store, App Store, Mac App Store, App Store for Apple TV and iBooks Stores and from our other products and services. Aggregated data is considered non‑personal information for the purposes of this Privacy Policy."
Ofc.
> "We may collect and store details of how you use our services, including search queries. [...] Except in limited instances to ensure quality of our services over the Internet, such information will not be associated with your IP address."
Ensuring "quality of services over the Internet" is _incredibly_ broad. For a company like Apple it could apply pretty much to anything tbh.
A lot of people don't know Apple collects all this data; and part of it is probably the fact you can't disable this collection. Only App usage, the one that might also be shared with 3rd party devs, is optional.
Thinking Apple doesn't take part on Google or Facebook scale data collecting because they sell phones and not ads is not only inaccurate (they do sell ads), but also a little naive. Data is very valuable. I'm not saying Apple doesn't care about privacy; their business model relies a less on individual targeting than Google or Facebook, but they're also in on the game of understanding users as much as possible, and given they control the phone they're in very deep.
[1]: https://arstechnica.com/gadgets/2011/04/how-apple-tracks-you...
Good luck to the author as now he will see generic AT&T and Galaxy S9 ads. Privacy has its costs and one should make an informed decision eitherways.
Why do you need help from targeted advertising?
What are the things where targeted advertising is a good way to get help?
Is targeted advertising on Instagram and Facebook the zenith of this help for you? Can you imagine another way of getting this help?
You're making the common mistake of assuming "targeted advertising" means targeted to what you want. The point is to allow marketing to target specific groups. Any overlap with your interests is just a coincidence.
You are not their customer; your interests matter only to the extent they provide more data points advertisers can target.
I spend >30min of my time on FB. As long as I find relevant content (which includes ads) I am good. I am not saying FB is the best source for all to find content but I do like to see updates from my friends and pages I follow and if relevant ads are sprinkled in between, I am a happy user.
Either you don't understand that FB isn't trying to target your relevant interests (they target you based on the categories[1] advertisers choose to target), or...
> I spend >30min of my time on FB [...] which includes ads
you are so entrenched in consumer culture and used to having your opinions manipulated by marketing departments that you no longer recognize the difference between "relevant content" and attempts to "nudge" your behavior in specific directions. I suggest taking a break from the internet/media/ads for a couple weeks.
edit: forgot url
[1] https://www.wordstream.com/blog/ws/2016/06/27/facebook-ad-ta...
I've had quite a few of those, and usually I can trace it back to me googling something, etc. But this time, nada.
She's received ads for Civic Type R (she hates cars), Senior Java Developer (she works in a totally unrelated field), cooking tools (I do all the cooking), and tons of other things. It creeps me out.
Edit: How Facebook would get your browsing data, even if you’ve disabled things like ads and those FB like buttons (like I assume most HN posters would), is beyond my wild speculation.
And none of this would have to happen while you were talking about cars, or anything else. Just enough times to make the correlation, and with data that could have been collected months to years ago. Google Now did that for my commute from work to home.
I think it is more likely Facebook is using location data to create edges on a shadow social network. We just bleed metadata.
I've picked two different, random topics (boats & umbrellas) to occasionally search for on my phone and computer respectively. So that should help me figure out what the source is.
Side anecdote: One weekday after I vacationed in Tahoe, I saw a BART (subway) ad for Tahoe, and I was like, "oh, great , probably because I just came back from ... wait, that's not possible!"
Any one of:
* OS-level confirmation (permission entitlements, cpu usage, etc)
* Packets resembling sound data being caught in flight.
* An internal leak of the method they'd be using to do so from one of three of the largest tech employers. Not even Apple can keep their secrets secret, and they're probably the most paranoid tech company in existence.
You know, literally anything concrete, rather than evidence-free accusations based on fallible memory.
So far, not one bit of these instances can't be explained by a combination of Baader-Meinhof and confirmation bias, with a mix of plenty of non-audio data that Facebook no doubt has. People are so willing to paint FB as this boogeyman that they're disregarding basic logic.
I understand that any one person’s anecdotes are weak evidence, but your comments are going much further and claiming that such tests can never be evidence, even though much scientific knowledge is similarly obtained.
Not even necessary. You could do keyword recognition on the device itself, pushing a list of keyword<->waveform maps, and sending an indicator when they're recognized.
I just asked her to check her Facebook and she doesn't see any car ads. I checked too but didn't see any ads for cars or that brand, but I don't have the Facebook app on my phone (she does) and our Facebook profiles aren't strongly linked (i.e. she's not listed as my "wife", we just friend each other). She uses her phone for navigation while driving, so it was in a position to clearly hear us.
I haven't done any research on that car brand, but I suspect that once I do a Google search, then the ads will start flooding in.
So maybe this is confirmation bias in the other direction, but I don't see any evidence that Amazon Alexa, Facebook, or Google Assistant are spying on us. Though it could just mean that this particular carmaker doesn't purchase ads based on keyword spying
Try doing different things. I don’t think name brand vendors do pervasive audio surveillance. I do think they broaden the scope of your intents. Use GBoard dictation or similar tools to write. Write stuff down in different contexts. Use apps in different ways.
Amazon and Facebook share in near real time. Anything you do in a consumer Amazon property is feeding context to FB.
You are just not noticing these ads until it is something you have called out. You would never have given a second thought to this totally random item otherwise.
I have tried the same test as you multiple times just for kicks and have never found it to be confirmed.
They don't have to collect information only directly from their FB/Instagram/WhatsApp apps: what they can do is buy information from other companies that publish thousands of "free" apps on appstores.
You have to wonder how so many of these free apps seem to sustain themselves since GDN advertising does not seem to be profitable enough.
FB group should be obligated to disclose whether they are buying information from these kinds of third parties.
More importantly they should disclose whether the price they pay is illogical, effectively making them silent partners in an indirect scheme to access your camera/mic information, while at the same time maintaining the allegation that "we do not access your mic through our apps".
Facebook isn’t spying on your microphone because they don’t have to. They know enough about you to monetize the shit out of you from things that are out in the open. When the populace trusts Facebook and Google with nearly their entire digital lives, and the DOJ lets these giants acquire the rest without a fight, why would they need to resort to clumsy subterfuge?
You keep using that word...
[0] Some company called softtonic has a "Privacy Badger" branded extension but I am skeptical about downloading from an unofficial source.
Key quote (from 1999 - Scott McNealy): "You have zero privacy anyway. Get over it."
Darn - I hate when they're right.
She gets related ads in the web wherever she goes. She has Facebook installed in the phone while I don't.
But the interesting thing is that this is not happening to me. Never. I do not get ads for scuba diving suits if we speak about it. She does. How can we explain this?
My tune would obviously change if that data were used in more malicious ways. But as long as it is advertising targets, I personally don't care.
I might be biased because I run ad campaigns on FB and the data is helpful (more efficient than Snapchat or Twitter where they have less data on you).
Then there is the the combined effect of everyones data that is threatening.
I don't consider myself to be paranoid, but I have never been willing to share details of my life with strangers. I recognize that any online activity is subject to surveillance, but I do what I can to minimize sharing that. There's a lot you can do along those lines, with really fairly minimal effort -- though I consider "minimal" to include not having a facebook account and not having any photos of myself on the internet, for example, which I know from having these conversations in the past is for many people some insurmountable hurdle.
Anyway, why would I want strangers to know the details of my personal preferences and tastes? There's no benefit to sharing it as far as I can tell. That is what has always stumped me when people say they don't care about their privacy.
Diamonds are not actually forever, but data is. There's little you can do to prevent Facebook from using the data you've already given them in malicious ways tomorrow. Throw in an economic crisis or major war and it's almost a certainty that they will do so due to desperation or legislation.
The one that springs to my mind right now is: Knowledge is power; which is also true in that case: the more an entity knows about you, the more power it has over you. And not only blackmail, but recent hints (AI-powered election meddlings, addictive user interfaces, etc.) have proven that your instinct can be, and will be used against you, whether you are conscious of it or not.
And of course, there are problems linked to physical security (if the wrong person can see you're spending a week abroad, your house might be broken into, or worse).
Let's not dwell on the ethics of financing a company that sells your personal data.
Private data is by definition not supposed to be made public, it is sensitive. Treat it like an attack surface: the more there is out there, the more likely it is that something can be used against you.
Once your data is collected, you have no control over it any more. Are you sure there's a valid retention policy, it is actually working and enforced, there are no secret agreements with e.g. the national security apparatus, and it is absolutely secure from malicious employees?
That is the big issue around data collection. Even if you're fine with it being collected now, you might not want a recording of this at a later point. For many different reasons. But a big one being that it's really easy to twist your words against you.
Anything that I want to keep private, I do as a compartmentalized persona. Separate hardware. Separate LAN. Separate Internet connection path. No overlapping Internet activity or interests.
Okeydokey, machine-learning, moving on.
2. Wouldn't surprise me if part of the app were heavily reliant on things that can be updated remotely. Chunks of big apps will sometimes be just views fetching some web components. Facebook created React Native, and iirc it can be updated dynamically, like a web site.
When you have lots of people working on an app you have a high probability of introducing bugs. I worked at a company that shipped a faulty update; it looked OK to users, but it was essentially DDoS-ing the servers. Having to wait for the App Store to approve your app to fix things like that is annoying and costly, so people tend to look for alternatives.
on wsj.com privacy badger and ublock origin go apeshit..
And that sort of explains everything.
I kept wondering how the info that keep to the Google-verse (search, Gmail) makes it to Facebook. Now I know.
Does it mean that Google collaborates with these data brokers? While it doesn't harm them directly, it seems like a myopic thing to do, arming a company that may undercut your sole major source of income. Yes, I imagine they take it directly from the device, but doesn't Google have control over the Android internals?
Also, I am in somewhat unique position. Being a Microsoft zealot, I still carry a Windows phone (v8.1) with slowly dying services. I almost don't use Facebook yet I still get these too relevant ads, mostly according to what I google on my desktop.
When it comes to the US and global content, there is virtually no difference in the results. Google is much better in the local content and the knowledge graph results though, as well as the maps. Video search is better in Bing.
When I worked in the OEM industry, there were rigorous standards and compliances that had to be met to ship our phones. If an app was invoking audio recording without the proper permissions, this would be a huge red flag. Google would never approve the phone to be shipped as it would be breaking their CDD.
Also, if audio data was being transmitted using some obscure APN that does not use mobile data or Wi-Fi, OEMs could still easily detect these from the modem side, no matter how encrypted the data is. After all, the app is still JUST an app, in the system folder with all the other apps like Candy Crush. Unless this specific audio recording feature was built into the Android framework, I will say it is 100% impossible.
Note: this is just for Android. I have no idea how iOS works.
You should check your sources again - that statement is false.... You only have to look at the release notes from the dev preview on Android P today to see that was possible for any background app to access the camera and microphone at any time:
https://developer.android.com/preview/behavior-changes.html#...
> Android P strengthens privacy by limiting the ability of background apps to access user input and sensor data. If your app is running in the background on a device running Android P, the system applies the following restrictions to your app: Your app cannot access the microphone or camera.
Note that this use case was extremely rare anyway, but I know Google did this simply to ease people's minds (I know several people on the Android framework team).
Call me cynical but surely there's some top-down pressure about adding a simple $0.07 component to $1000+ devices.
I personally skip step 2 because I'm addicted to being able to quickly search for things, read blogs. But even so: a) don't login to any websites, b) clear website data often, c) minimize use of apps, d) recall that when you browse the web from an app you're using the same cookie jar and other state as the main browser, and that you're letting the app track you.
If anything, go out of your way to ruin the creepy tracking, analytics and bullshit metrics. Add as much noise as possible to their databases.
What I always say is that this would not be possible. To constantly listen for audio, process it, and upload the results would use so much data and battery that you'd notice. Only a plugged in laptop/desktop could get away with this.
Am I wrong? Would it even be technically feasible for Facebook to listen all the time?
What people accuse Facebook of doing would require constant processing of the audio stream for random interesting tokens. Either process on device and then send metadata (more battery, less network) or send the raw stream and process on the server (less battery, more network).
A few years ago I uninstalled the official facebook app because it was using all of my battery and a fair bit of data, even when I wasn't actively using it. The battery usage record in the settings indicated the facebook app was the cause. After removal it felt like I'd bought a new phone, everything was running so much faster. I recommended the change to friends who saw similar results.
I don't know if it was maliciousness or incompetence behind the battery usage, but it would seem that you can get away with it largely unnoticed.
Processing and uploading doesn't have to happen all at once. They could analyze/process the audio while you're asleep. The power drain would be negligible - or completely unnoticeable if your phone is charging while you sleep. Data obtained from these processes could be uploaded in chunks, hidden inside the photos and videos you upload from the Facebook app.
So, in my opinion, not only is this all possible, it'd be not too difficult for Facebook to hide it from the general public.
You shouldn't assume that and actually do the math. GSM-HR (an early 90s codec that is fast to encode and designed to be low-power) combined with a trial noise gate will only require[1] uploading <100 kB/hr.
> Only a plugged in laptop/desktop
Audio compression existed before MP3/AAC. GSM-HR only uses 5.6 kbit/s, and was designed to be encoded on early (GSM) cellphone embedded CPUs/DSPs. Modern devices have several orders of magnitude more CPU/power/bandwidth available.
> all the time
It's only necessary to record when voice is actually present. A trivial noise-gate[2] will cut the recording time down to only a few percent of the time at most. Obviously a better filter should be possible.
There are many ways this could be covered, as many have said here (Above & Below). Noise gates... other Apps listening for key phrases etc.
E.g.
A Fitness App could be installed with a noise gate + setup to listen for specific fitness related terms... that triggers a simple pules to say person A is interested in that topic....
Another App does the same for topic B... and the list goes on...
That's just off the top of my head. (I don't work at Google)
I regularly go in the storage part of apps, and despite the fact that I don't use many apps, they still manage to generate megabytes of data, which are not clean by the "clean cache" functionality.
I'm not using it for texts and calls, I only use it with wifi, but still, I have low trust in the android ecosystem.
AFWall https://play.google.com/store/apps/details?id=dev.ukanth.ufi... patches iptables to block domains used by apps that are not allowed network access.
I did Android QA for a VPN/Proxy app. Had to benchmark battery usage and verify that we upgraded to SPDY properly. A lot of rooting, and tcpdumping. Eventually we got a WiFi network set up by sysadmins that mirrored traffic to one Ethernet socket. Plugged it in on a local SSH-accessible machine and I could tcpdump the whole WiFi traffic without fiddling with devices.
Facebook and Instagram are awful network hogs. Both send a lot of packets every 5-10s when the screen is on to Facebooks tracking domains.
I removed both and haven't looked back since I saw that.
Google and other tracking companies of course are not better, hence the apps mentioned at the top.
This is not a viable solution long-term though. They will find ways to eventually circumvent these restrictions.
Now I'm thinking about setting up the same thing I had at work at home, just to see what kind of crap goes through my network daily.
You're on Windows 10? Guess what Microsoft does. Install Wireshark and check for yourself how much MS domain hits you get. Not to mention HTTP traffic from random apps and/or websites that circumvent privacy protection domain lists.
The implication here is that you're the product but not for that company or another. You're for all of them, ISP included. The more crap leaks on your network the more they know and sell.
As a side note, I have no idea why DNS traffic still goes in pretty much plaintext...
Have you heard about tracking data exchanges?
The overall problem is that everyone is spying on you on multiple levels because there's money to be made on profiling. Facebook just grabs the headlines. Apps listening to you are scary but the fact is you don't have to speak to be heard by them.
How is this not the greatest threat to democracy in the entire history of democracy?
https://medium.com/@caityjohnstone/no-there-will-not-be-any-...
Edit: removed a section of this comment because it seemed too much to me.
Interesting piece, but was the rent of that place $4800 per month? How does that make any sense?
That is to say: it's an absolute bad thing and is a concern. But because it's so easily visualisable and conceivable, the actual threat is blown way out of proportion, and much less sexy threats that are much more important are ignored.
You're more likely to die slipping over in your bathtub than being killed by terrorism, etc.
I guess the "slipping in your bathtub" for privacy would be the consumerist / capitalist / deregulated culture that allows, normalises and incentivises trading away privacy for convenience with no real pushback. Which in turn actively discourages keeping your privacy, because it's like: why play some constrained rule-set that disadvantages you and no one else plays by?
Today it's no longer a paranoid fantasy but a reality, yet even those who aren't in denial about it are mostly not choosing to opt out of the surveillance.
More places should do this.
I know of a few places that have simple stamp loyalty cards. Oftentimes they are mom and pop shops, or chains still finding their footing. Most cards are electronic and want a name and phone number associated with them.
So, all incentives are not against consumers. The merchants avoid the credit card oligopoly fee and pass on the rewards to consumers.
I have no way of proving this but I'm pretty sure they used my camera to read the card because the next thing I knew I was seeing ads for Starbucks (a competing coffee chain). I have never said Starbucks out loud and suddenly I am seeing ads.
Loyalty apps/datasets literally print money in comparison.
FB then cross reference with any other devices nearby and any identified objects within the photos are cross referenced with their ad inventory.
Murdoch has money to make by gaining leverage over FB.
I haven’t had a Facebook account in almost 4 years and I block their traffic on my router, so I have no love for them. But I seriously doubt this story’s motives.
A "like" or "instagram this" widget is all they need to track basically everything you do even if you don't have fb.
ad display => subconscious influence, nudge => eventual product purchase, mention etc => recognized ad => spooky feeling
Do this: lie. 234-567-8901 works at a lot of US stores. Sometimes you can even use it for gas discounts.
If this behavior disturbs us so much, why does anyone continue using Facebook?
That said, this is a quite silly conspiracy theory. Believing the "recording" theory requires you to believe that Apple and Google both are in cahoots with Facebook to give them a rootkit-like API for mic and location access that override all of the OS-level controls and warnings about when that hardware is in use, and hide it from the data and battery usage stats on the phone.
This API, when (not if) found, would be a watershed moment for privacy legislation directed at all three companies. Little to gain, potentially the whole farm to lose.
Seems like you put conspiracy theories in two boxes. So what's on the other box?
You mean like when they gave Uber special access to grab screenshots even when the Uber app wasn't running? Yeah, totally impossible to believe.
I assume they are doing it through toilet cams.
Is this a rehash of the Reply All episode?
If you bring this up with someone who works in adtech, the stock response is that we need even more tracking to fix it. It's like the adtech version of "no true Scotsman".
Edit: That's not to mention what a huge waste of money this must be for the client buying this ad space. After ordering something from an Adafruit-like electronics shop, I saw ads for them around the web for quite a while. Often trying to sell me the very thing I had just ordered.
How did I find them in the first place? Second or third page on Google search.