On the hunt for Facebook’s army of fake likes
benthamsgaze.org
benthamsgaze.org
One effective strategy we've employed not mentioned here is category mapping: if an account of type A, only targets accounts of type B for likes (especially if they ignore categories C, D, etc.), this is usually a high indicator of fraud. For example, one very common strategy is to create a fake account for an attractive female to friend many male accounts (especially relatively new accounts unaware of these tactic). This can be easily detected by analyzing the gender and account age of all targets and coming up with a diversity score. Low diversity score = likely fraudster.
The incentive to make fake accounts on Facebook is orders of magnitude greater than almost any other social network.
> One effective strategy we've employed not mentioned here is category mapping: if an account of type A, only targets accounts of type B for likes (especially if they ignore categories C, D, etc.), this is usually a high indicator of fraud. For example, one very common strategy is to create a fake account for an attractive female to friend many male accounts (especially relatively new accounts unaware of these tactic). This can be easily detected by analyzing the gender and account age of all targets and coming up with a diversity score. Low diversity score = likely fraudster.
Facebook has methods that radically exceed this method in both complexity, precision, and recall.
Not true. In the article, the writer pays Russell $15 for 1,000 likes. Being generous and assuming each of Russell's fake accounts can farm out 100 fake likes, he's making $1.50 per fake account before it gets shut down. Compare that to social networks where you can directly extract payments from other members by listing fake items for sale, laundering payments from fake credit cards (on other fake profiles) to yourself, or link-baiting other users. A single successful fake account on those networks can easily net you $100.
> Facebook has methods that radically exceed this method in both complexity, precision, and recall.
Agreed, and indeed Simility's models have much more complex methods too, but a) I wanted to post an interesting example everyone here would understand and b) I still say Facebook is not using anywhere near its full ability to stop these fake profiles given how rampant this fraud scheme is on their platform. (Again, follow the money, FB has very little incentive to stop these fraudsters who are only inflating their own numbers. It's important to keep them in check, but there's no incentive to waste resources stopping them.)
You're comparing apples to oranges: fake accounts used for "like spam", etc. are different with regard to their complexity and scalability than accounts used for phishing. There are phishing accounts on Facebook as well.
> Again, follow the money, FB has very little incentive to stop these fraudsters who are only inflating their own numbers.
Facebook has a massive incentive to stop fake accounts: fake accounts decrease meaningful conversions that lower the ROI for advertisers, which is tracked carefully both by Facebook and advertisers. This directly lowers the price for ad space on Facebook, and makes Facebook look noisier and less impactful than other channels.
Following the money leads a direct, unmistakeable path to a strong incentive to shut down fake accounts.
It's also very bad to accidentally shut down real accounts, especially in cases where users could be confused enough not to return.
I think the difference in opinion on that is that if you look at the long term, which you hope FB are, then yes fake accounts that reduce ROI for advertisers are bad. Unfortunately, they lead to a short term increase in FB ad revenue, which disincentivizes stopping fake accounts too effectively, as it may actually be a noticeable dip in revenue, depending on the scope of the problem.
In a worst case scenario, FB might be in a situation where 20% of ad revenue is from bad impressions, and completely stopping that, if they had the power, would have major negative repercussions for the company. There would need to be some hard choice made about the best path out of that situation. Not that I think this is necessarily the case, but it is an example of how the incentives may not be as clear as they seem.
As an advertiser, I am pretty certain this is not the quite case in (current) reality. A large part of FB's proposition for more money from us includes:
A) Pay more for increased reach and engagement.
B) Our traffic isn't decreasing (despite outside reports/indications to the contrary) and you would be missing a massive and engaged audience if you didn't spend with FB.
This combined with the fact that a _lot_ of ad-spend isn't directly attributable to conversions (often by design), means that more "activity" whether its fake or not, drives up ad-revenue for FB.
You see the same issue occur with other publishers by the way. -It is not uncommon for a publisher (or other related party) to purchase a swarm of fake bot traffic to boost impression and engagement numbers of an ad buy they've sold. -Advertisers un-aware of how much of the traffic to their ads are bots vs legitimate humans (read: "publishers stealing money from advertisers") is a major problem for advertisers, but the bigger the publisher, the harder it is to 'not' be on their platform too. (and FB is _very_ big)
It's not. Even the "leaks" make clear that overall traffic is still increasing, both overall and per person.
> This combined with the fact that a _lot_ of ad-spend isn't directly attributable to conversions
I've seen direct reports from advertisers at my last job (doing social media analytics) that show how well they can quantify ROI for ad spend. Fake accounts would negatively impact this number, and it would be extremely obvious immediately.
But I don't think FB has put lot of efforts. I was being targetted by some fake account which was a Facebook profile of a company (created as user). I complained and reported the user several times. Facebook has not taken any action. From what I can see a simple regex on name should tell that "Taylor Swift Lover Group Admin" is not a human being and cant have a facebook account.
I'll go a step further and give you some unsolicited advice. The anti-abuse community amongst internet/game/tech companies is actually fairly close knit since it's one of the few places where everyone is on the same side and lots of "secrets" are shared (including, even, at the spam fighting conference we organized last year). I would bet a lot of people just rolled their eyes while learning of your company for the first time. You're already entangled in one argument from someone calling you on this silliness, but I assure they're not alone. I'd probably suggest reconsidering this approach.
We recently bought some likes for a page via FBs internal system. The likes we eventually received were nearly 100% identical in terms of names, looks (mostly arabian or oriental), even though the region we targeted was within central Europe - and lot's of obvious fake accounts in there.
We didn't set it to 0.
There's sometimes positive value in spam. Ex Instagram users get a boost when their pictures are liked, by someone real or not.
"The issue with these methods, however, is that stealthier (and more expensive) like farms — which likely do not rely on fake/compromised accounts — can successfully circumvent them, as a result of spreading likes over longer timespans and liking popular pages to mimic normal users. Our recent preliminary study confirms this hypothesis on accounts used by BoostLikes.com, showing that tools similar to those deployed by Facebook (which rely on graph co-clustering) fail to accurately detect fraud."
The posts getting liked are from fan pages with <5 likes, yet the post gets 10,000 likes within minutes. Is there some reason that isn't easily detectable?
Same goes for their other product, Instagram. I want people to find my pictures I post on there but the amount of notifications I get not from interactions but from spam accounts adding me and liking my photos is very annoying.
I feel bad for whoever has to handle the reports on the other side.
I think a lot of the problems people have with Facebook are based in how Facebook gets used. If you have 400 friends, you're simply not getting the same experience out of Facebook as a person with 60 friends is getting.
Maybe he does too
I get a few more than that, but the pattern is always the same. Profile picture is of an arguably attractive girl and we have a mutual friend. The mutual friend, in my case, is always one of a few guys who historically hasn't done well with women.
As for IG, the amount of spam accounts is almost unbearable. To me it is next to Twitter with the number of accounts. IG I average about 3-4 spam (porn or sell me followers) accounts per day. With Twitter it seems to come in chunks. I'll go a few weeks without any and then one day I will get 50-60 follower notices from obvious spam accounts. With both it makes the user experience unpleasant.
They're not really selling advertising, they're selling KPIs. Facebook is the best platform because they deliver the best KPIs and so get a larger proportion of ad spend.
Think about this: Company spends $X on digital branding campaign, then to make sure they can justify it to bosses etc. they spend $Y to get the views to go with the spend. Did customers benefit from or like the campaign? That's besides the point. Companies essentially pay people to watch their videos (AutoPlay in their FB feed).
However if that ratio is not stable then I think it will be a serious problem for marketeers because we would not have any metric to determine how much budget we should scale.
This is the reason I suspect FB is going slow on the killing the fake likes. Their strategy will manifest over much longer period than usual.
If i changed the language of my fake accounts in response to this, would they perform better because they now no longer work in your filter?
That's an example inductive reasoning. It's quite flawed as a way of extrapolating from known information in to the unknown because it doesn't account for the things you don't know. In the same way, saying "Fake Facebook accounts post with this frequency or these few words" doesn't work, because it only considers the accounts that do. It can't detect the bots that don't.
Now talking about false positives or false negatives might be a more insightful conversation.
Please keep incivility and personal attacks out of HN comments.
"I've run 10,000 experiments that show toy bricks are green, therefore I can say with great confidence that all toy bricks are green."
Knowing things with absolute certainty is super hard.
"I've run 10,000 experiments that show toy bricks I found on the ground are green, therefore I can say with great confidence that all toy bricks are green." What about bricks in the toy chest? What if someone was building something green before you tested?
As this applies to Facebook's statement, all they can assert is that they've found a way to detect fake Facebook accounts that they've tested against the accounts they've been able to ascertain are fake. What about the fake accounts they could not confirm were fake, or that weren't identified in any way?
Facebook is looking at tackling farming. In that context, the vast majority of accounts are likely to be operated in the same way. They need to avoid detection and they need to do it as efficiently as possible in order to make the most profit. Facebook has now cracked this particular code and an arms race is likely to ensue. But at the moment Facebook discovered this, it was unlikely there were a significant number of farming operations doing things any differently.
No one is claiming this applies to all fake accounts (or if they are, it's definitely not in the article). Obviously a catfishing account will be run quite differently than a farming account.
You're essentially getting at why science is both inductive and deductive. Sure this article is inductive, but it's conclusion is falsifiable, so people can test the hypothesis over and over again.
Pointing out that they haven't tested the everything in the universe, including the celestial teapot, doesn't mean that it's false. It's just that it has never been falsified yet.
1: The original title was "Fake Facebook accounts can be detected by when they post and the vocabulary used", which without additional qualification I think is quite a strong statement, and has some problems because of that.
At most you can say some fake Facebook accounts can be detected by when they post and the vocabulary used, which is what I think the top level comment was trying to note.
I imagine this to be even easier on Facebook, as it's not realtime.
(The submitted title was "Fake Facebook accounts can be detected by when they post and the vocabulary used".)