Twitter Followers Vanish Amid Inquiries into Fake Accounts
nytimes.com
nytimes.com
I realized that API query results were mostly news bots, retweet bots, corporate PR bots, social media aggregator platforms like Buffer, and just plain old spam bots.
How bad was it? After filtering 1,000 tweets per query, I barely found 10-20 real human users. That signal to noise ratio is dismal, and detrimental to the core product experience. Twitter must be forced to maintain this fake high activity to prop up the share price.
BONUS: Guess who else is spamming their post feed: Tumblr. Tumblr didn't allow any adult content or keyword search; since Marissa Mayer took over she seems to have loosened that policy to fluff the numbers. Tumblr today is drowning in porn.
Twitter does take a pretty dim view of DM and @reply spam though because it is genuinely annoying and not opt-in.
Careful here. This appears to be referring to the concept of fair use, but firstly there are multiple criteria used to judge fair use (https://en.wikipedia.org/wiki/Fair_use#U.S._fair_use_factors), secondly these criteria are (intentionally) subjective and up to a judge's interpretation, and thirdly fair use is US-only.
That seems fine. In the interest of their userbase, companies should aim for the least copyright-damaged user experience possible; this means picking a single country (ideally one with liberal copyright law) and ignoring copyright law in other countries they aren’t based out of. If countries want to force their censorship standards, they can at least be honest about it and block the website (rather than silently deflecting the responsibility of censorship to the website itself).
I'm not saying they can't skirt the laws a little bit and get away with it, all I'm saying is they need to have awareness.
Due to the Safe Harbor rules in US copyright laws they don't really need to remove copyrighted content proactively. There's tons of subreddits exclusively dedicated to piracy that they don't care about.
They brought the heat on themselves when they decided to mock DMCA requests and cease and desists instead of accommodating them.
There's a reason 4chan, Reddit, imgur, Tumblr and its ilk all still exist in the age of copyright. None of them produce their own content-- it's all user submitted, and mostly in violation of some copyright or another.
Generally speaking, my understanding is that hosts aren't liable for user-uploaded content unless they curate or promote it (deletion notwithstanding). That changes their role to that of a content distributor/publisher instead of a mere platform. This is why backpage's CEO got arrested for human trafficking whereas craigslist's did not-- when challenged craigslist shut down its prostitution ads, but Backpage actively reworded and posted them and in doing so became their publisher.
I've used it for a while, and what I got is that it's goos for (and people use it) to spam others about your projects or show off. However, if you try to use it to get news or updates on anything it is the least efficient, most stressful thing I've ever used.
I see Twitter as a good tool for outages, natural disasters, and protests. That's pretty much it.
Depends heavily on the set of people you follow. I've found it to be a great source of news, and I typically see news show up there hours to days before I see it show up in places like HN.
I have no idea how people follow more than say a hundred people profitably.
What I find awful is that HN and Reddit save me time by sorting out low-quality content, while on Twitter you're the one that has to work hard to do it (if you can), and Twitter itself just adds more noise in the meantime.
Twitter is what you make it. I'm sincerely worried their VC backed top heavy company is going to topple over and carry with it to oblivion a service that could probably run on 1/20th of it's infrastructure.
That said, it took months to get my feed to where it is today. It's not easy for each person to find the mix of accounts that is best for them. My recommendation is, be fast to follow folks who look interesting, and fast to unfollow folks if they are boring or you don't like them. When you find folks you like, see who they retweet, reply to, follow, etc. and follow all those folks to see what you think.
It's called ESPN.
I guess you can find other people's lists and follow them? E.g. a sports journalist's sports lists etc. I have not tried to do that so not sure if there are hurdles to overcome. And maybe someone can aggregate these public lists for others to find/follow?
This is particularly relevant for hyper-local media which basically has zero media footprint outside of dead-tree newspapers and talk radio.
Basically: follow your city councilor on Twitter. Start from there.
But I see it as a way to keep track of what a columnist, journalist, or public figure is doing.
'Drowning' in porn? In the sense that that's a bad thing?
...what?
Tumblr was known for porn long before it sold to Yahoo. If anything, Tumblr started cracking down on blogs with adult content afterwards (for example, requiring users to log in before visiting them). There was a huge backlash from artists and bloggers with non-pornographic gay-themed content, because many of them were caught in the ripple effects from these changes.
I had friends who worked at Tumblr before 2012, and the running joke was that Tumblr was 50% porn and 30% pictures of cats. (Don't take those figures too seriously, but clearly porn was a large part of Tumblr long before the acquisition).
Also, Tumblr definitely did allow adult keywords in searches. I know this because of a rather unfortunate incident at a student hackathon my company sponsored, in which a student thought it would be a "funny" idea to search for a risque phrase when demonstrating his weekend hack (an aggregator of Tumblr posts).
At boarding school many moons ago, we all used Tumblr to get our rocks off because most porn sites were blocked.
Thankfully more and more websites are based out of other countries now and don't need to swim in the puritan current of the US.
I used to have a much bigger problem on Twitter with fake accounts and bot likes than I did, but despite my decades of willingness to bitch about spam, I have to concede that they've gotten better lately.
There's also a user benefit to letting bots run when they're not hurting anyone: you don't give hints to the spammers on how you're finding and nuking them. Indeed, if you identify them but don't block them, you can get a lot of data on what is spam. For example, on my mail server, I noticed I was getting a lot of dictionary attacks. So I took the hundred most common first names not in use on my domain and fed them all into the spam training system. That means odds are very good that my spam trainer will have seen a piece of spam before they try my actual account.
Very nice!
Likes, Followers and accounts yelling into some hashtags can be dominated bots. But if you look at the feeds of individual people and those that they talk with that's a totally different experience.
> Tumblr today is drowning in porn.
Tumblr has had NSFW for a long time now. If anything some of it may have migrated to patreon now that it is easier to monetize.
Twitter's model is basically "shouting at the universe" but if nobody listens it doesn't matter.
I 100% guarantee twitter knows the exact ratio of real to unreal users and their content.
They're also trying to slowly crack down on bots.
A year ago they insta locked your account if it looked like automated liking of content.
Now they're insta locking if you look like a follow/unfollow bot.
Its a balancing act and I feel like its analogous to getting out of extreme debt. They're trying to replace bots as their real user base grows internally.
I don't know about Tumblr wrt porn - my impression is that there's a lot of sexual material there but that it's community-driven rather than commercial, partly due to the demographics of its user base.
Would also be interesting to run similar analysis in FB, I'd expect similar rates here.
For this reason, trending topics and keyword search are essentially hijacked features.
However, because I believe there are many useful bots that have organic followings in the millions - I don't believe they need to be simply removed from the ecosystem.
Instead, my suggestion would be a 'bots' account type. Some ideas:
- A robot version of the 'blue checkmark'. This would allow users to quickly identify a tweet as sent from a bot.
- This account type could be linked to a real owners account, much like Twitter apps are. Accounts flagged and failed to register as a bot could be subject to deletion.
- Bots would automatically receive low ranking in search queries, and trending topics. Perhaps they would be completely delisted.
- Bots cannot follow other users.
- Bots cannot tweet at* other users.
More extreme:
- Bots cannot tweet without some sort of spend. Maybe they can only tweet in some ratio from real Likes they receive. This is a bit extreme, but would mitigate a lot of problems.
I really believe a happy median could be found - and currently think that a well-curated Twitter timeline is amazing, but as I stated search results and trending topics are completely broken.
Weather bots. For any city within 100 miles of where I am. Plus bots posting job listings. Plus companies posting those same job listings.
There was basically no signal to find, it was all noise. The few ‘legitimate’ ones I found were from local PD/FD.
I think they just need to ban all bots/automated postings. Or make them filters le and require a $100/mo account and $1/tweet. Something to discourage the absolute garbage.
The other issue is that the distinction between a bot and a human isn't clear-cut: there's plenty of shades of grey in between something which spits out badly-scraped listings all day and an actual human having noncommercial conversations with Twitter friends. It's more classifying what's bad behaviour (using the pornbot tactic of mass follow/unfollow to attract attention even if you're a human marketer tweeting actual content, and even if you haven't written a script to do it) and what's perfectly acceptable automation like tweet schedulers that could use some work
I would dispute that. I'd argue that the bots exist because spamming twitter is free. If it costs me $0 and I get even a tiny benefit out of it then it is to my advantage.
Its essentially the same problem as spam email.
I think this is a legitimate use of a bot. We even mention the host club because they want us to.
What's funny is that the account got squelched 3 times before we got a human at twitter to officially prevent us from getting flagged. So they do definitely have some measures in place to prevent spam accounts. I suspect it's become non-trivial to identify all the bad actors.
I know lots of people also use bots for cross posting from Instagram or something else. Or to post when they put a new article up on their site.
I’m sure you’d have to allow them to some degree. But there are some really noisy bots out there that need a fee attached to em.
We do request access so you can send Direct Messages from your dashboard. We are considering removing that functionality for the sake of privacy.
Unfortunately, the Twitter application permission model is not granular:
https://developer.twitter.com/en/docs/basics/authentication/...
We would prefer to just have write functionalities and not read, but this is not possible in their model.
I follow specific people that I've discovered outside of the network, or by referral, and as a result my feed is basically "all signal".
This seems like the only reasonable way to use Twitter now..
I am kind of surprised the number of human users is so high. I know a lot of bloggers of various sizes. To the best of my knowledge virtually all of them have hooked up one of the available services to post on their behalf. They either spend a little time scheduling out their tweets for the next week/month then forget about twitter until their schedule runs dry OR they set something up to randomly pick from some pool of blog posts and spam links.
Either way they then essentially never go on twitter again once things are up and running. The whole thing is full of bots talking to each other.
How many of David Brock's bots have been removed from Twitter?
Doesn't really work everywhere in the world.
What about my pet open source projects, shall I shut down their account because they don't have a company backing them?
What makes you think Facebook isn't prioritizing this? I'm not questioning whether you're right about this, just curious about what made you draw this conclusion.
Someone I am intimate with left Facebook in the last month, in anger at the way senior management up to and including MZ first dismissed the severity of the problem of their platform being used to push propaganda during the election cycle.
And then stonewalled and foot dragged efforts to correct this. By the account I got the only incentive for action has been unflattering attention at the national press level.
This from someone with a personal relationship with Z.
Make of it what you will. The poison is real.
Intelligent and reasonable people can disagree - strongly as to the extent of how much of a problem this is, whether or not it is a problem, or whether the cure is worse then the disease.
This also has little to do with fake user accounts.
I haven't seen a claim that fake users played a large role in spreading propaganda or fake news. As far as I understand it the issue lies in what people were sharing and what Facebook identifies as trending news.
That's a pretty different issue from Twitters fake accounts.
I've seen fake Twitter tweets as well that have bad spelling and other stuff in them that it is obviously a fake tweet made by some website fake tweet generators out there. Then shared on Facebook as an uploaded image. If they don't link to the original tweet on Twitter consider it a fake tweet.
"So how many of them are real and how many of them are fake accounts?"
"We work very hard to make sure Facebook is free of fake accounts."
"That's not the answer to my question.".
Disclaimer - I work for facebook.
Twitter doesn't even take them down if you complain about them. Very annoying.
If your profile is public, and they decide to follow it, in part, to make you aware of their existence, perhaps that's a feature not a bug, given that the main use for the thing is to connect with accounts and follow them and discover content.
I want people to follow me because they're interested in what I post, not because they're looking for me to follow them. Perhaps I'm asking too much of the platform, but... it is what it is.
A telling anecdote: A security researcher friend of mine found a somewhat small botnet of twitter accounts (~7000). Reported it to twitter, a few months passed and he noticed twitter hadn't done anything. So he turned it over to a journalist who eventually poked someone at twitter and... poof all 7000+ accounts were gone 6 hours later.
Maybe simply no dedicated people or team for the problem?
I have no idea what credentials I used to create it so I can't delete it, but it still exists (unlike the person it's pretending to be).
You creating a fake college account and forgetting about it probably passes as human enough, there's no concerted effort or agenda to that account besides existing and adding friends at your college.
I think what makes bots and fake accounts generally detectable is consistently pushing certain messages in ways that exceed normal human behavior, as well as showing patterns across many fake accounts.
It's hard to see a pattern in one fake account, but easy to spot it across many.
Except that Facebook knows who is real and who fake. Just like my mail provider knows who sends spam and who doesn't.
Facebook can display ads to fake users and still make the campaigner pay for the impression, or make a group pay for reaching more users, fakes included.
Facebook can easily filter requests from fake profiles, so no, fake users do not necessary worsen the experience for real ones.
You say that so definitively. Have you worked on a product with millions of new users a week? It's an extremely hard, constantly shifting problem.
"Facebook can display ads to fake users and still make the campaigner pay for the impression"
Facebook's entire business relies on user and advertise trust. Why would they sacrifice that for some short term growth that would inevitably kill the business by eroding trust?
Facebook has hundreds of engineers and enough data to do match patterns against. Facebook can largely identify who's who.
> Facebook's entire business relies on user and advertise trust. Why would they sacrifice that for some short term growth that would inevitably kill the business by eroding trust?
Yes, the infamous "they trust me, dumb fucks"?
"Yes, the infamous 'they trust me, dumb fucks'?" You didn't do or say stupid things when you were 19? You don't believe in giving second chances, let alone to a teenager?
For starters, encouraging advertisers to spend money to build up an audience with the assumption they could continue reaching that audience much like email, only to throttle organic reach to zero was pretty bad.
There have been other things such as being extremely...generous...with the definitions of how some ad metrics are defined and what defaults are presented. Even as an experienced advertiser who knows to look for those things, the lengths to which some of it it is buried is astounding to the point of it being hard to trust that it wasn't intentional. And the recent lawsuits around such things shows I'm not alone in that feeling.
That sounds like a sweeping statement with little forethought when you're talking about a platform that enabled protests in dictatorial regimes.
Step back here: Fake users, sold by the thousands to give credibility and a megaphone to whoever shells out money: bad. Identity verification: A solution with many consequences to be weighed before jumping on the bandwagon.
And dictatures aside, I have no desire to give Twitter my ID or passport.
The ability of state actors to use social media to spread propaganda should also be considered in our tally of social media impact. The jury is still out on whether they are a net good.
And there are many people who lack a driver's license and passport. Are they disenfranchised?
What prevents me from using the same id in bot accounts then? Verification implies they have to keep some form of personal identification.
We don't need a passport or state-issued ID to determine if someone is a real user. That's the laziest solution any tech company has ever come up with, and it's deliberately lazy.
Facebook and Twitter can already tell if you're a human or not, shit Google lets you just click a button to tell them you're a human and then they determine if they believe you or not. All based on info they already have.
We're the best software engineers the world has ever seen on the cusp of an AI revolution... we don't need your passport. We just don't want to find out the truth, so we make it so hard no one will do it.
Best of the best, top of the class is not infallible. It's good enough to sell, because even something as low as 80% accuracy is good enough for things like "Do you want to subscribe to weekly pop tv news", it's not good enough when at the other end you get your account banned.
Google has, on this very site, built up a horrible reputation of using automated processes and having too many users to give those affected by those processes some good recovery. What you're describing is a recipe to replicate that.
99% is not good enough. 99.99% accuracy gets you 0.01% false positives. That's crazy low, right? It's also 100k users when you have 1bn users. Those fancy AI processes are nowhere near that.
I think a lot of their abilities are overstated.
I've met three people in my life who were employed primarily to place ads on the internet. All three were glorified secretarial drones who fell into the position because nobody else wanted to do it.
Any person who knows how to effectively place Internet ads has such high-status skills that they don't have to actually do it.
I'm increasingly seeing this annoying trend of people hand waving about how a given problem can be solved with data/AI/ML/DL.
Maybe some of you are, but people like me read HN too ;-)
- now just imagine the speech was pointed at the bad actors. and that's why this is not as easy as you make it sound.
We^W They may have shitloads of data, but I'm very skeptical we^W they can use this data in any meaningful way. At least, for now.
I don't know what's wrong but despite all that data giant AD companies are supposed to have, I never ever saw any relevant ads, even when I've specifically wanted to see one. Unless I've already bought something - then, sure thing, I'm spammed with more of the same (which is, again absolutely useless - I've already made a purchase).
> We don't need a passport or state-issued ID to determine if someone is a real user.
I don't think we are even capable to come to a consensus what does this mean - to be a "real user".
Am I? What about my alter ego, posting about something I don't feel like publicly associating with (like, porn)? What about a whistle-blowing throwaway account I may create if I learn something fishy? Or what about a "thoughts are my own but went through editorial" corporate representative persona account? And that's just the obvious cases.
Relatedly: if you participate or "observe" any online communities of marginalized people, you may have noticed there are starting to be earnest conversations of whether or not "free speech" is actually more easily wielded to abuse them, than they can use it to have a voice.
Just food for thought. A lot of built-in assumptions are being tested right now.
Let's assume for a second Twitter has a magic wand that, without consequences, automatically deletes fake users used for marketing purposes.
Now what does identity verification actually do for the remainder of the users? In what way does it benefit them? In what ways does it inconvenience them? In what ways does it put them at risk?
Should there be a process in place, that is required of the entire userbase, just to fix the fake user problem? A problem that most users aren't even sort of aware of (otherwise the NYT article wouldn't have been as successful as it was).
We do a lot of reactionary things as a society that follow this exact pattern: Put massive processes in place to remove tiny bits of risk. Governments do that and justify it all the time, cf. the "Terrorism" or "For the children" memes. You often read people's complaints here on the ridiculous security theater that the TSA is, how it doesn't solve anything and inconveniences everyone for something that happens to people less often than winning the lottery.
This, is that. The pattern is: You have a tiny problem, you fix it with something which impacts your entire userbase because you didn't stop to think whether there are more targeted solutions, or even whether it's worth it.
[Note: I argue lower down that AI and statistics aren't the correct fix for this either... gotta keep digging. Maybe the "correct fix" is a mix of identity verification and statistical analysis.]
That exact thing is happening. The Reply All podcast did an excellent piece on the dominant political party in Mexico using armies of people on Twitter - not bots, but people - to drown out news they didn't like, and eventually, to harass people. Show, with full transcript: https://gimletmedia.com/episode/112-the-prophet/
Authentication adds a potential cost-of-reputation to social media content that you post, which in turn may raise the value of the content. If there were a safe, secure way to do it, then why not?
It seems like it only matters from an investment stance.
-Upton Sinclair
Twitter’s main strength is its low barrier to entry, for end users and developers alike.
Fun story, years ago in my office, before buying followers was well known, folks would prank each other by buying fake followers for our co-workers. They'd wake up and be so happy and surprised and then have to spend the weekend manually blocking each one.
At the time it seemed pretty harmless, but now it is definitely a threat to someone's credibility.
I was on Uncyclopedia when someone did that to get them removed from Google.
"Link to this post in three years and ponder in its prescients... Web 3.0 will be born in the death of the heavily botted social networks."
I got a new iPhone and on iOS Safari I needed to sign in to all my accounts.
Except for some reason Amazon recognised me with one-click enabled. I never used one-click to buy anything before and accidentally bought a kindle book while browsing the site.
More strangely, when I went to turn off one-click in my settings I was forced to log in.
So I could one-click buy without explicitly authenticating, but needed to authenticate to disable it. Very strange and/or shady.
Btw - is there an easy way to cancel an accidental one-click buy? In my case it was a local author I wanted to support anyway, so I’ll keep the purchase. But surprised it’s so easy to accidentally purchase something from the mobile site if you swipe to scroll on the wrong place.
https://blog.plan99.net/did-russian-bots-impact-brexit-ad66f...
The problem is that some phrases that appear on the surface to have one meaning have been grabbed and redefined by particular political groups, almost used as code words. "Bot" and especially "Russian Twitter Bot" for example isn't used by Twitter or others in the way you'd always expect:
https://www.projectveritas.com/2018/01/11/undercover-video-t...
"Just go to a random [Trump] tweet, and just look at the followers," Singh says. "They'll be like guns, God, America, like, and with the American flag and like the cross. Who says that? Who talks like that? It's for sure a bot."
The idea that Twitter bots can change society in fundamental ways is one that seems to obsess journalists, who all seem to spend half their day on Twitter anyway, but I've yet to see evidence that it's true.
The graph isn't a "this is how many followers the user had at this point in time" but rather "right now, here are all the followers this user."
A point is "the Xth user following had a join date of Y" and thus the X axis is in effect a time axis (though not a linear time axis).
From an eye scanning view, the horizontal bands are easier to follow and notice than vertical ones - and that is part of the goal of the graphic (to emphasize those bands).
It's a kind of weird visualization, but the patterns which emerge are pretty striking (and obviously artificial).
For those not familiar with the graphs, the graph shows date followed on x and date created on y. Where there's horizontal stratification, NYT noticed it means the follower accounts that all began following the account at the same time were also created around the same time, indicating a strain of bots. However, when there's vertical stratification, there's a shift in the rate at which accounts are following the account. When the vertical stratification is followed by horizontal stratification, it's an indication that both sets are bots - for example, one scenario is that instead of providing a mix of bots created at different time, someone got lazy and just grabbed a list of bots all created at the same time. However, vertical stratification could also just indicate that the person did something good or bad to change the rate of acquiring new followers so it isn't clear cut. That being said, I'm not sure that justifies labeling the initial section which lacks horizontal stratification as organic.
We’ve had AIM, ICQ, MSN, Yahoo Messenger and many others that can be seen in the list of protocols supported by pidgin.im. Now we have Facebook Messenger, Whatsapp, Skype, Hangouts and the thing from Apple.
Or at least there should be a standard everyone that wants to sell to public institutions should follow for instant messaging.
1. Twitter is selling content feeds to TV news networks. With enough bots out there your tweet may very well end up on TV.
2. It counts for SEO. I once met a guy in 2010-ish who was into casino SEO. He was running a network of ~150k FB/Twitter accounts to promote articles that quoted the oddball news outfits that quoted his clients' press releases.
3. Some people actually read what's going on on Twitter. In particular journalists and swaths of opinion leaders. See any late night show, really, for ample Twitter coverage.
4. Some people with tons of followers occasionally retweet garbage memes on Twitter, including racist videos tweeted by white supremacist UK groups that turn out to be fake.
The problem being that even if you yourself don't view or access Twitter, it is influencing, and largely for the worse, the world in which you live.
two young siblings... earn a combined $100,000 a year as influencers, working with brands such as Amazon, Disney, Louis Vuitton and Nintendo. Arabella, who is 14, tweets under the name Amazing Arabella.
But her Twitter account — and her brother’s — are boosted by thousands of retweets purchased by their mother and manager, Shadia Daho, according to Devumi records.
The idea that dubious businesses are possibly skimming money off big brands (or even more disturbingly, possibly with the knowledge of those brands) by using fake followers is interesting, and newsworthy. Add widespread identity theft into the mix, and it’s even more newsworthy.
They're just looking for any blame they can shift for their culpability in helping throw the 2016 POTUS election to Trump.
This has been a known problem for a while and Twitter has done little to fix it at scale. Given that they make their money on these paid “engagements” there are going to be a lot of people taking a real close look at this. Interesting days for Twitter ahead.