Nearly 20% of active Twitter accounts likely to be fake or spam
sparktoro.com
sparktoro.com
The indicators of being a spambot they have in their post seem VERY iffy to me. "Not tweeting in the past 120 days", "Location set to a non resolving location", "Small number of followers", "default profile image", "No URL in bio or non-resolving URL in bio", "Not on many lists", "tweets in a different language than the person they're following" - Those all seem like extremely weak signals to me. My profile matches 6 of those, and I'm a human. I would like to see them hand-verify a subset of their results and see if their algorithm matches reality.
Also note that they define "active" differently than Twitter. They define "active" as having tweeted recently. Twitter gives spambot numbers as a percent of monetizable daily active users. I wonder if Twitter's given bot numbers are low because bots don't typically lurk or load ads. I can believe that the total bot count as a percentage of users or as a percentage of recently-tweeting-users is higher than 5%, but that only 5% of daily visitors seeing ads are bots.
Just out of interest, imagine you were in a hot desert. There is a tortoise in front of you. You reach down and you flip the tortoise over on its back. The turtoise lays on its back, its belly baking in the hot sun, beating its legs, trying to turn itself over but it can't, not without for your help. But you're not helping. Why?
The real question is how strong am I to flip a hundred fifty pound tortoise without injury?
Bots tweet and they usually have some sort of generic profile picture, so their methodology wouldn't even account for real bots. Bad.
Regardless, I do think that there are a lot of bots in TW and they are definitely more than 5% of total users.
Meanwhile, several accounts of mine that were predominantly run with bots would have passed as human with ease when they were active.
Are you sure you aren't a bot?
But hey maybe with these kind of analysis, and rando computer generated / un-appealable bans in the the future the "real accounts" will just mean "very elaborate bot".
This is a terrible metric. Real people use the location field for all sorts of non-location purposes, as well as more freeform descriptions of their location that wouldn't resolve mechanically.
"WTF, my account was closed because I didn't tweet in four months."
You can see it here (https://twitter.com/paraga/status/1526237578843672576).
Noteworthy highlights:
* Twitter estimates its <5% number from human analysis of multi-thousand user random samplings of mDAU
* Twitter allows that number to remain so high to avoid introducing friction like captcha into real users' experiences
* Twitter uses all sorts of internal private data in its analysis
* Parag says you cannot get a reliable indication of bot/not bot without this internal private data
Having just finished building a Twitter analysis tool, I agree with Parag that the Twitter API doesn't provide sufficient clarity to make decisions about spam. This article's analysis doesn't hold up - just because you can name several features you're going to use to generate a spam confidence score about an account does not mean that spam confidence score will have any precision.
Twitter CEO: “Let’s talk about spam, with the benefit of data, context” - https://news.ycombinator.com/item?id=31399913 - May 2022 (13 comments)
This doesn't pass the smell test in my opinion. Given that everyone who tries to create an account without a phone number has to go through the friction of getting their account locked immediately, they clearly don't care about this sort of friction. Not to mention the friction of just trying to view a tweet which has been discussed at length on HN before.
Yes
> Not to mention the friction of just trying to view a tweet which has been discussed at length on HN before.
Yes
Now imagine filling out 3 captchas every time you open the twitter app on your phone, 1 for ever time you tweet and 1 for every person you follow.
Most users of twitter, use either an app on their phone or the cookies in their browser suffer the friction of forgetting their password (likely password1 btw) because they have to log in so infrequently.
That's convenient isn't it?
This report made headlines because it aligns with everyone's experience with Twitter: almost everyone on Twitter is either a bot or a corporate managed account.
Even accounting for the breakup fee he comes ahead financially by many billions in cash, while also building his image as the free speech edgelord. Win-win.
It's much easier for the board if they settle privately with Musk: "Yeah, pay the breakup fee and we'll forget this ever happened."
Ironically, Twitter deal plus related announcements hammered the stock more than the overall market downturn. Besides that, I believe no one wants this deal to go through including US government. Earlier, The board could not legally decline the deal
More than one person has pointed out that Musk likely wanted to sell his Twitter stock, and was using the official-looking, low-ball buy offer to pump it. Twitter board smartly called his bluff, and now Musk is trying to weasel his way out of paying anything.
Could you point me to those instances?
His offer wasn't a lowball offer at all considering the tech stock bloodbath we have seen as of late, and especially so because the offer was made just right before the massive downturn we have witnessed. I think it's quite the contrary; 44 billion is way higher than its worth imo.
> Musk is trying to weasel his way out of paying anything.
He would need to pay a $1 billion break up fee. I don't understand what point you are trying to make here.
The real bloodbath came after he bought it, at the time it was a modest premium.
> He would need to pay a $1 billion break up fee.
He is laying out the case at the moment to walk away and pay no break up fee.
Here's one [1].
[1]: https://techcrunch.com/2018/08/08/the-sec-wants-tesla-to-exp...
You asked for an example, and received one. No True Scotsman'ing the violation doesn't change the fact that Musk pumped the stock on Twitter with false and misleading statements.
Another example includes crypto: https://www.bloomberg.com/news/newsletters/2021-05-13/money-...
The example you linked does not showcase nor prove your original point.
At this point I'm beginning to wonder what your burden of proof is for Musk doing anything that could be considered wrong.
Musk only needs to pay the $1B breakup fee if he can't procure enough financing, among other strings. Musk has definitely has enough financing, he just doesn't want to pay it, so he's trying to weasel out of it by laying the ground work for a settlement to walk away.
Either way, Twitter's board is getting Musk's money, either selling at a now-premium or settling for wanting to break, due to Musk's errors.
No he does not. He has no history of that what so ever.
Seeing how badly twitter has been managed, (for a laugh, check out their "R&D" expenditure), how mush of a loss making enterprise it has been, and how it always at risk of a take over, is it that surprising?
If it has been Elliot Management (the previous rumored takeover threat for twitter) a group far less prone to public display than Musk... would things have been any less different? The only difference is that Musk is being open about what he has been doing, which I see a public good, frankly.
Elliot's track record shows it is far more vicious in layoffs of cuts.
----
https://fortune.com/2013/10/25/why-is-twitter-spending-so-mu...
https://www.rndtoday.co.uk/latest-news/is-twitters-rd-provid...
https://www.axios.com/2021/11/30/jack-dorsey-twitter-departu...
https://www.forbes.com/sites/kevindowd/2022/02/27/wall-stree...
Riiiight, and Craig Wright claims to have proof of being Satoshi Nakamoto but won't show anybody.
I don't know how good of a CEO Parag is, but he's not a very good bullshitter.
> The hard challenge is that many accounts which look fake superficially – are actually real people. And some of the spam accounts which are actually the most dangerous – and cause the most harm to our users – can look totally legitimate on the surface.
He's not talking about detecting bots. I.e. fake and automated accounts. He's talking about twitter users/bots that cause what they perceive to be harmful content. Which is a very different thing, and was the whole point of Musk's intended involvement in the first place.
Outside of that, bots cause the same problems that real people do, making twitter a place people don't want to spend time on/view ads on.
The advertisers need both bad groups removed
>* Parag says you cannot get a reliable indication of bot/not bot without this internal private data
"I have secret information so trust me" is an excellent reason to reject an assertion every time, whether it is made by an individual, a corporation, or a government. It doesn't mean it isn't true, but it means that absolutely nobody should put any credence in the assertion at all.
My markup. If I understand correctly, not having a public tweet is a marker for being a spam account. Isn't that kind of a lot of people? I know from other forums that there's a large ratio between lurkers and active posters.
Given that you need an account to customise your timeline, and, these days, pretty much for just reading a tweet, there may be loads of real reader accounts that never post and never bother customizing their profile.
I'd almost certainly be marked as a bot.
A spam account is gonna spam right? But some real users of twitter may only tweet once a year. This study just doesn't include you. It isn't saying you are spam, just not including you in the count.
Anecdotally, I know plenty of non-robot people who have twitter accounts but don't tweet.
1) They talk about "active" accounts (meaning have tweeted in the last 9 weeks), and do a bunch of filtering against that. That seems like a huge bias - lurkers exist, and in my experience are usually the majority of users...this step removes them or ignores them entirely. Frankly, until recently, my twitter account would have been one of the ones they would have discarded as inactive. This one thing alone makes me question all of the rest of their results.
2) By the same token, the rate or frequency with which a user sends tweets has no relation to whether a user is monetizable. If they're seeing ads, they're monetizable...lurkers are just as monetizable as high-volume posters.
Edit: I guess it's true that lurkers won't be bots, unless they are clicking on ads or trying to simulate engagement to help certain twitter accounts seem popular.
As long as lurkers are "liking" content, their local network will see an engagement increase.
There are other ways to achieve the goal, such as making ads more relevant (targeted advertising), having users consume more of the same content (recommendation), having the same content take longer to consume (periscope). Growing the number of human posters is definitely not a requirement.
I've been so surprised at how effective the advertising has been on me. I've never experienced this level of engagement with online marketing. Ads for TV shows, movies, live shows, musicians and comedians have been particularly effective.
I've found myself following a lot of show writers I've never heard of, and I even signed up for some new streaming services because of it. Google and Facebook ads never felt like they impacted me, though I know how important and dominant they are to business marketers. I've never clicked on a banner ad and my eyes glaze over sponsored links. Twitter's level of engagement with their marketing content is new to me, and I'm impressed.
Those are indeed incredibly valuable. Engaged audience = your real audience.
Where I might agree with you is a lurk mode account could become collateral damage in being considered fake. Lurkers don't retweet though. An account with a million followers isn't seen by everyone. Having a portion of that million like/retweet amplifies even further with their network now possibly seeing something from someone they are not following directly.
I'd be willing to accept that the number of lurkers that get lumped in with fake accounts when deciding the percentage of actual eyeballs on posts is not harmful. Those numbers are made up stats anyways. Like the old days of TV/Radio stations that covered large cities with millions of citizens. They would claim they have an audience in the millions even though a small fraction were actually watching/listening.
A lurker isn't an active user in my opinion. Maybe that's not the same understanding as accepted definition. The lurkers might be absorbing some of the ad content, but they are not helping create new avenues for ads to be shared. Twitter's ad share surface area would increase tremendously if every user was actively producing tweets. That's the only metric that they are concerned. They don't care about how many people actually see the ads once they are there. They make their money on the potenial eyeballs alone. Lurkers are not helping increase those numbers.
I don’t follow this.. Lurkers are they eyeballs presumably.
If everyone on twitter tweeted the same amount it would probably just drown out the popular accounts and create a more diffuse and less profitable ad space I think.
The number of eyeballs allows for the price per ad to increase while the number of places ads can be placed increase the volume of ads. If lurkers are not helping to increase the volume, it doesn't make the platform as much money. Proving the lurkers are actually consuming the ads and making the ad buyer happy is non-trivial. Proving the lurkers are worth increasing the price per ad is also non-trivial. In the end, I personally feel like it is a wash by lurkers being overly represented in the fake account numbers.
As an extreme example, a single monetized tweet with a billion viewers generates money. A billion monetized tweet with one viewer obviously does not..
I am technically "logged into" twitter so I can click through and read the postage stamp-sized charts linked to through various articles and blogs, or watch a video about a riot in some far flung part of the planet. Once a year I tweet at airlines when they lose my luggage or whatever but otherwise don't tweet. Twitter isn't a good social media service, it just happens to be the image/video sharing platform of choice for journalists to promote themselves.
20% is huge and I am curios if there will ever be some comparable "official" numbers to that.
It's possible they are posting MUCH more or less than 20% of the content.
If these are skewed toward the high end of producers - the 80/20 rule would say that as much as 80% of the content could come from them. Still - it's possible this content isn't interacted with much outside of other bots. You can't draw many conclusions from such a limited data point.
I created an account 5 years ago, followed one or two people, got bored and never logged in again.
Presumably their intention is to exclude abandoned accounts, like mine - is there any way they, viewing Twitter externally, could tell lurker accounts like yours and abandoned accounts like mine apart?
That's part of why I find articles like this frustrating: I don't think they have the data to actually answer they question they're attempting to answer. Knowing that, what's the purpose of the article?
It's impossible to disprove Twitter's assertion because they never claimed that less than 5% of their accounts are spam. From their quarterly earnings:
>We define monetizable daily active usage or users (mDAU) as Twitter users who logged in or were otherwise authenticated and accessed Twitter on any given day through Twitter.com or Twitter applications that are able to show ads.
>... mDAU does not include users accessing Twitter through third-party applications.
Their statement said that less than 5% of their monetizeable daily active users are spam. There very well could be 50% of the entire user base as bots or spam, but that doesn't negate the metric Twitter releases.
No idea if I should be counted or not in any particular bucket, or how anyone would know.
I have an account that is logged in, but it has only sent 7 tweets since 2014 (and they're only to customer service accounts).
No. Which is why the only reasonable thing to say as an external party is "we don't know."
Their definition:
> “Spam or Fake Twitter accounts are those that do not regularly have a human being personally composing the content of their tweets, consuming the activity on their timeline, or engaging in the Twitter ecosystem.”
They note the following to differentiate fake and spam: > Many “fake” accounts under this definition are neither nefarious nor problematic. ... By contrast, most “spam” accounts are an unwanted nuisance.
Some general data analytics notes from their post:
* Then lump together fake and spam in their analysis - and this really matters! somewhere like NYT is both 'fake' meaning it isn't a real person and A HIGHLY VALUABLE ACCOUNT for twitter to have.
* They use a sample of 44,058 accounts (of ~1.047B)
* They look at a number of classifying variables (17), spam accounts met 10+ of those 17 criteria. They don't list all 17.
* The criteria were developed from a "machine learning process" that is undescribed, and was developed from a sample of 35,000 'known' fake twitter followers bought from 3 vendors and 50,000 claimed non-spam accounts. They appear (imply?) to have used 50% training 50% real data but dont't specify explicitly.
* They say their model is about 65% accurate, and unlikely to produce false positives ("almost never includes false positives") - however they don't list any specificity, sensitivity, etc. that would be useful to evaluating that claim.
* The analysis does no statistical tests, no confidence intervals, minimal information about how the model was tested or validated.
* Critically: they note, but do not describe or quantify, that a lot of the criteria are highly correlated
* then later in the article they suddenly seem to switch to a 10 point scale for quality away from their 17 point scale? with a threshold of 3 or below as low quality?
* My personal twitter account meets most of the metrics where they have listed a quantifiable threshold. And their fake followers tool lists it as pretty f'ing suspicious - i.e., low quality.
I'm not saying there wrong but I am saying good luck getting this from a blog post to any sort of respectable science publication. As they note at the end, they aren't even calculating the same metric - twitter uses monetizable daily active users - remember NYtimes? Absolutely a monetizable account - even if it isn't a real person.
anyone who thinks this is proof of Elon's 4D chess based on this article is, to me, frankly delusional.
The fact that they add a .42% is a red flag in itself, especially when they admit in their own post that they agree that their analysis is deserving of critique. Very misleading stuff.
Their analysis using purchased bots seems a bit more reasonable.
Similarly I don’t think there is any way to separate active vs abandoned passive accounts as a 3rd party.
Even paid for Twitter Blue...still thinks I'm not real. Support is unreachable.
My current plan is to wait til Elon completes the takeover and then build an entire site dedicated to getting Elon's attention to unlock my account...because that's the only way to contact somebody apparently.
Edit add: I find it horrible that we have companies that you can not contact, in fact they seem to be going out of their way to make hard to contact them.
Even things you pay money for, like airline tickets. They want you to email them, make the phone number hard to find. So you do, they don't respond and then you have to search and call them, wait an hour or more on hold. The agents are nice but the entire process is terrible.
Earlier I had to do that for a damaged luggage claim. Went through the automated phone assistant to get to damaged luggage claims and it gave the option to use text messages. So I give it a try, nope. They can't resolve the issue through text, has to be on the phone. So I had to call back, re-enter all the info through the automated system and then ignore it's pleadings to use the text system.
I believe the author writing prompt was just: a headline about fake Twitter accounts showing a number significantly higher that 5%. That’s it. Whatever the methodology, that was the author’s goal.
The article achieved this goal. Otherwise is completely irrelevant. Even for the person who wrote it.
It's almost like Inception isn't it? A PR stunt within a PR stunt within a PR stunt.
Sure that's a different question from what proportion of all users are fake/spam, but this is still a perfectly valid question to ask, and the fact that they're only considering active users is in the title so I really don't get your complaint.
If you want an analysis that attempts to answer a different question go find or write one that addresses the question you want answered...
The article clearly states (emphasis mine):
> This represents the largest set of accounts on Twitter we could acquire, but it includes analysis of many older accounts that haven’t sent tweets in the last 90 days and thus, likely don’t fit Twitter’s definition of mDAUs (monetizable Daily Active Users).
From the linked Twitter earnings report:
> We define monetizable daily active usage or users (mDAU) as Twitter users who logged in or were otherwise authenticated and accessed Twitter on any given day through Twitter.com or Twitter applications that are able to show ads.
EDIT: rephrased "accounts that are active" to "accounts that actively send tweets" to clarify what the article addresses.
You can file this in the 'pedo guy' cabinet of his life story where his child-ego got the better of his undoubted business skills.
Given what Musk does to the personal lives of his opponents, I'm not sure I would want to fight him. But given how many laws and rules he's broken at the point, I think there is a clear failure of justice if he can just do whatever he feels like without repercussions due to his common popularity.
You could probably argue that most of the world read Twitter and hence are users, account or none. It's that pervasive.
But then there's the next question: "am I a user that reportedly matters to Twitter's business?". What people are trying to land on, in light of Elon's tweet that the deal is on hold pending investigation of Twitter's metrics reporting, seems to be a framework for carving out what exactly constitutes a user that brings the platform revenue that shows up in quarterly reports and hence would directly relate to the tangible value of the enterprise.
In reality, nobody knows what numbers are being thrown around behind closed doors. This article is just one framing.
But there's no reason to pin the whole frame of the conversation to the one question for which Twitter corporate chose to publish an answer, unless the only question we are interested in is "did Twitter technically lie" which is the most uninteresting question in this whole situation. If this is the sole context you are using to frame this issue then maybe you should consider if you're following the current news cycle a little too closely.
The idea that there is such a thing as an 'inaccurate definition of active' is silly.
I don't know, that seems more interesting than most questions that could be asked about Twitter.
Why? Twitter is a for profit corporation. If, on the balance, lying serves their interests (I'm sorry, I meant "is consistent with their fiduciary duty to their shareholders") more than edging up to the line without crossing it, that's what they will do.
Even the watchdog organizations such as the FTC and SEC that police the speech of corporations more or less limit themselves to material statements that move markets or influence consumer behavior in ways that can be considered fraudulent. The FTC, FDA, and others are concerned with a fairly narrow reading of consumer harm, the SEC is motivated by the health and trustworthiness of the public market. In any case, there pretty much always has to be some sort of alleged harm. Lying per-se is hardly ever forbidden. So if the advantages of a lie outweigh the (risk adjusted) penalties and reputational risks, that's that.
I agree, that would also be a much more interesting conversation than "did Twitter technically lie."
"Active, as in logs in regularly? Wait, what is 'regularly'? Once a week? Once a month? Every day? Does 'active user' mean, online right now?"
Etc, etc...
Their definition of "active user" is relative, not inaccurate.
The big story in the news last week was "Elon Musk says deal on hold while verifying twitter's 5% Monthly Active Users stat", or something to that effect.
That's the context this article was published in. It is transparently obvious they are re-using the word "active twitter accounts" to cause confusion with the definition of "active" that has been being bandied around. The post is using such a title as a clickbait, to hop aboard a trend.
I think the title, and lack of significant clarification in the article, make it clearly misleading, and I don't think pedantic "well technically active can have multiple definitions" changes the reality of the situation meaningfully.
Let's take both their numbers at face-value and assume they're true.
Twitter has reported: 396.5 million logged-in-this-month users, of which 5% are fake/spam (19.8 million fake users)
This article reported: Looked at 44,058 tweeted-recently accounts, of which 20% are fake (8,800 fake)
Which of those stats looks worse for them?
> The high number of lurkers would make the percentage of fake accounts smaller
Why? Twitter included lurkers in its dataset, this article didn't, why should that impact stats in the direction of fake accounts being smaller?
Because you usually don't create fake accounts to lurk, but to do "something".
I'm speculating, but even when you create bots to boost follower counts you'd probably make them post now and then so as to seem "active".
It makes sense that the proportion of tweeting accounts being bots is much higher than the proportion of lurkers. And since there are also more lurkers in turn than posters, I would say that the real number is much lower than that.
Let's say I own a twitter bot farm. I make 20k accounts, have a system setup that logs into each of them from a unique IP each month at random times to make sure they're not banned yet, and advertise it out. On month 1, someone buys 1000 of them as followers. On month 2, someone buys 1000 of them to tweet spam. etc etc.
Each month, there's 20k active bot accounts (logged in to verify they weren't banned). Only a small number may actually tweet though since buyers may have not gotten them yet. Bot accounts lurk too, for months on end, before ever acting.
I'm not claiming this is accurate, but I am claiming this is a reasonable alternative which doesn't align with the view of bot accounts being more prevalent in tweeting accounts than lurking accounts.
A metric they've artificially inflated by gating tweets, which works to their advantage when calculating spam. With that in mind, I think I'm more inclined to look at spam as a percentage of active tweeters and ignore lurkers.
If you're more interested in Twitter's ecosystem as a whole, it is less interesting.
Nope. ~20% of accounts that tweets are fake. A lurker (aka read-only) is by all meanings an active account.
And since they are trying most probably to get some PR for their company, they use their specific definition of "active Twitter account".
Sure, clickbait headlines are the norm and the devil lurks in the details, but still, many comments have been spent on this, because it's clearly misleading.
~80% of email is spam, it doesn't surprise anyone, because it's so cheap to send spam. Similarly it's easy to create fake accounts and spam, yet it doesn't mean much.
> it includes analysis of many older accounts that haven’t sent tweets in the last 90 days and thus, likely don’t fit
> Twitter’s definition of mDAUs (monetizable Daily Active Users)
As implying that they think accounts that haven't tweeted in the past 90 days don't fit Twitter's mDAU definition. Given the placement of the qualifying phrase, I think that's a reasonable parsing of the sentence, but I see your point that they could be trying to imply their set doesn't fit the definition. If so, that sentence is very badly constructed.
> Followerwonk selected a random sample from only those accounts that had public tweets published to their profile in the last 90 days, a clear indication of “activity.” Further, Followerwonk regularly updates its profile database (every 30 days) to remove any protected or deleted accounts. We believe this sample is both large enough in size to be statistically significant, and curated to most closely resemble what Twitter might consider a monetizable Daily Active User (mDAU).
The fact that they don't even consider the concept of a non-tweeting lurker to be an mDAU brings their entire analysis into question. Let's face it - Twitter is an emotionally-charged enough place, and tweets have such a way of living forever and being taken out of context, that there are many who use it to consume (and perhaps Like) content but will not tweet publicly. These people are still viewing and engaging with advertisements! Twitter absolutely should consider them monetizable!
But of course, engagement data on lurkers is internal only, and Likes data counts against global API caps: https://developer.twitter.com/en/docs/twitter-api/tweets/lik.... Which means that SparkToro and Followerwonk are incentivized to ignore these users. That they do ignore them, and don't address it anywhere in their methodology, is highly suspect.
But then use your own definition of active and write only a one liner on the difference with no reflection on the impact it might have and no warning on the fact you are answering a different question. Then my conclusion is you want people to make this mistake.
> EDIT: rephrased "accounts that are active" to "accounts that actively send tweets" to clarify what the article addresses.
Made me laugh because you had to add it and made more effort than the author of the article to prevent the confusion :D.
The fact that you had to do this proves the point. Nobody defines "active" the way they have here. The claim is nonsense.
Though personally I think filtering specifically for users that actively send tweets makes sense, since that's really what matters when it comes to measuring how healthy and authentic the discourse is
It seems like everyone is arguing about different metrics and it makes more sense to discuss different, specific measures that might fall into a range of behaviors that are "active" in some sense rather than focusing on which definition of "active" is somehow the best one.
What would be more interesting would be to adapt this and answer several different questions about the proportion of spam among accounts with different metrics of activity to see how things change. For example, does the percentage of spam accounts go down a lot if we lower the bar for "active"? How much & how fast?
Twitter's quarterly earnings define active users thusly:
> Twitter defines monetizable daily active usage or users (mDAU) as people, organizations, or other accounts who logged in or were otherwise authenticated and accessed Twitter on any given day through twitter.com, Twitter applications that are able to show ads, or paid Twitter products, including subscriptions.
https://s22.q4cdn.com/826641620/files/doc_financials/2022/q1...
I'm pretty sure I've heard a similar definition from Facebook.
This definition supports g-clef's critique that the article picks an unorthodox way to measure active users, resulting in an inflated percentage of accounts being measured as spam/fake accounts, vs what the percentage would be if measured against Twitter's definition of 'active', which includes lurkers.
> “Spam or Fake Twitter accounts are those that do not regularly have a human being personally composing the content of their tweets, consuming the activity on their timeline, or engaging in the Twitter ecosystem.”
Ok, but "consuming the activity on their timeline" is essentially unknowable outside of Twitter, since you can't see what tweets people are viewing. It turns out they're trying to infer this through some other signals like follower count, etc. But you can imagine why that might be sketchy.
Then they constrain the analysis: > A more fair assessment of Mr. Musk’s Twitter following would only include accounts that have tweeted in the past 90 days
Let's be real, if you look at a list of Elon tweet replies, they might as well all be spam. Just search @elonmusk and sort by latest. Then compare that to the sorted tweet replies under an actual tweet. IDK how many millions of dollars and man-hours went into the AI that sorted this list, but it seems to just be putting the blue checks at the top and shrugging at the rest. I doubt this three man team is doing any better at spam detection.
I've spent many, many hours lurking on twitter, don't have an account at all, and mostly access it through nitter instances. Are they "biased" for not including me?
edit: should inactive users be counted as active users?
I do wonder about, given perfect knowledge, how the bot accounts would shake up. What percentage produce content (presumably propaganda, automatic tweets using it as an RSS like announcement service, and spam) vs follow people (boost follow accounts, sell likes)?
But they probably should expand more on this and reflect on how much inaccuracy it adds. With a quick search you can find that less 50% of US users tweet five times a month (https://www.pewresearch.org/fact-tank/2022/03/16/5-facts-abo...). Or the study which, reported that the top 25% of user produce 97% of the content, the median user of the bottom 75% as posting 0 tweet a month (https://www.pewresearch.org/internet/2021/11/15/2-comparing-...). Those studies were done using survey I believe so should include only active users and no spam/bot.
So with random invalid maths, if you make the assumption that the 25% less active users might not even post every two month (exponential decrease of activity ?) then you need to add back a quarter of the 80% they found as active.
Not to say I believe the 5% number from twitter; and I was going to use the price for a thousands follower as an example, but seeing it appears to be at 30$ now (https://socialboss.org/buy-twitter-followers/ ?) when I remembered it at like 5$ then the twitter team might have done some good work ;).
Probably picking the sample is still challenging but at least can somewhat tell if the accounts in the sample are genuine.
The method in this article is so flawed that Larry Ellison, founder of a famous law firm, would count as an inactive account since haven't tweeted since 2012[0] and that person apparently looks into investing in Twitter[1]. How can be investing a billion in Twitter when he doesn't use Twitter at all?
[0]https://twitter.com/larryellison?lang=en
[1]https://www.grid.news/story/politics/2022/05/16/larry-elliso...
This is not their definition, that's what Twitter considers an active account in their revenue reports.
> has no relation
It has some relation, no? I wouldn't be surprised if there is a strong correlation between how frequently a user sends tweets and how monetizable that user is.
All true. However, do you really believe that a bot is more likely to be active than a real user? If so, fair play to you. If not, then we would expect inactive users to be bots in an even greater proportion than what we see among active users.
I don't see how you can arrive at this conclusion. It depends on who you are following, with some additions by the algorithm (unless you use the chronological feed) and (speculating here) the algo pushes content from real humans.
Except - the baseline that they choose is entirely NOT comparable to that of Twitter's baseline. The study says:
> Followerwonk selected a random sample from only those accounts that had public tweets published to their profile in the last 90 days, a clear indication of “activity.” Further, Followerwonk regularly updates its profile database (every 30 days) to remove any protected or deleted accounts. We believe this sample is both large enough in size to be statistically significant, and curated to most closely resemble what Twitter might consider a monetizable Daily Active User (mDAU).
Except that we know what Twitter defines as a monetizable DAU:
> We define monetizable daily active usage or users (mDAU) as Twitter users who logged in and accessed Twitter on any given day through Twitter.com or Twitter applications that are able to show ads.
Nothing about posting, nothing about engagement at all - simply: were you able to see an ad?
So there isn't any reason to claim that this "might" represent what Twitter uses as an mDAU - we know, in fact, that is not how they measure it. A more honest statement would have been:
"We selected a random sample of (etc. etc.). We believe that this sample is large enough to be significantly significant, however, it can not be compared to Twitter's mDAU set, as it does not count passive consumers of Twitter content. Instead, this data can be used to suggest that a significant amount of the total posted content on Twitter is delivered by bots"
My guess is the number of consumers of content is greater than the posters of content by several orders of magnitude, though some of that would be mitigated by the longer time horizon.
Both sides are bullshitting to negotiate a better price.
https://d18rn0p25nwr6d.cloudfront.net/CIK-0001418091/cb1d93d...
"We have performed an internal review of a sample of accounts and estimate that the average of false or spam accounts during the third quarter of 2020 represented fewer than 5% of our mDAU during the quarter. The false or spam accounts for a period represents the average of false or spam accounts in the samples during each monthly analysis period during the quarter"
This is not a new stat or new information.
Are you saying Twitter isn't bullshitting? You think that Twitter put in their best effort to get an accurate number?
"both sides are bullshitting to get a better price"
Which implies that this number is related to Twitter trying to get a better price. However, if this is the number they have always said, regardless of the methodology, clearly they can't be trying to get a better price - as they started doing this long before there was a price to make better.
That is separate and independent from whether the number is accurate to begin with - if it's remained consistent and publicly available, there's no chicanery here, the risk around that stat has always been baked into the market price of the company.
"Are you challenging Twitter’s earnings report, saying that <5% of mDAUs are fake/spam?
We are not disputing Twitter’s claim. There’s no way to know what criteria Twitter uses to identify a “monetizable daily active user” (mDAU) nor how they classify “fake/spam” accounts. We believe our methodology (detailed above) to be the best system available to public researchers. But, internally, Twitter likely has unknowable processes that we cannot replicate with only their public data."This guy (Rand Fishkin) has been selling SEO as a religion for the better part of this century, and is in no small part responsible for all the search-result-garbage style websites everyone is complaining about elsewhere on HN today and every other day.
He's a third-rate market-bro hack that's been taking advantage of web professionals who get thrown into SEO/Marketing jobs and have no idea what they're doing by relentlessly shoving half-assed corporate strategies through moz.com and now his new sparktoro.com, and calling himself the great SEO redeemer.
Wanna question his methodology? There is none. Wanna question his science? Totally devoid.
https://blog.plan99.net/fake-science-part-ii-bots-that-are-n...
> This methodology likely undercounts spam and fake accounts, but almost never includes false positives (i.e. claiming an account is fake when it isn’t).
In other words, their model performs well on their training set, and they don’t acknowledge that it may be over fitted or mislabeled, and they hand wave mistakes
meanwhile, saudi bot armies apparently basically run rampant across the platform
Less than 2 days later I am banned for "spewing hatred based on sex, gender or religion". No amount of replies that my tweet was same as another but much much more popular account helped. Banned for life.
Somewhat alarming, if accurate.
I'm active on twitter, check it every other day, follow @ElonMusk but I don't tweet. Perhaps I'm a unique case, or their assumptions are a bit off.
Some metrics they consider suspicious (from the article):
1) Accounts that didn't tweet recently
2) Accounts with low number of tweets
3) Accounts with a low number of followers
4) Accounts that didn't set up their own profile image
Lurking != bot, and these data-points would all hit high for lurkers. I'm somewhat suspicious of their results, especially given the results from this pew study suggesting the majority of twitter users don't tweet very much.
25% of Twitter Users Produce 97% of All Tweets: https://www.pewresearch.org/internet/2021/11/15/2-comparing-...
I treat it like RSS not Social Media
If it's 1-1 then 20% of tweets being from bots isn't great - but it could even be more than that.
I think Youtube's UI "encourages" this behaviour even more as it has a highly algorithmic homepage rather than a feed of people you follow.
I've been meaning to write a blog post regarding this endeavor. But moreso, I've come to the conclusion that social media needs to have verifiable audits for their userbase; similar to how there are audits done for financials. A lot of the value of these companies IS derived from their DAU/MAU and or userbase in general (example: WhatsApp - $19 billion for their 1 billion users).
Twitter is different, they seem to have a more robust banning system in place. Instagram, not so much.
You expect them to ban you based on your account gaining fake followers? Wouldn't such a policy make it very easy for any third party to get a user banned?
Of course an active Twitter account can be one that never tweeted in 10 years. Such an account may as well see advertisements etc... So I think any outside studies are fundamentally flawed since they don't have access to internal data like last login time.
The chance of no digit repeating out of 7 would be about 20%
> Our systems do not, however, attempt to identify Twitter accounts that may be irregularly operated by a human but have some automated behaviors (e.g. a company account with multiple users, like our own @SparkToro, or a community account run by a single person, like Aleyda Solis’ @CrawlingMondays). We cannot know how Twitter (or Mr. Musk) might choose to classify these accounts, but we bias to a relatively conservative interpretation of “Spam/Fake.”
So this means:
* @EmojiAquarium - spam/fake * @threateningcake - not spam/fake * @CanYouPetTheDog - spam/fake * @ChuckGrassley - not spam/fake (?? - what fraction are staff generated vs Chuck?) * @Wendys - spam/fake * @Twitter - spam/fake
To quote Matt Levine's "Money Stuff" newsletter:
> “Temporarily on hold” is not a thing. Elon Musk has signed a binding contract requiring him to buy Twitter.
> That contract does not allow Musk to walk away if it turns out that “spam/fake accounts” represent more than 5% of Twitter users... The merger agreement contains a provision that allows Musk to walk away if Twitter’s securities filings are wrong ... but only if the inaccuracy would have a “Material Adverse Effect” on the company. That is an incredibly high standard: Delaware courts have almost never found an MAE.
> Musk ... had the opportunity to do due diligence on these numbers before signing the deal. (He declined.) He can’t now go to Twitter and say “actually now you need to prove that your user numbers are right.”
[0]https://www.bloomberg.com/opinion/articles/2022-05-13/elon-m...
How many more of these types of "events" until people stop treating Elon like some super genius god? Stop giving/loaning him money enabling this garbage.
Perhaps all this uncertainty is intentional.
It appears hes angling for a lower price with the threat of a court battle. He doesnt have much basis but it would take time, money, and probably wouldnt be great for Twitter.
1. Select 1000 accounts uniformly at random. Either from among all twitter accounts, or from active twitter accounts for whatever definition of "active".
2. Classify these 1000 by hand. Do as much investigation into them as you need to classify them accurately; no need to use heuristics here.
You will (with very high probability) get an estimate accurate to within a percent or so. If you do statistics you could find the actual bounds.
Users are denoted by numerical ID, you can sample using this.
Anyway I think it’s pretty goofy to try to make claims around %s of Twitter accounts, active or otherwise, only absolute numbers make sense. What really matters at this point besides revenue and revenue trend?
"Through trial and error (and, of course, pattern-fitting) we crafted a scoring system that could correctly identify over 65% of the spam accounts."
65% is not actually very accurate for a binary classifier...
"Applying this model to the ~44K random, recently-active accounts provided Followerwonk produces a quality score for each account, visualized below:"
Many real twitter uses are likely not to be "active" aside from reading stuff. So this methodology would clearly overestimate the number of spam/fake accounts (which all would be active).
Also, this is an important point:
"The other potential critique is our spam/fake follower calculation methodology. Because we crafted it in 2018, based off sample sets of purchased spam accounts, it’s likely that more sophisticated spammers and fake accounts go unidentified by our system"
The features collected are certainly outdated by now.
A bot operated to post advert messages might want to post them on high profile accounts.
A bot for a "500 followers for $5" company might be more likely to follow low-profile accounts, with some random activity/follows thrown in as camouflage.
A bot operated to amplify certain messages/opinions might follow accounts of middling visibility, and focus on likes rather than more visible activity.
I hope they like law suits:
> Our analysis found that 19.42%, nearly four times Twitter’s Q4 2021 estimate, fit a conservative definition of fake or spam accounts
Ok great, I hope you have fun proving that in court. Especially the part where you have to prove that Twitter's definitions which you don't know match yours.
>SparkToro is a tiny team of just three
Then I applaud your bold decision to interfere with a $45Bn merger.
>Our definition (which may differ from Twitter’s own
Any lawyers in the house? How obviously do you have to renege on your libelous claims before you're in the clear?
Are these are the fake accounts that are lurking around without posting a tweet, but impacting the mDAU?
"Active" definition needs to be more precise than simply tweeting.
If I were an investor or an advertiser, I would drill into these details.
No Twitter, you are not getting my phone number.
Twitter, however, does call these accounts "fake" and has further rules/policies concerning Misleading & Deceptive Identities.
https://help.twitter.com/en/rules-and-policies/twitter-imper...
'Fake' accounts exist to fraudulently increase follower count. This is what the study is claiming to measure. These typically have low activity and engagement profiles.
'Spam' or 'bot' accounts, on the other hand, generally have high activity. Whether trying to influence political opinions, or engaging in astroturfing or phishing activities. They probably have very, very high ratios of replies to original tweets, and overall tweets to # followers.
Perhaps Mr. Musks' team did an analysis of their own and saw a high number of potential fakes/bots, and that made them question the authenticity of Twitter's numbers.
If it can be shown with some accuracy that Twitter underreported the number of fake/spam accounts, how does this effect Musk's acquisition? Could he lower the price by saying you gave me incorrect data?
What about "bots" that aren't "bad bots"? For example, a news organization twitter account that automatically tweets out new articles. This is clearly a "bot", but is it a "fake account"?
Yes, most studies published specify what their own interpretation of these are, those definitions tend to be wildly inconsistent, and make any sort of comparison of the rest of the methodology impossible.
The delay in closing the deal is just that, a delay, in the hope that the market will recover and the price of Tesla stock will recover, thereby making Musk's financing position much easier.
Nobody anticipated the downturn, otherwise Musk never would have made a $54 per share offer for Twitter in the first place, because obviously it's much lower now, as is much of the tech market.
It's hard to believe that making a binding $54 offer on a stock that's only $39 a month later was "4D chess". I do think $54 was a reasonable, maybe even lowball offer a month ago. Twitter itself hasn't changed much in a month operationally speaking, but the whole market changed in valuation.
It doesn't matter if 5% or 20% or even 40% of Twitter users are bots. If Musk manages to double the number of subscribers by making Twitter a household name then the purchase was a good investment.
One definition isn’t better than the other but it is worth noting that opening Twitter, looking at a lot of Tweets, seeing a lot of ads, and not posting any Tweets is a behavioral pattern much more commonly seen in humans than in spambots.
Which features were they? How reliable are those features? What method was used to overlap the features? What thresholds within that method were used to classify the overlap(s) as spam, or not spam?
Unfortunately this research, while compelling, fails the falsifiability test.
The reason why it seems not as effective (at lower price brakets) is because the entire industry is based on fake numbers. This fake degree of effectivity allow them to balloon the prices,
1st By alleging that the ad is "as effective as what you are willing to pay" 2nd. Their "massive" userbases.
In reality lower priced ads are not effective precisely because their userbase is not as massive.
as such what should be defined as normal tier is now the highest priced type of advertisement campaign.
The people hardly dig into this as there is no real incentive, billion dollar companies are not being sold all the time.
The % of bots was known to be between 15 and 20 before he made the deal.
If your plan is to buy a soda company to sell soda I would not believe a good strat is to reveal how cheap the product is, in reality, you are basically cornering yourself.
If he manages to buy it,the issue will immediately dissapear.
If he would genuinly made an effort to unmask the industry to help the consumer, I would be behind that.
These are active accounts.
But active accounts usually are a small percentage of total accounts, especially when it comes to bot networks, because bot spammers usually create huge quantities of accounts in advance, so that as soon as one bot goes down, another spins up to take its place.
I'd be surprised if 80% of the total accounts on twitter aren't bots.
It is entirely feasible that at any point in time 20% of accounts are bots. It is whether they have counted them in their mDAU that matters.
On the flip-side, you do realise this is a way to justify complete verification of users on the platform to significantly reduce the bots, fake accounts, etc? I won't be surprised to see that implemented soon.
Probably by the end of this month, things are going to get very chaotic at Twitter.
We'll see what happens.
If you took a random sample of websites you'd find that 99.995% of them are spam. That's why we have to rank search results.
https://mobile.twitter.com/McDonalds/status/1521143128073424...
it s very hard to say anything about bots without IP addresses. The most active bots get sold / change subject etc.
I was just explaining the law of small numbers to my daughter for her science fair. I think this would qualify.
At some point bots will have enough humanity and intelligence breathed into them to deserve some dignity. Like books, books deserve respect, you can't just burn them or step on the pages, Bibles and Torahs in particular. That was a huge thing in the Holocaust, protecting many Torah from destruction, in one case I read Jews protected them by burying them.
I would pay $44bn to own email.
Bring back the bots, they were great
Which is home field advantage for Musk really, considering what he got away with in the past.
Kind of "hey guys, look, I really think your platform costs around 52$ per share, but ... let's try to see the actual price. How many bots do you have again? :)" Whenever this guy speaks, a whole freaking market moves.
It's like "cleaning up the mess" even before joining the company as their CEO. In the worst case scenario, he dodged a huge bullet. So, who cares.
My perception of Twitter is that some of the worst behavior on Twitter is coming from the 'well-respected authentic accounts of government and media personalities', particularly when it comes to disinformation, marketing, and indeed, intimidation of those with non-conformant opinions on various topics, such as, for example, the wisdom of unconditionally flooding Ukraine with high-tech ordnance in the name of repelling the Russian invasion (or is that really more about delivering a cash cow to US weapons manufacturers)? That's the kind of thing that could lead to doxing by corporate media reporters and charges of being 'a Russian asset' even if such concerns are reasonable.
Also, how many of the 'registered authentic accounts' include tweets not written by the actual personality they claim to represent, but rather by some kind of in-house PR/image management consultants? Just because there's a paid human somewhere in the chain doesn't mean the account is all that 'authentic'.
The valuation the board has reached, for Twitter shares, is predicated upon human users on their system and not spam accounts.
Twitter's valuation of its stock is way-overblown if there are 20% bot accounts on their platform.
Go Elon!
Re:parroting, while that seems to have been a pejorative, I wasn't actually copying from anywhere, just a smb executive sensitive to this sort of thing during m&a as standard business discussions, so it's interesting that it's being repeated enough to accuse people here of 'parrot'ing!
The SEC filings thing has been repeated even by Musk himself. It doesn't really help him.
Is it way above fair market value (especially in the time between his offer and today)? Absolutely.
Absolutely! When the board demanded that Musk pay $54.20 per share – which is a completely rational number that they demanded from Musk and not a meme number that he made up himself – they underestimated his ability to do Math, and he's about to show them what's what.
One of the core duties of a Board of Directors for a publicly traded company is to arrive at a valuation of the stock each morning before the markets open. It's a shame that this board is either incompetent or corrupt and came up with the number $54.20 based on bot accounts. Musk is having his people do a complete audit of 100 random followers of the @Twitter account to find the true number. I expect the terms of this deal to change significantly in Musk's favor.
It's too bad that this distracts him from shipping the Cyber Truck, going to Mars, and then building a 1.6km tunnel under Las Vegas (a feat which has never before been achieved).
Praying for a quick resolution so we can finally get some Free Speech on the Twitter platform.
If the figures for Elon Musk and Donald Trump's followers are at all accurate, it's over 50% for large accounts. Which means that the people who probably care most about whether or not their followers are real (e.g. those who might be charged for the privilege of using Twitter to communicate with large numbers easily), are the ones most likely to have mostly bots.
I could see why this might impact the value of Twitter as a company.
[1] - https://www.bloomberg.com/opinion/articles/2022-05-13/elon-m...
PROTIP: You can search Twitter by URL. Just search for a popular news article on Twitter, and you can see the massive number of obviously non-human accounts who are tweeting the article.