College Kids Doing What Twitter Won't
wired.com
wired.com
Hillary herself is sponsoring since years a small army of geeks in full time that blast biased tweets and comments, posing themselves as normal posters instead of being identified as paid comments: http://www.motherjones.com/politics/2014/09/david-brock-hill...
This is a perversion of democracy, not even taking weight the proven federal crimes to which this former presidential candidate seems immune.
The means to bring back balance to democracy are certainly in our hands. This also includes exposing what is wrong with the upper echelons of western society.
You do understand how campaigning works, right?
Not how you want it to work, but how it actually works.
Really? I'm asking because I don't know the answer; don't you have to identify yourself (like TV ads that say "I'm Obama, and I approve this message").
If there are punishable crimes why hasn't the justice department filed charges?
It seems as though some people are starting to realize the weapon they thought was their own, was never theirs at all… inserts something about chickens coming home and roosting
All in all, every step twitter takes to "protect" its platform (euphemism for those with voices most likely to be received and implemented by them, to achieve some arbitrary censorship ends) will be a step forward for federated and decentralized social networks.
Though, who are we kidding, lets hope that in the next election, even more money is wasted by the entrenched and ossified political orders and even more for the one that looses that needs some PR to cover their ass from the carnage that comes after, after all, its just good business.
[0] http://web.archive.org/web/20140531073143/http://www.cfr.org...
To get their training data, they identified 100 accounts that appeared to be bots, and then assumed that all their followers were also bots. They used verified accounts as their sample of non-bot accounts.
They claim they "identify bots 93.5 percent of the time." No data is given about false positives.
They have demonstrated a proof of concept that hunting bots on Twitter can be achieved, and that Twitter could do it easily if it wanted
You will basically be put at risk for following 'controversial' accounts on Twitter of having your account delete accidentally. And businesses will have to jump through bureaucratic hoops to use legitimate automated scripts to manage accounts.
It also seriously deters anonymous online speech if they force a big percentage of users to provide a unique phone number to prove they aren't a bot (the current Twitter strategy for countering spam). Not to mention the liberal sharing of information with nation states, surreptitiously or otherwise, even for people not trying to be anonymous. A phone number is a critical identifier in the surveillance industry.
Getting this accurate is very important and I feel like that is being heavily downplayed in this article.
No false positives ship has already sailed. What you can best hope now is timely response to mistakes (and as few of them as possible).
Captchas are certainly not foolproof, but they up the difficulty of running a huge number of phony accounts while having a minimal impact on normal Twitter usage.
They started by identifying "left- or right-leaning" tweets. Then they decided that everything in-between were bots. Who decides what is left and what is right, and who decided that anything other than those extremes is "fake news"? There are ethical questions here.
It scares me that we now find it acceptable to classify accounts as "bots" or "Russian-backed" because of any content in the tweets, even if it's #LockHerUp or #BuildTheWall. If people really believe these things, shouldn't they be allowed to say them? What if tomorrow, something that you strongly believe in becomes labelled "fake news" or "propaganda", and your own accounts are shut down?
The methodology is flawed in the most fundamental way.
There are other technologies that can be used for detecting bots[0]. Twitter should use it.
[0] https://security.googleblog.com/2014/12/are-you-robot-introd...
In our model, we look at hundreds of features per account. Polarizing tweet content is not the only thing the model looks at. If an account is tweeting OC content, it's a clear indication that account is human.
Twitter's policy allows bots of their platform, and most are harmless! There are bots that tweet out the upcoming songs on a radio station. Unfortunately, a CAPTCHA would eliminate all bots, not the ones that spread fake news or claim compromised accounts.
Give users the ability to designate their account as a self-reported bot. Accounts designated this way would be identified as a bot in the Twitter UI, and would be exempt from captchas.
There's the rub, Twitter doesn't want to do that. Why? Good question, but the WSJ adage applies: Follow the money.
I think it's because bots make up most of their users and that getting rid of them is a net negative to them. Not just because they can make stockholders and advertisers think they have more people online than they do. But because trying to continually cull the bots would cost too much as the bot makers up the ante every time. It's not profitable to chase them and it seems that the people that pay Twitter don't care about them.
Also, by bot, I do not mean the hydration sensor in your tomato garden or the code that pings your followers to watch you on Twitch. I mean bots that are trying to masquerade as real humans to get you to buys stuff or vote some way or another. Essentially, things making an attempt at the Turing test.
How long until their blind eye towards harassment, racism, abuse, and booting collapses any value they have? How long until the average person perceives Twitter as a place where racists congregate and no one else can bring themselves to visit?
And when it does, will they pull a reddit and suddenly grow a conscience, miraculously saying "Oh there's a problem on our service and we're going to charge to the rescue!", or are they going to go 4chan and hide behind "freedom of speech" to excuse the filth their platform breeds?
I mean, I don't think 4chan is doing that at all. You give the sons-of-moot too much brainpower. It's a bathroom stall's wall, policing it is silly.
That said, when will Twitter crash over the bot stuff? Wall Street doesn't think it will crash, so it won't. When Wall Street thinks it will crash, then it will. However, Twitter is at the foundation of the ad-scam that most of the Big 4 are running. When that goes belly up, Twitter will be along for the ride, but it's more of a Haunted House than a Log-ride. Scary.
Again, this is my thought process, who knows.
This seems like a fatal flaw, are the subsequent layers also bots? I mean, if bots have no human followers, what is the problem?
For a long time, the play book to improve your online reputation was to get rid of those bad links. However, a company (and I can't remember the name) decided it might be easier to just create better links/stories about the person and let google bury the bad stuff. Therefore when someone searches Bill Smith, they see the good stuff.
Anyways - back to the topic at hand here. Perhaps the solution to these bad bots isn't to try and stop them, but to build better bots that spread #realnews and aren't as toxic? Maybe it wouldn't work, clickbait and catchy headlines are more "share worthy" than the latest WSJ front page article, but maybe we need a new strategy?
[0]: https://en.wikipedia.org/wiki/Journalism
[1]: https://en.wikipedia.org/wiki/Media_bias
As for non-toxic, I definitely think that's possible, but involves a bigger conversation as to how people communicate with each other: the same self-awareness that's necessary for good journalism is necessary when people talk with each other day-to-day. It's not clear to me how we rebuild our ability to engage each other in a non-toxic way about important and divisive topics, but I think it's necessary and well-worth trying to figure out.
There is a such thing as journalistic integrity, it is possessed in wildly different amounts by different news sources and it matters. A lot.
The internet has broken mainstream media’s monopoly on “the truth” so it seems a natural next step that they are viewed objectively based on their actions instead of their perceived and heavily marketed position of authority.
Postmodernism itself is the byproduct of the Information Age giving people access to multiple viewpoints. I’m glad to see the playing field leveled.
I despise propaganda as much as... Well, I despise it.
Unfortunately, the winners of 'breaking the MSM's monopoly on Truth' aren't any better then what they replaced (And are in many ways, much worse.) Their output is, by far and large, the worst kind of yellow journalism.
Unsurprisingly, nobody actually wants the truth. The people shouting from the rooftops about corruption and bias, and collusion in the NYT turn around and uncritically read about how Hillary is in cahoots with gay alien pedophiles operating out of the back room of a pizza joint. Or, alternatively, how this is totally the week that Trump's finished. (And did you see the great burn some celebrity gave him?)
There is no "objective" news; people will always have different ideas of truth; however, i think in this case it's easy to discern a difference by the objective of the actors.
Some actors are purposely attempting to distort colloquial understanding via persuasion and misinformation, others are attempting to describe that colloquial understanding to the best of their abilities.
Our modern shortcoming is that we've just lumped all of these things under the shorthand "News" and thereby implicitly gave them all the same credibility and importance. It used to be that the Editorials and Opinions were clearly denoted and segregated in the newspaper... not anymore.
Though we may disagree on the semantics of what constitutes #realnews determining whether someone is attempting to describe the world through the prism of their own perspective or proactively trying to persuade others to think the same as them is pretty easy to see.
We can get mired in the idea that nothing is perfect, so why bother, or we can make the best of what we have & understand that it will never be perfect.
(apologies for the offtopic nature of this diatribe, this is something I've been spending a lot of time thinking about lately)
This kind of false equivalence is almost as toxic as the verifiable lies being promulgated by various "media" sources.
edit: https://www.nytimes.com/2015/02/15/magazine/how-one-stupid-t... maybe?
I haven't had a lot of disagreements with liberals, but I've had a fair number of discussions with republicans/conservatives who discount anything that opposes their worldview no matter how grounded in facts or science it is, and accept anything that reinforces it regardless of how unreliable or artificial the source is.
In fact, studies have shown that being shown facts that contradict our biases reinforces our biases rather than weaken them[1].
So bots that tweet facts that contradict the fake news people want to believe makes things worse and helps no one. It feels as though the only way to help someone overcome those biases is to prevent them from experiencing reinforcing information and to expose them to situations that contradict them.
[1] http://archive.boston.com/bostonglobe/ideas/articles/2010/07...
Twitter has every incentive to lie, to minimize, to shove this under the rug. As a fairly recent IPO with virtually flat/negative user growth, and lots of fed up people (like me) abandoning the platform all together, it is desperate to squelch any negative info that Wall Street might use against it.
Unfortunately, there's no favorable outcome for Twitter shareholders in either case. Twitter fesses up about its actual percent of bots (reality is likely closer to 50 percent than the 5 percent it claims) and its numbers go down even more. Twitter continues to lie and folks like the ones in this article expose them ... not good either because every advertising dollar it's getting is "truthfully" reaching fewer actual humans.
Do these kids not think that Twitter has the ability to do this.
This article comes off as if Twitter does not know what they are doing, which I think is hard to believe.
What I think these kids are going to find out, it's not that easy as it sounds.
Couldn't anyone do what these people are doing with a couple hundred bucks using Google cloud services and their natural langage API to label positive and negative tweets.
Seeing news like this which is not news, makes me realize how gullable people are about what goes into 'models'.
People have no idea how hard nlp is. That is all lol.
I think the "won't" may be accurate as if Twitter did it and chose not to act that would look terrible. And the scope of the bot problem may make their user base look a lot less attractive if they admitted it.
Granted everyone suspects, but there may be a valid reason they don't want to "know" at Twitter.
We agree - as independent of Twitter we have a considerable amount of freedom in how we build this. However Twitter can and should be doing much more. For example with a model they can start placing captchas before Tweets.
In addition to building the model we went about trying build our own bots. To do this we went on forums and contacted individuals selling "aged" accounts. It turns out its as simple as sending $4 over paypal to get a compromised account. These are accounts with histories, real followers, and real people behind them.
We bought 11 of these and were able to automate them within the hour of purchasing them. They also started receiving replies to their retweets and content almost immediately from all over twitter.
The ease of setting up these compromised accounts as bots was also incredibly worrisome. We've found high confidence heuristics to determine that an account has been compromised. If we can - Twitter should be able to as well.
We're ultimately a bit confused over Twitter's inactivity here. We also haven't heard anything from the company.
Here's our analysis if you want to read more.
analysis: https://medium.com/@robhat/an-analysis-of-propaganda-bots-on...
Don't know, but they aren't.
Twitter has a much harder post-discovery choice. Do they disable or publicly flag accounts, and catch flack for hitting real ones? Or do they monitor internally and wait for high confidence, then get in trouble for not doing enough?
I don't see how Twitter could offer a comparable service even with a comparable tool; being the official overseer of the question leaves them with too much responsibility.
First, Twitter can't do what the students are doing - an informal, unreliable service going "hey , bot!" is fine for some random people to offer, but the site owner can't chance high error rates on public statements. Twitter keeps suggesting they're doing something like this internally, which would make sense.
Second, the article offers no real evidence that this service works. It's a machine learning problem with no proven classifications in the dataset, so the learner will at best reproduce the developers opinions of what a Russian bot looks like.
I'll play around with the tool, but my initial guess is that it just learned "patriotic, inflammatory language, retweets only" as "bot", which isn't a breakthrough anyone can actually act on.
The model looks at hundreds of features of each profile, not just a few features such as patriotic, inflammatory language, and retweets as you suggested. As a result, the model has predicted profiles with very sketchy behavior, such as accounts of normal people that were compromised and become political propaganda accounts a a few years later. We wanted to bring up these accounts and the analysis behind it so we could show the techniques behind these organizations who run these bot-like accounts. We offer our own analysis of these accounts here: https://medium.com/@robhat/an-analysis-of-propaganda-bots-on...
The Wired take definitely left me cynical; it's all too easy to write a human-interest piece about ML approaches that don't actually work. But this is a much more concrete explanation of results, and I'm intrigued.
If you don't mind, could you offer any more clarification on your test/training set? I see the Medium piece talks about wanting to avoid selection bias from hand-classification, but the Wired summary just described hand-selecting 100 "ground truth" bots and then adding their followers. How did you know the followers were also bots? And how did you try to ensure the bots you selected were a reasonable sample?
Your right - twitter could do (and probably are doing) everything we're doing. They have billions of dollars and hundreds of engineers.
The value of building a model is that we can do a wide analysis on bot like activity. Separately launching botcheck.me as something that users can use is incredibly valuable from the ML side. Users essentially hand classify a bunch of false positives for us (to further train on) and also give us an idea of how are model is doing.
We aren't just doing sentiment analysis and you're right - NLP is hard. Fortunately at UC Berkeley we have some amazing CS professors that have been incredibly helpful in advising us while building this.
We're using LSTMs to learn the weights of various words. We've been using high confidence heuristics to generate our training data that aren't based primarily on tweet content.
One such example is looking at compromised accounts that have had their usernames changed.
Here's an analysis: https://medium.com/@robhat/an-analysis-of-propaganda-bots-on...
Methodology: https://medium.com/@robhat/identifying-propaganda-bots-on-tw...
We want to release a portion of training data so that others can build similar services. Let me know if you have any more questions :)
Sadly, there's a drawback to open systems (like this one), that the robot controllers can keep probing for weaknesses and keep changing their style.
Of course, such extension also begs the question: how will they (the creators) make money? "Volume", you say?
As I remember it, that's a major piece of what got Palantir started - PayPal's fraud-fighting system tracked known attacks algorithmically, but they needed visualization tools to locate new attacks.
That said, there's definitely an argument for attacking the ease of creation and longevity for bot accounts. Payment scammers are an endless problem because they profit on individual wins; bot networks need some sort of protracted presence to influence people.
For example there could be a Twitter account that replies to every bot tweet and those included in the bot's tweets. Accounts that have a high likelihood of being a bot would get a reply stating that the account is likely a bot.
I'm just spitballing—certainly, this would push up on Twitter's API limitations. But it seems like there a much better way to identify fake accounts than through a Chrome extension.
The chrome extension and website release also helps improve our model significantly. We get feedback on how it works in the wild and false positives that we can use to improve our model on.
Both my own and my father's accounts were classified as "exhibit patterns conducive to a political bot or highly moderated account tweeting political propaganda.", which is in a sense kinda accurate because they are solely political outlets for us. I think this is a feature, not a bug, and probably the only use case the service has provided for me.
I feel like twitter, tumblr, and other semi-anonymous networks have always been pretty political. I think the only thing that has changed is that people see twitter as news - stations actively report about what's going on on twitter - and that online political discussions are no longer dominated by liberal / social justice voices.
Because the cost to the kids for a false positive is 0. The cost to Twitter could be company destroying lawsuits.
They can't test against a large array of public bots because they're only detecting political bots, not everything automated. So they'd have to train/test on "accounts which are definitely known to be bots, but trying to hide it". Meaning, presumably, the least-convincing bots or bots specific to a previously-exposed network.
For some reason I feel like there are other ways......
They’re also pushing the “Russian interference” narrative, but as far as I can tell your product can’t point out the source.
That's the spirit! What was this saying again? "Better to convict 10 innocent men than to let at least 1 guilty go unpunished." That must be it...
If we assume that it's important to get rid off fake users, it must be because accessing a platform like Twitter is important.