Botrnot: an R package for detecting Twitter bots
mikewk.shinyapps.io
mikewk.shinyapps.io
Given the preposterous error rate, I deem there is no actual classification logic in R, and instead it uses the (very fallible) humans to do the actual classification via a Mechanical Turk-style API.
Wait, do I need to add "/s" to this post, or is it obvious enough?
Dread to think how they came up with this
There are (of course) some useful bots, but lots of incredibly harmful bots, they they should be treated differently to actual humans.
But Twitter can't ship product, so it's not really worth suggesting what they should do.
In the mean time, my colleagues and I got a nice WWW18 conference paper about a new unsupervised (!) way of detecting some type of bots on Twitter. Like most things it's completely obvious in retrospect...
EDIT: Perhaps it's this one ? https://arxiv.org/pdf/1711.09025.pdf
That's supervised for one thing. Also this is bot detection, not attempting to detect fake news.
I'm not the first author and I'm not sure what the convention is for pre-prints at WWW.
It's the kind of thing I'd love to talk about because it's so damn amazingly effective, but it isn't really my place to say anything.
> Uses machine learning to classify Twitter accounts as bots or not bots. The default model is 93.53% accurate when classifying bots and 95.32% accurate when classifying non-bots. The fast model is 91.78% accurate when classifying bots and 92.61% accurate when classifying non-bots.
Overall, the default model is correct 93.8% of the time.
Overall, the fast model is correct 91.9% of the time.
How is this accuracy determined? There is no information available explaining how this determination is quantified, nor what the caveats are.
However, the percents are from the training set; there’s no test/validation set, which is a problem when working with bespoke text data as a feature.
> The default [gradient boosted] model uses both users-level (bio, location, number of followers and friends, etc.) and tweets-level (number of hashtags, mentions, capital letters, etc. in a user's most recent 100 tweets) data to estimate the probability that users are bots.
Not an exact science, but shows what you can do and deploy quickly with R/Shiny.
The author’s rtweet package is very good for making quick Twitter data visualizations.
Plus it said I was a bot (85% chance anyway), and I'm pretty sure I'm not.
[1] https://twitter.com/mkearney (I assume this is his account?)
If you haven’t:
1. Downloaded RStudio IDE
2. Built a hello word Shiny App (better still for a flavor of the thing a hello world app using the shiny dashboard package)
3. Deployed your app to shinyapps.io
I highly encourage you to do so if for no reason than to see how streamlined RStudio has managed to make web app deployment for people who often don’t have much of a programming background.
I’m continually impressed with the work RStudio does, even if I’m a curmudgeon and still write all my code in Emacs instead of their IDE. If RStudio expanded to support Python similarly well, I imagine they could really be the place most data scientists work.
Definitely worth a watch: https://media.ccc.de/v/34c3-9268-social_bots_fake_news_und_f... (The video is German but there should be a translated version of it on the site)
While “bots” featured large in initial discussions, the details that have emerged since, especially in the recent indictment, actually point at software-assisted humans as the main avenue for these activities.
I believe the “bot” monicker because (a) it’s how most people would tackle the task, perhaps ignoring the cheap labor pool this shadowy PR operation had access to, and (b) the bot-like behavior exhibited by many of these accounts, which was actually a result of not being native speakers, following a script, and having no deeper knowledge of politics and culture to fall back on.
What I'd want to also see however, would be curation of a relevant corpus of tweet content that could be integrated in the classification process in terms of actual tweet content (bag of words, sentence structure, & other NLP features, images, other media) being used to determine a given account's similarity between bots and non-bots -- instead of relying essentially on metadata similarity in addition to crude auto-extracted text features like simple total number of words and character counts. I can't say for sure if it'd improve predictions but would seem key for this type of problem.
Link to github: https://github.com/mkearney/botrnot