It basically says "we rebuked one study, therefore bots aren't real."
Anyone who has been on Reddit for the last year has seen government run botnets posting spintext by the hundreds.
Thinking there are no bots is wildly naive
It basically says "we rebuked one study, therefore bots aren't real."
Anyone who has been on Reddit for the last year has seen government run botnets posting spintext by the hundreds.
Thinking there are no bots is wildly naive
https://blog.plan99.net/fake-science-part-ii-bots-that-are-n...
One of the points I make is that social bot research is a prolific academic field. There are nearly 10,000 papers indexed on Google Scholar on the topic and that's not all of them. And basically all of them are fantasy. Clearly, bots do exist, but the techniques in these papers are not able to detect them reliably and the authors don't seem to care.
They show some papers use a bad bot detector with even worse (and unrecommended) settings. Worse, they explicitly skip bot accounts that twitter confirmed, and thus evaluate on accounts twitter didn't think were bots. That kind of mistep means it fails to show botometer is a good/bad classifier (even if we all know that), so until they fix their methodology, their misreproduction is bad scientific grounds for disqualifying all papers using it.
This is why we have peer review and you can pick what level of peer you trust your work at. I like the intent of the paper. But, I wouldn't have let it pass peer review as-is.
I do agree social scientists deserve better tools because botometer is more of a crude footgun at this point. A lot has happened since botometer, and what may have made some sense 5-10 years ago doesn't make sense today. They could have written a scientific paper showing this and I'd 100% believe it, but instead they used garbage methodolgy and people are inferring whatever they want from that.
Some of the papers mentioned in the analysis don't use Botornot, but are still deeply flawed. I should point out here that I'm cited by the authors (Hearn2017) and the paper I looked at doesn't use Botometer at all, nor did I exclude accounts identified by Twitter for the simple reason that no specific accounts were named by the researchers at all.
"I like the intent of the paper. But, I wouldn't have let it pass peer review as-is."
Then you will be pleased to know that despite peer reviewers signing off on thousands of incorrect, evidence-free papers talking about social bots, this paper - which provides ample proof of its claims - remains a preprint because none of the journals interested in social bots are willing to publish it. What a surprise.
You sound like someone who is or was a social scientist? I think you should take it from a guy who did get paid to fight bots at Google: none of this research matters to or is read by the people whom it's supposed to benefit. Industry ignores it or occasionally points out that it's flawed, as Twitter did here: https://blog.twitter.com/en_us/topics/company/2020/bot-or-no...
"What about tools like Botometer & Bot Sentinel? ... This is an extremely limited approach ... These tools do not account for these common Twitter use cases, how far we’ve come, and how things have evolved. As a result nuance can be lost. The outcome? Binary judgments of who’s a “bot or not”, which have real potential to poison our public discourse — particularly when they are pushed out through the media."
Furthermore, it is questionable to me to know 10% were already confirmed removed and then state the reproducers found all of the remaining were great.
If I went deeper, I bet even more issues. As is, this is already enough to remove any word like "all" and various weasel word phrasings.
---
Ad-hominem shouldn't matter. But:
Professionally, I work with top gov agencies, top 3 cloud providers, F500s, startups, misinfo teams, etc. to provide pluggable core tech for their core graph intelligence pipelines (Graphistry), such as for sigint tasks. Botnets, hackers, account take overs, account abuse, etc - threat hunting, threat research, AML teams, digital human trafficking. As a #data4good effort, we help non-profits & researchers on the same things, such as via our free tier & some of our own hands-on help. We're just a tiny startup, but helped with massive international takedowns, flagged one of the Jan 6 indictees back in December (that FB and friends didn't take down util after more pressure), regularly help folks publish breaking news, etc.
Before, I was a well-cited academic (Test of Time award, Best of Years, etc), and work spanned early R&D for engineering tools your colleagues likely use today, and empirical & social research stuff for being able to judge work like this. Which I am, harshly. It looked like an easy win of a paper if they were dilligent, but ended up needlessly weasely.
---
I agree with you on the tools weaknesses and big companies ignoring researchers, but for different reasons than researchers being wrong.
The Twitter response is stock corporate CYA, so I don't take that at face value.
Instead, consistent with what I've written, tools like botometer are largely irrelevant to modern practitioners here because reasons like:
- Abuse teams at free social platforms are at battle with themselves and their surrounding organization & CEOs. The Facebook whistleblower accounts may particularly resonate well with you on why being factually right/wrong has little to do with what platforms do.
- Botometer is dumb. Twitter went as far as acquiring the startup of one of the top GNN researchers: In contrast, botometer is from the dinosaur era.
- Botometer is data-starved. Data is thin (no IPs, ...) and at low volumes (capped APIs). When you have APIs, app use logs, etc., same as someone self-hosting even the tiniest and most boring online store, you can do much better informed analyses. It's like a crypto app fraud & compliance team who only looks at blockchain data to decide good bad account/transaction, and ignores all the app log data & threat intel data.
You've been implying that this is a methodological flaw of the paper we're talking about but Section 3.2 is an attempt to replicate someone else's study, one which claims Botometer isn't flawed. As they clearly point out the study wasn't replicable because the original authors didn't record/supply the data used to do the original classification, just numeric user IDs. Non-replicable research is exactly the kind of problem they're calling out here!
Gallwitz and Kreil have pointed out by this paragraph that the way the accounts were pre-filtered will already avoid Botometer's biggest weaknesses. Yet, despite this fact, "A surprisingly high number of the 121 alleged bots in our sample are in reality individuals with academic or professional credentials, many of them directly related to the topic of vaccines" and "A number of the 121 alleged bots are, in reality, the official Twitter accounts of health-related organisations", and "some of the alleged vaccine bots are, in reality, Twitter accounts related to agricultural topics and pets".
Section 3.2 points out that the study isn't replicable, which is by itself sufficient to discard it. Yet even when this problem is ignored and a replication is attempted anyway, a convincing takedown is easy to mount. It shows clearly that the authors never bothered to look at their "true positives" and is thus a great example of the problem in action.
If you think it's easy to do a much better rebuttal to all these bad papers than these two German authors managed, great, let's see it? It sounds like you would be pointing out flaws in a widely used competitor, after all. I agree completely that Botometer is dumb and data starved, but where we seem to disagree is that you think it's possible to give social 'scientists' better tools. I don't think it is. Users consent to sharing their data with services when they tweet, post, etc but they don't consent to sharing it with random academics. The whole field of academic social bot research is DOA because of this problem, yet, academic institutions like universities, journals and granting bodies systematically cannot accept this. Anything they produce will always be a data starved PRNG.
Half of them being flagged for inauthentic behavior is arguably NOT a false positive. I doubt the authors checked whether it is really Bernie Sanders or a gaggle of twenty-somethings working with some DNC tooling.
This paper is frustrating.
Nearly all is only something that you would say if you had no familiarity with the field
The authors of the paper concluded that, in the future, any claims that withhold the list of account handles classified as bots should be ignored because the claims can’t be independently verified.
So they didn’t rebuke just one study but datasets underlying most studies on the topic, they analyzed Twitter rather than Reddit, and they did not claim that bots don’t exist. If I am mistaken, someone who’s analyzed the paper in more depth than I did can correct me.
> The field of “social bot” research is fundamentally flawed. While “social bot” researchers have received an enormous amount of public attention, their methods are highly dubious. And, as we and others have demonstrated, they fail miserably and consistently when evaluated under real-world conditions. Studies claiming to investigate the prevalence, properties, or influence of “social bots” have, in reality, just investigated false positives and artifacts of the flawed detection methods employed.