Writing an automated service that classifies if a twitter account is a political bot. No word on their error rates.
They can't test against a large array of public bots because they're only detecting political bots, not everything automated. So they'd have to train/test on "accounts which are definitely known to be bots, but trying to hide it". Meaning, presumably, the least-convincing bots or bots specific to a previously-exposed network.
For some reason I feel like there are other ways......