I don't see a reason to ban someone over this behavior as long as you can make sure that the fraudulent votes don't actually add any value to the content, which they shouldn't.
I don't see a reason to ban someone over this behavior as long as you can make sure that the fraudulent votes don't actually add any value to the content, which they shouldn't.
Visibility of the content is essentially predicated on early votes.
The problem is that an initial sort is built on limited data (quality content or just easy to upvote produces same results and latter is more common), and the subsequent feedback loop is built on initial visibility decisions.
Why do you say that? I've worked hard on this problem and would not call it easy.
In the less general case, you can still determine that a particular group of accounts never take action in the same small time windows.
The way I've approached this problem in the past relies heavily on meta data about the user. It's easy to determine if an IP is a proxy/VPN, which should immediately make a client more suspicious (when you're already suspicious about fraudulent voting).
Initially, account ages are important but if someone is able to get away with it for a long time, they can get complacent and keep using the same accounts, and at that point you can look at things like upvote/downvote ratios (or generally the lack of diversity in voting activity) compared to some expected average activity.
There's a lot of other tricks, but I don't know what your approach has been when working on this, and I know HN is a much different platform from what I worked with :P
> Why do you say that? I've worked hard on this problem and would not call it easy.
Are there reference data sets for this problem? If not, you or reddit should publish data sets.
I suppose the fact that there's already a detection algorithm biases the data, but it'd still be better than nothing.
That takes 30 people and it requires discipline to oney the die, but it'd be pretty hard to detect. I think, maybe I'm wrong?
What makes you so sure it doesn't?
Also, isn't the reddit source public? I presume that if it had this, it'd be easy to find it there, but I haven't looked, maybe I should :)
The problem with banning fraudulent users is that they are tenacious, thats's a big factor in their committing fraud in the first place, they are usually on a "mission," see: https://en.wikipedia.org/wiki/Wikipedia:Single-purpose_accou... (not really the same as a fraudster, but this is a common archetype)
They will just make another account or increase their efforts, better to keep them in their container within the system.
At which point, they create a new account.
> while shouldn't be that negative on the system as a whole, especially if there is suspected fraud.
Assuming it isn't computationally expensive to do this on a per-vote basis. If you have to do any sophisticated analysis, that may not be true.