Using HyperLogLog to detect voting rings (2013)
opensourceconnections.com
opensourceconnections.com
or an even more basic / stupid question - what exactly is a "voting ring"? does it mean people having some kind of a pact to always vote for each other?
While there are a few clever voting rings, the vast majority are just one guy with many sockpuppet accounts upvoting self-serving crap.
HN doesn't need anything as fancy as HyperLogLog. Most voting rings are apparent within the first 20 votes so elementary data structures work fine.
(I know how the HN system works and anecdotally how Reddit's used to work, but I'm being a bit generic here to avoid revealing that One Weird Trick to Get to the Top of HN.)
With that though, of course this is better than a model that requires no noise. Forcing noise from the edges reduces the set of entries requiring human review and, still, human review might reveal with sufficient weight the existence of a possible ring. But, this is an evolution to the cat+mouse stage. Hence, rings could still survive if clever?
With that question (ring survival with measured noise) set aside, it seems if the cat+mouse were to come in earnest that reputation dependent models would be a driver for mapping user "types" (social, political, behavioral models) in a much more accurate manner than the current driver of advertising (google wanting to increase price of adds)
1. Pick U out of N. Submit the story as U.
2. Pick [U1, U2...UX] out of N, and upvote the story enough to carry it to the front page. Let the random front page upvotes do the rest.
3. ...
4. Profit.
Also, why not write a bot that randomly upvotes stores for your N controlled users (presumably, they are not real people)?
For example: if you buy likes on Facebook, the likebots won't just like your page and other spammers, but will also like many innoculous pages as well (Coke, Obama, "Facebook Needs a Dislike Button!1!", etc.)
http://research.neustar.biz/2012/10/25/sketch-of-the-day-hyperloglog-cornerstone-of-a-big-data-infrastructure/