Edit: Here are the original threads, I don't think the project got very far. http://www.reddit.com/r/announcements/comments/ddz0s/reddit_...
http://www.reddit.com/r/redditdev/comments/dtg4j/want_to_hel...
Edit: Here are the original threads, I don't think the project got very far. http://www.reddit.com/r/announcements/comments/ddz0s/reddit_...
http://www.reddit.com/r/redditdev/comments/dtg4j/want_to_hel...
EDIT: (Sorry for all the parentheticals.)
There are startups selling health data this way, I don't think it would be so bad for subreddit subscription data.
> It's guaranteed not to screw up your observations because, by definition, if something is statistically significant it has to show up often enough that it CANT be used to single out a source.
What's 'statistically significant' here? The usual p<0.05 convention? You realize that there can be multiple measurements or pieces of data all of which individually have p>0.05 but together have p<<0.05... Information leakage should be measured in bits, not p-values.
(This kind of aggregation is one of the benefits of approaches like meta-analysis.)
I just used post and comment histories, which suited my purposes fairly well because the larger project was looking into how memes spread.
http://www.reddit.com/r/redditdev/comments/dtg4j/want_to_hel...
Here's another version: