What is non-linear data?
33 karma · joined November 15, 2013
What is non-linear data?
This ignores the fact that Wells Fargo is a publicly traded company and has an incentive to make their numbers look good for their shareholders.
... And to be clear, your assertion is that thousands of employees across different branches all independently decided to commit fraud and risk their jobs all for an additional 450 per pay check? That is your alternative explanation?
They go on to give supporting analysis in form of the false positive and false negative rates by race, which is pretty compelling evidence. You claim to not believe that because you cant find it in the notebook but its literally right underneath the Cox model section.
I was intrigued by this article and went a step further to plot the ROC curves and the evidence is solid. It's messy, but you can see it here https://github.com/stoddardg/compas-analysis/blob/master/my_... in cell 78. Its quite clear that the algorithm is choosing a different point of optimization on the ROC for white people (a more lenient one) than for black people. A white defendant with a risk score of 5 is as likely to commit a crime as a black defendant is with a score of 7. That's an obvious case where you could simply relabel and be more fair but their algorithm chooses not to.
I also hate when people abuse bad statistics and reasoning to sell page views.
http://economix.blogs.nytimes.com/2010/10/11/the-work-behind...
LASSO is a diamond because it represents the constraint that w_1 + w_2 <= 1. The region of (w_1,w_2) that satisfy that inequality is a square.
Ridge is a circle because it represents the constraint that w_1^2 + w_2^2 <= 1. The region of (w_1,w_2) that satisfy that inequality is a circle.
However if SSRN is just a paper-hosting service, then everyone should move to the arXiv immediately.
However I can also see it as a sickness in the sense that we fetishize this extreme success. Imagine a hypothetical person that has to make a choice between doing some risky start-up that will consume his/her life or getting a solid job that allows them to work 40 hours per week and then enjoy their personal lives. Nobody would explicitly blame the person who chooses the latter option but we celebrate (especially on HN) the people who choose the former option. We celebrate it to the point that there are many people who choose the harder option simply because they feel they are a loser for not going after it and throwing everything they have at it. I don't see that being a good thing that everyone who is on the fence decides to go for the risky/hard path simply because of some societal zeitgeist.
Then again, as you point out, those individual choices lead to our collective advancement, so it's a really hard balance.
It's a simple game where I show you two images from Reddit (SFW of course) and ask you to guess which was more popular. I'm using the data from the site as part of my research into the dynamics on internet popularity.
It's awesome because it's showing that Reddit is a pretty random/fickle thing. In the first iteration of the experiment (we're on the second now), people couldn't really do much better than randomly guessing. If one image had 10,000 upvotes and the other 10, people could only guess the popular one about 55% of the time. I wrote up some quick results in this blogpost: https://medium.com/@gregstod/guess-the-karma-2-0-82a224a691f...
I'd very much appreciate it if you played the game and donated a few data points :-)
https://onedrive.live.com/?cid=351055fdce7d0ff1&id=351055FDC...
Now of course this is not for all universities but I think it applies to almost any R1 university, top teaching college, etc. Its not competition for their job that drives them, its just competition against their self-image, competition against peers for prestige, etc.
In terms of creating your own site, a few researchers [1,2] have already done this and have some interesting work. But even with that, you still have the problem of accounting for position bias within the site (like HN doesn't really know if you skimmed the title of an article and decided to ignore it, or never read the title at all). But the experimental power you get with that is pretty cool.
And I have totally read that Facebook cascades paper and have more than a few thoughts about it. In fact, I have adapted their prediction-style results to the MusicLab data and you get really strong predictive accuracy (like 90-95% in terms of predicting whether a song will eventually be above the median popularity). However the accuracy you would achieve on Reddit or Hacker News data is considerably lower. I didn't really include those results in my paper because I'm not sure how they fit yet.
If I had one critique (which is not really a critique but a comment) is that the Facebook study doesn't really contradict Watts' point that popularity is hard to predict. The Facebook study shows that if you can observe the "initial conditions", then you can predict eventual outcomes pretty well but that's directly in line with the rich-get-richer effect that Watts et el demonstrate. To put it pithily, its easy to predict who gets richer if you observe who is rich.
Anyway, I could geek out about this for a long time but feel free to drop me an email at stoddardg [at] gmail.com if you're interested in chatting some more.
[1]: http://arxiv.org/pdf/1410.6744.pdf [2]: http://journals.plos.org/plosone/article?id=10.1371/journal....
http://arxiv.org/abs/1501.07860
The work isn't complete yet (always more to do) but the TL;DR is:
1. Yes, randomness governs a lot of article outcomes. Whether something hits the front page or not is pretty arbitrary.
2. However, conditioned on making the front page, popularity is actually a good reflection of "intrinsic quality". I think the ultimate relationship between popularity and quality is stronger than the MusicLab experiment suggests.
Of course not every example of a targeted ad is going to be controversial (like advertising merchandise of a sports team that I like versus the one that is just popular in my area) but its problematic that we don't have any sort of machinery in place to control/modify the examples that are politically/societally important.
At least if that machinery were in place, then we could transfer the agency of decisions to the humans that run google, rather than the algorithms that are just running in the background.
Epsilon greedy is a method for minimizing regret, that is the expected loss you occur from choosing options that are sub-optimal.
A/B testing's goal (or one of many goals) is to maximize the chance that, after the test is over, you select the best option going forth.
So e-greedy makes a conscience choice to not maximize its statistical confidence in certain options because it is trying to exploit the things it knows to be good. Meanwhile A/B testing is trying to balance the exploration so it can have that statistical confidence.
Hopefully someone with more expertise can chime in but I think this is the gist of it.
[1]: http://www.amazon.com/Everything-Obvious-Common-Sense-Fails/...
My best resource is foodwishes.blogspot.com. His videos focus completely on watching the ingredient, never himself, and I feel like he does a pretty good job of always explaining why we are adding in a spice or choosing a particular cooking method, or when you can feel free to improvise with your own preferred flavors. And there's also a good variety of different dishes, so its not just one type of cooking. I have been following recipes from foodwishes for about a year and I feel I've learned a lot (for an amateur) about how to cook.
It seems a bit silly to observe some behavior in a game and say "See that's a Nash Equilibrium, so game theory works", and then to turn around and observe some non-NE behavior and say "well in this game, BNE is clearly the right model, so game theory still works". And then yet again to observe some more behavior that conflicts with the theory (or to get rid of silly equilibrium) and say "ah, now we simply use perfect Bayes" or trembling hand equilibrium, or actually we were totally using correlated equilibrium this whole time.
In any case, it feels weird that a theory should behave like the "No true Scotsman" fallacy. We can always get the equilibrium by simply redefining what we mean by equilibrium.
And I think in the cases where the NE are predictive of actual play, then there's something intuitively obvious about the NE, and that you could have arrived at the same conclusions without using the tools and machinery of game theory. This is the point that Ariel Rubenstein makes (http://arielrubinstein.tau.ac.il/papers/74.pdf) when he says that game theory is useless.
That's not to say that I believe that game theory is never useful (for certain limited settings, like repeated auctions, it works well), nor do I think that it can never be useful (recent work in behavioral game theory is promising in my opinion) but I'm skeptical using these counterintuitive claims as examples of its use.
Honestly, I don't think that estimating number of submissions is a very good metric for the growth of Reddit (or Hackernews). If you want some proxy for its influence, you care about readership more than anything else. This thread [1] on Reddit (and the references therein) shows that 50% of Reddit activity comes from users that aren't even logged in. So even if you were just able to measure logged in users (which would still be a far greater number than submissions), you would still only estimate half the influence.
[1] http://www.reddit.com/r/TheoryOfReddit/comments/1khp85/logge...