HNHacker News
TopNewBestAskShowJobs

marketforlemmas

33 karma · joined November 15, 2013

submissionscomments
marketforlemmas··on Being a Data Scientist: My Experience and Toolset
> dealing with non-linear data

What is non-linear data?

marketforlemmas··on Wells Fargo’s ‘Living Will’ Plan Is Rejected Again by Regulators
> WF lost money on this

This ignores the fact that Wells Fargo is a publicly traded company and has an incentive to make their numbers look good for their shareholders.

... And to be clear, your assertion is that thousands of employees across different branches all independently decided to commit fraud and risk their jobs all for an additional 450 per pay check? That is your alternative explanation?

marketforlemmas··on Anil Dash Is the New CEO of Fog Creek Software
Is being chosen by Joel Spolsky not enough of an endorsement for you?
marketforlemmas··on To study possibly racist algorithms, professors have to sue the US
So I've read the PP piece and your article, and I think your criticism is way off-base. The most technical part of your argument relies on a p-value being .057 vs .05, which is not a good one. No one seriously believes that .05 is magical number that determines true from false; things that are close to it but not quite below .05 are not automatically false.

They go on to give supporting analysis in form of the false positive and false negative rates by race, which is pretty compelling evidence. You claim to not believe that because you cant find it in the notebook but its literally right underneath the Cox model section.

I was intrigued by this article and went a step further to plot the ROC curves and the evidence is solid. It's messy, but you can see it here https://github.com/stoddardg/compas-analysis/blob/master/my_... in cell 78. Its quite clear that the algorithm is choosing a different point of optimization on the ROC for white people (a more lenient one) than for black people. A white defendant with a risk score of 5 is as likely to commit a crime as a black defendant is with a score of 7. That's an obvious case where you could simply relabel and be more fair but their algorithm chooses not to.

I also hate when people abuse bad statistics and reasoning to sell page views.

marketforlemmas··on One Year as a Data Scientist at Stack Overflow
I don't think advertising for jobs is a zero-sum game, especially if the matching algorithm is good enough to match employers and employees that had no knowledge of each other before. If you are able to reduce those search frictions, you have created value. Several economic professors won the Nobel prize for their work in this area:

http://economix.blogs.nytimes.com/2010/10/11/the-work-behind...

https://en.wikipedia.org/wiki/Search_theory

marketforlemmas··on Machine Learning 101: What Is Regularization?
To be a little more clear about the diamond and circle...

LASSO is a diamond because it represents the constraint that w_1 + w_2 <= 1. The region of (w_1,w_2) that satisfy that inequality is a square.

Ridge is a circle because it represents the constraint that w_1^2 + w_2^2 <= 1. The region of (w_1,w_2) that satisfy that inequality is a circle.

marketforlemmas··on SSRN sold to Elsevier
Does SSRN come with any guarantee of peer review or minimum quality threshold? The only downside with arXiv is that anything can be posted. It works fine as a paper-hosting service but is terrible as a "social proof that my paper is OK". If SSRN comes along with such a reputation, then it might be a hard transition.

However if SSRN is just a paper-hosting service, then everyone should move to the arXiv immediately.

marketforlemmas··on Don’t Blame Silicon Valley for Theranos
Totally agree; this is basically a text-book example of the "No True Scotsman" argument.
marketforlemmas··on Why people still go to grad school
I think it's possible that its both a sickness and a boon. I fully agree with your point that the "fear of being a loser" does drive people to work harder, do better, etc. All of that has (generally) positive effects (well, assuming that you are willing to ignore the fact that many people will use less-tan-ethical means of making more money).

However I can also see it as a sickness in the sense that we fetishize this extreme success. Imagine a hypothetical person that has to make a choice between doing some risky start-up that will consume his/her life or getting a solid job that allows them to work 40 hours per week and then enjoy their personal lives. Nobody would explicitly blame the person who chooses the latter option but we celebrate (especially on HN) the people who choose the former option. We celebrate it to the point that there are many people who choose the harder option simply because they feel they are a loser for not going after it and throwing everything they have at it. I don't see that being a good thing that everyone who is on the fence decides to go for the risky/hard path simply because of some societal zeitgeist.

Then again, as you point out, those individual choices lead to our collective advancement, so it's a really hard balance.

marketforlemmas··on Ask HN: What are you working on and why is it awesome? Please include URL
Thanks for the suggestion. If you click on them, they'll pop out to the full size. Otherwise, I couldn't figure out how to get it to appear OK on both desktop and mobile (because I"m not really a web developer).
marketforlemmas··on Ask HN: What are you working on and why is it awesome? Please include URL
www.guessthekarma.com

It's a simple game where I show you two images from Reddit (SFW of course) and ask you to guess which was more popular. I'm using the data from the site as part of my research into the dynamics on internet popularity.

It's awesome because it's showing that Reddit is a pretty random/fickle thing. In the first iteration of the experiment (we're on the second now), people couldn't really do much better than randomly guessing. If one image had 10,000 upvotes and the other 10, people could only guess the popular one about 55% of the time. I wrote up some quick results in this blogpost: https://medium.com/@gregstod/guess-the-karma-2-0-82a224a691f...

I'd very much appreciate it if you played the game and donated a few data points :-)

marketforlemmas··on How a matchmaking algorithm saved lives
You are right that the general problem is intractable; a lot of the advances in these match making algorithms come from better heuristics and computing power.
marketforlemmas··on Economics Has a Math Problem
Here's a link to the slides that she used for the NBER talks. I've flipped through them and there seems to be some references to working papers (that Googling should be able to surface):

https://onedrive.live.com/?cid=351055fdce7d0ff1&id=351055FDC...

marketforlemmas··on University labour strife underscores cost of tenured academics
I agree with the original point. In my experience, there's such a strong selection process in who even tries to get tenure that most people simply can't turn off their drive because its so fundamentally a part of them. Do they change the shape of their efforts? Sure. They might stop trying to go for many small publications and go for a longer view, try a new subfield, etc. But the base level of effort never really diminishes.

Now of course this is not for all universities but I think it applies to almost any R1 university, top teaching college, etc. Its not competition for their job that drives them, its just competition against their self-image, competition against peers for prestige, etc.

marketforlemmas··on What Makes Hacker News Fame?
Thanks for the feedback (and the Watts' reference). Building a plugin is definitely an interesting idea; I hadn't thought of that before. I guess the problems would be three-fold. First, I don't know how to do that :-). Second, there are probably a bunch of ethical concerns/IRB issues that would stand in the way of academic publishing (but thats not huge). Third, and the only fundamental issue, is that the self-selection into using that plug-in would bias the estimates of intrinsic quality. Still its a pretty good idea but I'm currently trying to get access to more fine-grained data in other ways, so we'll see.

In terms of creating your own site, a few researchers [1,2] have already done this and have some interesting work. But even with that, you still have the problem of accounting for position bias within the site (like HN doesn't really know if you skimmed the title of an article and decided to ignore it, or never read the title at all). But the experimental power you get with that is pretty cool.

And I have totally read that Facebook cascades paper and have more than a few thoughts about it. In fact, I have adapted their prediction-style results to the MusicLab data and you get really strong predictive accuracy (like 90-95% in terms of predicting whether a song will eventually be above the median popularity). However the accuracy you would achieve on Reddit or Hacker News data is considerably lower. I didn't really include those results in my paper because I'm not sure how they fit yet.

If I had one critique (which is not really a critique but a comment) is that the Facebook study doesn't really contradict Watts' point that popularity is hard to predict. The Facebook study shows that if you can observe the "initial conditions", then you can predict eventual outcomes pretty well but that's directly in line with the rich-get-richer effect that Watts et el demonstrate. To put it pithily, its easy to predict who gets richer if you observe who is rich.

Anyway, I could geek out about this for a long time but feel free to drop me an email at stoddardg [at] gmail.com if you're interested in chatting some more.

[1]: http://arxiv.org/pdf/1410.6744.pdf [2]: http://journals.plos.org/plosone/article?id=10.1371/journal....

marketforlemmas··on What Makes Hacker News Fame?
At the risk of plugging my own work, I did a follow-up to the (really awesome) Watt's experiment using data from reddit and Hacker News:

http://arxiv.org/abs/1501.07860

The work isn't complete yet (always more to do) but the TL;DR is:

1. Yes, randomness governs a lot of article outcomes. Whether something hits the front page or not is pretty arbitrary.

2. However, conditioned on making the front page, popularity is actually a good reflection of "intrinsic quality". I think the ultimate relationship between popularity and quality is stronger than the MusicLab experiment suggests.

marketforlemmas··on Netflix’s Secret Special Algorithm Is a Human
You might also be interested in Duncan Watts' book "Everything is Obvious", which explores this tendency we have to assume that observed success equates to some skill. Its a nice read.
marketforlemmas··on The Code We Can't Control
I don't think the issue is whether ads should prompt them to lose weight, or they shouldn't. I think the crux of the issue is that we (humans) are ceding this judgment (which have political and societal ramifications) to the ad matching algorithm. Some people (myself included) find that bothersome.

Of course not every example of a targeted ad is going to be controversial (like advertising merchandise of a sports team that I like versus the one that is just popular in my area) but its problematic that we don't have any sort of machinery in place to control/modify the examples that are politically/societally important.

At least if that machinery were in place, then we could transfer the agency of decisions to the humans that run google, rather than the algorithms that are just running in the background.

marketforlemmas··on Show HN: The problem with the epsilon greedy method
Interesting comparison but, by my understanding, epsilon greedy and A/B testing do not solve the same problem.

Epsilon greedy is a method for minimizing regret, that is the expected loss you occur from choosing options that are sub-optimal.

A/B testing's goal (or one of many goals) is to maximize the chance that, after the test is over, you select the best option going forth.

So e-greedy makes a conscience choice to not maximize its statistical confidence in certain options because it is trying to exploit the things it knows to be good. Meanwhile A/B testing is trying to balance the exploration so it can have that statistical confidence.

Hopefully someone with more expertise can chime in but I think this is the gist of it.

marketforlemmas··on Leaving Academia (2013)
To put this more bluntly: You should not do a phd in the sciences (in the US) if you are not getting paid. Its indicative that either A) the department has insufficient resources (bad situation) or B) they really don't want you as a graduate student (another bad situation).
marketforlemmas··on Homo Narrativus and the Trouble with Fame
For people who enjoyed this article, I highly recommend Duncan Watts' book "Everything is Obvious" [1]. Its basically the thesis of this article expanded out over the course of a book with a lot of interesting examples and lessons drawn from scholarship (both Watts' own work and many others).

[1]: http://www.amazon.com/Everything-Obvious-Common-Sense-Fails/...

marketforlemmas··on Make Dinner: A Home Cooking Manifesto
Completely agree; learning to cook shouldn't be about cooking different recipes, it should be about generalizing the lessons from each one. That said, I think an effective way to learn it is just to try cooking a million different things and keeping notes (mental or physical ones).

My best resource is foodwishes.blogspot.com. His videos focus completely on watching the ingredient, never himself, and I feel like he does a pretty good job of always explaining why we are adding in a spice or choosing a particular cooking method, or when you can feel free to improvise with your own preferred flavors. And there's also a good variety of different dishes, so its not just one type of cooking. I have been following recipes from foodwishes for about a year and I feel I've learned a lot (for an amateur) about how to cook.

marketforlemmas··on Game Theory Is Counterintuitive
I take some issue with the idea that we can simply just rely on another equilibrium refinement and say "no big deal".

It seems a bit silly to observe some behavior in a game and say "See that's a Nash Equilibrium, so game theory works", and then to turn around and observe some non-NE behavior and say "well in this game, BNE is clearly the right model, so game theory still works". And then yet again to observe some more behavior that conflicts with the theory (or to get rid of silly equilibrium) and say "ah, now we simply use perfect Bayes" or trembling hand equilibrium, or actually we were totally using correlated equilibrium this whole time.

In any case, it feels weird that a theory should behave like the "No true Scotsman" fallacy. We can always get the equilibrium by simply redefining what we mean by equilibrium.

marketforlemmas··on Game Theory Is Counterintuitive
These are interesting examples of using game theory to model certain situations but none of these examples show that game theory "tells us" something about the world. In broad strokes, game theory claims to have a descriptive model of human behavior but people very rarely follow the predictions of game theory (and that is even when you're in situations that are simple enough to have a stab at applying game theory). This is particularly true for Nash Equilibrium, which it seems most of your examples rely on.

And I think in the cases where the NE are predictive of actual play, then there's something intuitively obvious about the NE, and that you could have arrived at the same conclusions without using the tools and machinery of game theory. This is the point that Ariel Rubenstein makes (http://arielrubinstein.tau.ac.il/papers/74.pdf) when he says that game theory is useless.

That's not to say that I believe that game theory is never useful (for certain limited settings, like repeated auctions, it works well), nor do I think that it can never be useful (recent work in behavioral game theory is promising in my opinion) but I'm skeptical using these counterintuitive claims as examples of its use.

marketforlemmas··on Blue Eyes Logic Puzzle
The islanders don't know that there are only blue and brown eyes. For all they know, their eye color could be green, red, etc. Hence, the logic only applies for the blue-eyed people.
marketforlemmas··on Reddit is Growing Slowly, but Surely
Have you looked at number of comments instead of number of submissions? That would seem to be a better estimator of user growth.

Honestly, I don't think that estimating number of submissions is a very good metric for the growth of Reddit (or Hackernews). If you want some proxy for its influence, you care about readership more than anything else. This thread [1] on Reddit (and the references therein) shows that 50% of Reddit activity comes from users that aren't even logged in. So even if you were just able to measure logged in users (which would still be a far greater number than submissions), you would still only estimate half the influence.

[1] http://www.reddit.com/r/TheoryOfReddit/comments/1khp85/logge...

marketforlemmas··on Alexis Ohanian on Colbert Report [video]
So which one of your start-ups is one of the top 100 websites in the world?