481 karma · joined January 30, 2014
Listen in on the phone conversations a plumber has and you'd get the impressions that no one's pipes work.
1. Underpowered tests are likely to exaggerate differences, since E(abs(truth - result)) increases as the sample size shrinks.
2. The much bigger problem I've seen a lot: when users see a new layout they aren't accustomed to they often respond better, but when they get used to it, they can begin responding worse than with the old design. Two ways to deal with this are long term testing (let people get used to it) and testing on new users. Or, embrace the novelty effect and just keep changing shit up to keep users guessing - this seems to be FB's solution.
Explanation: suppose there are only two things Google observes, GPA and coding ability, and that Google uses some correct decision rule to only hire those people where the sum GPA + coding ability > some threshold. Those who have lower GPAs will thus tend to have higher coding ability, otherwise they wouldn't have met the threshold to make it into the pool of hires they're analyzing - and, therefore, comparing "those with low GPAs that Google hired" and "those with high GPAs that Google hired" is not an apples-to-apples comparison.
In order to assess whether GPA should be used at all, they would need to look at how the people they didn't hire because of their existing policy would have performed.
More reading: http://beerbrarian.blogspot.com/2013/11/the-subtle-joys-of-s... and http://www.jamesmahoney.org/articles/Insights%20and%20Pitfal...
89k is actually a very large number.