Hacking Hacker News Headlines
metamarketsgroup.com
metamarketsgroup.com
Why showing the future is essential to acquiring data
Noted.
I have a feeling that headline wouldn't do all that well, but it does seem to be keeping with google's culture.
From the article, what's the difference between "data |" and "data -" ?
Also, no clue if the factors you pulled out are orthogonal.
Only thing I'll add as a data critique, the negative factors are reported as things to avoid. But, in fact, all of the reported on titles actually made it onto the Hacker News front page (1). There are an awful lot of submissions that never make it that far. In fact, the significance of the findings indicate that those terms make it onto the front page A LOT (2). I don't think the negatively correlated terms should necessarily be viewed as failures. Just less successful. My own suspicion is that those titles do draw eye-balls, but someone using titles like those is also likely to be kind of a bad writer, preventing those stories from getting upvotes. It would be very hard to prove a correlation between quality of title and quality of writing, though.
(1) I believe. Hard to tell from the post.
(2) Otherwise there wouldn't be enough data for them to be significant.
Especially since the author of the linked article isn't always the one who submits the article and therefore gets to choose the HN title.
Another way of looking at it would be the amount of time that a post spends on the front page.
"Essential Lessons Showing How to Hack Hacker News with Data Visualization"
and it will be ranked #1 in no tim... never mind :D
This headline uses none of the hacks described in the article, yet it is ranking quite well.
Perhaps people should focus on letting the content speak for itself rather then using tricks like this?
My proposal for a good headline according to the numbers in this article: Showing why impossible future controversy survived the problem could hire data. Score: 1.3 (could) + 1.2 (problem) + 1.3 (survived the) + 1.0 (controversy) + 0.9 (impossible) + 0.7 (why ___ future) - 3.3 (11 words) + 2.6 (showing) + 0.5 (hire) + 1.9 (data [END]) = 8.1. For comparison, Why showing the future is essential to acquiring data gets 1.4 (essential) + 0.7 (why ___ future) - 2.7 (9 words) + 2.6 (showing) + 1.7 (acquiring) + 1.9 (data) = 5.6 -- except that it doesn't really get the points for "essential" (not at start) or "why ___ future" (two words in between) or "acquiring" (not in second place, word isn't quite right). Of course my headline has the little drawback of being total nonsense.
[[ Edit: deleted question about what 'k' is for the discretized 1{ rank <= k } response. It's mentioned in the article ]]
I'm genuinely curious how to do coef significance testing for L1-regularized models. I once saw someone ask this at a Tibshirani talk and he said "oh we have no idea, we've resorted to the bootstrap before".
% of replicates with (coef==0) is potentially much more clever, especially since that's the test we want to perform anyway. i'll run that over the data and see what changes.
From a frequentist perspective, counting zeroes don't make much sense because under the null of coef=0 there is still a chance you don't estimate coef=0, even after regularization.
I think the question is these don't look like NormalCDF(coef/se) p-values given the coef and se you report. They tend to be too small.
right that's my questionFind someone who reads Hacker News [1] and blogs. Subscribe to their RSS feed. There's a trick to [1], no doubt. But for most people the time savings far outweigh the occasional mismatch between your interests and the delegate's.
[1] For the same types of articles you do, ideally.