If anyone wants to try to train a filter to detect this sort of comment, I'd be very interested to see the result.
If anyone wants to try to train a filter to detect this sort of comment, I'd be very interested to see the result.
To any that have experience getting comments data from HN -- what's the fastest, most polite way to do this? And am I correct in remembering that there's some aggressive rate-limiting for crawling the site?
http://www.btscene.eu/details/2240774/Hacker+News+Database+o...
(That might introduce a confounding factor, though--namely that by alleviating people's urge to downvote the article by giving them a [nonfunctional] button to do just that, people might stop upvoting the dismissive comments. Hmm....)
Wouldn't that require real AI though? I thought for a minute that NLP (Natural Language Processing, not the other meaning(s) of the acronym) might help, but then thought that it may not work for cases where the comment is quoting another comment. Note: I'm not at all an expert in any of those fields, just interested.
You could probably find a way to mark negative and positive comments. Whether the resulting algorithm would be fine-grained enough to semi-reliably mark 'middlebrow dismissal,' I really don't know. Actually, as somebody who has worked on that stuff in the past, I don't think it would be very easy.
Agreed I don't think it would be an easy task, but I wonder how would perform a "bag of word" approach.
Harder part as I see it would be to categorize the comments on middlebrow dismissal / Not dismissal. It seems like we would be spending more time preparing the data than in the algorithm itself.
The hardest part is always getting the data into usable form. Its not as much fun as fitting models, but its definitely the majority of any role where people pay you to do this kind of stuff.
There's a lot of good research on forums (pm me if you want a bibliography i collected for a previous role), and short texts have become a bigger deal post Twitter. I completely agree with pg on the somewhat annoying nature of comments such as the GP.
I did have worked with ML before but mostly with images which are (IMHO) way easier to put in a format depending on the problem.