For example, the comment in question contains a url. I could easily imagine that turning out to be a valuable predictor. The defining quality of the middlebrow dismissal is that it's a cache dump of the writer's prejudices, and someone doing that doesn't even take the time to think, let alone look up urls; they're not even really writing to inform.
He was doing some research on detecting which comments are authoritative based upon textual analysis (no username or social analysis). They made a complicated topic model, but found that the following heuristic is almost as good for automatically detecting authority in comments:
Favor the person with the broadest vocabulary compared to other people in the thread.
This was evaluated on Yelp and Goodreads. IIRC it may have also been tested on HN data.
(reference: Alexandre Passos, Jacques Wainer, Aria Haghighi, What do you know? A topic-model approach to authority identification.)
The problem with disclosing these heuristics as part of your filtering algorithm means that people will try and game the system. They'll include URLs and expanded vocabulary to get higher-ranked comments. And then we win. (Relevant: http://xkcd.com/810/)
But also potentially looking at things like the urls that get linked into regular discussions such as weight loss or political topics. That might go beyond the immediate scope, but comments with those links might have other factors similar to middlebrow dismissals, so they might be worth building into the model.
Trying to catch something like this post would be interesting to add. Maybe an anti-indicator if the link adds value. Then you get into figuring the value of the content of the link. Maybe compare it to the content of the original link, which you would want to be similar, but not too similar. Yeah, that might get involved.
Google's Pagerank didn't have to look at the page contents to understand which pages are higher quality. At Quantcast, we categorize related content based on who is visiting, no need to look at the content itself.
The best way to find out is to open up the HN data, in either a closed or open research project. I'd help fund a closed Kaggle competition with NDAs if that's whats wanted. Just consider it a way to learn what works. pg himself would be the final arbiter and likely implementor of what works.
For instance, this article seems to have been spawned by the recent Businessweek article re: financial forecasts for the airline industry and specifically what it may do to VA: http://www.businessweek.com/news/2012-10-17/virgin-america-t...
The tone of this particular write-up sounded very doom and gloom when in actuality VA seems to be following the annual trends of the airline industry while also pursuing an aggressive growth strategy. Not unlike many extremely popular startups (whether it works out or not, well, that's still quite a few years away).
(BTW, thank you for saving me the necessity of reading past the first page of the article.)