Vowpal Wabbit is made to scale. You set a fixed bitsize and words and n-grams are hashed. So if you expect 2^32 unique words you set the bitsize to around 32. More data is usually better. Linear speed-ups by adding parallel machines.
I too think that HN's user patterns could be different than other web estates. With large scale spam filters like at Yahoo mail, I believe they employ two (or more) models: One fitted on your inbox, and one fitted on everyone's inbox. That ensemble model should be able to specialize on your behavior, yet still be able to detect general spam that it has already seen in other boxes.
>Is this the basic assumption that close documents also have close information entropy?
Yes, that is the gist of it. The better the compressor, the closer NCD will approximate NID.
A simple principle: Compressors do a better job on repeating data patterns. If two files or documents share data patterns, then adding these together and compressing, will result in a smaller filesize, than if you concatenate and compress two files that don't share any data patterns.
NCD works on text, but not as good as other algo's for NLP. Sometimes PAQ (very slow, but efficient compressor) is used on genome data, or bzip on binary files like virusses. I don't think it will be practical here, since for a comparison every other file would need to be concatenated and compressed. If not using a fast compressor like Snappy or Gzip this would take a while, over a simple cosine distance between tokens.