>The accuracy is often equal or slightly worse than in-memory techniques.
Slightly worse? because it has to go to different nodes to do the match, and hence incurs a slight delay which then incurs a bit inconsistency of the sample base?
> I too think that HN's user patterns could be different than other web estates.
That's a better generalization...
perhaps some more mining could be done based on this, such as the social clustering ; and also perhaps construct different language/wording models for the different interest/area groups (by posts) -- after all, there is no language model that fits them all...
And as in terms of spam detection, I assume if an ID, or originating IP, replies to almost every post, then it's unlikely the quality of his post would be high...Well, this returns to that fundamental question, how do you define "spam" in the space of HN?
> With large scale spam filters like at Yahoo mail, I believe they employ two (or more) models: One fitted on your inbox, and one fitted on everyone's inbox. That ensemble model should be able to specialize on your behavior, yet still be able to detect general spam that it has already seen in other boxes.
Interesting...It would be interesting to see how often the two models produce inconsistent results...Then if you biased towards one, then what's the use of the other one? Or perhaps they devised some strategy to combine the two models...
>NCD works on text, but not as good as other algo's for NLP. Sometimes PAQ (very slow, but efficient compressor) is used on genome data, or bzip on binary files like virusses. I don't think it will be practical here, since for a comparison every other file would need to be concatenated and compressed...
perhaps
NCD and NID are theoretically charming...but perhaps they are only good for pure coding without much context to depend on...there are just tons of information outside the analysis target, especially a natural language, that would be hard to take into the calculation of the entropy model...By the way, in terms of the similarity analysis based on compressed binary, there was a paper (peHash) published a few years ago that was interesting....
https://www.usenix.org/legacy/event/leet09/tech/full_papers/...
Another issue with the entropy is, it's easy to inject some tokens to manipulate the frequency...In this case, some preprocessing should be in place...