Sentence = "This sentence is semantically and syntactically valid."
P(Sentence) = log(p(START,START,This)) + log(p(START,This,sentence)) + log(p(This,sentence,is)) + log(p(sentence,is,semantically)) + log(p(is,semantically,and)) + log(p(semantically,and,syntactically)) + log(p(and,syntactically,valid)) + log(p(syntactically,valid,.)) + log(p(valid,.,STOP)) + log(p(.,STOP,STOP))
where START and STOP are special symbols that aid in determining the proximity of a word to the beginning and end of a sentence.
If your training set fails to sufficiently generalize, you could use Bayesian inference to estimate the likelihood that the sentence is spam. Under this framework, you'd be calculating the posterior probability of the sentence being spam given the observed sequence of n-grams, which combines (i) the inherent likelihood that any sequence of words is spam and (ii) the compatibility of an observed sequence with (i), which is proportional to the impact it has on (i).
[1] http://storage.googleapis.com/books/ngrams/books/datasetsv2....
Then, run the model over your data and start playing whack-a-mole (and refining the model).
You can also turn off links in comment bodies and the URL field of the comment form to try and prevent scrapers from even finding you a worthy target. Won't help identify spam though.
Finally, centralized spam identification systems like Akismet work really damn well because they are watching the whole site network at once and can use those heuristics to identify spammers rather than the actual spam content itself.