What about using baysian filtering? We already have a lot of data on good and bad comment styles, baysian filtering could give us an indication if a comment is violating HN guidelines.
I am not so optimistic about submissions because the data we currently have is tainted by submissions we don't want. However by using the technique you have described, we could probably achieve better results.