I was planning to build a pretty sophisticated bayesian filter on top of that but it turned out to be unnecessary. The strategy I ended up with, which has served me well for many years now, is:
1. Block a small number of extremely spammy TLDs. There are about a dozen of these, including .biz, .casa, etc.
2. Anything received from someone I've sent mail to, or a sender I've previously marked as good, is assumed good.
3. Anything received from an address whose name contains an English word on a relatively short list of spammy words (discount, offer, etc.) is assumed spam.
4. Anything received from an address that I have received email from before and not marked as good is assumed spam.
That leaves emails from addresses that are sending me email for the first time. These are overwhelmingly spam, but after the four filters above there are few enough of these that I just scan them manually once a day or so.
The real key here is treating the first email from any given address as essentially a "contact request" with the default action being "deny in the future". That just turns out empirically to work incredibly well. More than 90% of my spam is from repeat offenders.