I run my own email server in part so I could implement a spam filtering strategy that had access to my outgoing mail. Anything received from someone I sent mail to is assumed non-spam, and that is the anchor for the rest of the filtering.
I was planning to build a pretty sophisticated bayesian filter on top of that but it turned out to be unnecessary. The strategy I ended up with, which has served me well for many years now, is:
1. Block a small number of extremely spammy TLDs. There are about a dozen of these, including .biz, .casa, etc.
2. Anything received from someone I've sent mail to, or a sender I've previously marked as good, is assumed good.
3. Anything received from an address whose name contains an English word on a relatively short list of spammy words (discount, offer, etc.) is assumed spam.
4. Anything received from an address that I have received email from before and not marked as good is assumed spam.
That leaves emails from addresses that are sending me email for the first time. These are overwhelmingly spam, but after the four filters above there are few enough of these that I just scan them manually once a day or so.
The real key here is treating the first email from any given address as essentially a "contact request" with the default action being "deny in the future". That just turns out empirically to work incredibly well. More than 90% of my spam is from repeat offenders.