Paul Graham provides answer to spam emails (2002)
infoworld.com
infoworld.com
http://www.hashcash.org/papers/hashcash.pdf
Hashcash requires the sender to expend a small quantity of computational work, and attach a proof of this to the email before a recipient even opens it. The underlying assumption is that the burden would be insignificant for a real email sender, but onerous for a spammer.
The approach inspired the powerful anti-spam system at the center of Bitcoin.
It might be interesting to catalog all of the most innovative early approaches to combating spam, and the unrelated technologies that later arose from them.
Put another way, this is the unfortunate flip side of free unlimited talk and text.
[edit: I didn't realize the original article was from 2002. I agree the article is a bit obsolete at that point.]
No it doesn’t, and pg explains why in his essay. (Don’t know if the article states this too as since I’ve already read the essay before I didn’t bother to read a summarizing article about it. The essay is really excellent though.)
Quote from the essay:
> I'm more hopeful about Bayesian filters, because they evolve with the spam. So as spammers start using "c0ck" instead of "cock" to evade simple-minded spam filters based on individual words, Bayesian filters automatically notice. Indeed, "c0ck" is far more damning evidence than "cock", and Bayesian filters know precisely how much more.
> [...]
> To beat Bayesian filters, it would not be enough for spammers to make their emails unique or to stop using individual naughty words. They'd have to make their mails indistinguishable from your ordinary mail. And this I think would severely constrain them. Spam is mostly sales pitches, so unless your regular mail is all sales pitches, spams will inevitably have a different character. And the spammers would also, of course, have to change (and keep changing) their whole infrastructure, because otherwise the headers would look as bad to the Bayesian filters as ever, no matter what they did to the message body. I don't know enough about the infrastructure that spammers use to know how hard it would be to make the headers look innocent, but my guess is that it would be even harder than making the message look innocent.
http://www.paulgraham.com/spam.html
And as for your point about text in image I don’t know of any email client today that defaults to showing images from unkown senders.
I receive a lot of spam and it is all very distinct in nature and the Bayesian approach is still the way to go for fighting it I think.
Cool circa-2002 game for Unix nerds: you have 60 seconds to telnet to port 25 and generate the highest spamassassin score.
Paul Graham however definitely championed and popularized the idea of Bayesian filtering.
http://www.catb.org/~esr/bogofilter/bogofilter.html
https://en.m.wikipedia.org/wiki/Apache_SpamAssassin
Spamassassin's Bayesian classifier credits Graham.
https://metacpan.org/release/Mail-SpamAssassin/source/lib/Ma...
Also, Yahoo! Mail implemented at least partly Bayesian based filtering not long after the original PG essay.
I'm guessing GMail is using some kind of statistical filter for not only spam, but general categorization, though as others have pointed out GM is pretty aggressive in filtering. I think that GM might be using collaborative data, aggregating across user behavior and I believe many users deal with signups they're tired of by marking as spam, sharing that behavior across the userbase would cause things that you willingly signed up for to be filtered out.
And I'm pretty sure that some years ago gmail was much better at this.
I only use a bayes filter (I use bogofilter) and out of some 15k spam e-mails and 150k ham e-mails I received in the last 9 months or so, I've seen like ~10 false negatives and around the same number of false positives. That's way better than gmail can ever dream of, where I always had issues with automated messages ending up in spam. With bogofilter, the training is always explicit, and otherwise the model doesn't change, so I can be sure that when my automated messages pass, they will always pass, and don't randomly stop passing "just because".
I started with a big archive (around 100k msgs each) of SPAM and HAM for the initial training. I learned from the early days of using e-mail to archive my spam, for the eventual training, since my first Linux job involved setting up a DSPAM installation.
Another issue is the carbon/energetic cost of spamming. I'm no expert, but given the probably high figures, should this issue be escalated to a higher level, e.g. an international agreement along with heavy penalty if caught running a larger spamming farm? Compared to the difficult issues of drug trafficking etc. I don't see how there could not be a consensus for spam.
I only filter them into a pre-spam folder and batch spamflag later, that way I don't miss anything hit by a false positive (e.g. a Signal v. Noise digest had the phrase "Why the hurry?" at one point)
Using this method I don't think I had a single black friday email hit my inbox last year :)
Oh and of course "Greeting(s?) of the day" has never been a legitimate email. You can safely auto-spam that phrase.
But can AI stop politicians like Cuomo from censoring unpalatable speech (Usenet/Reddit/Voat) to death?