Honestly, why would you spider and scrape sites when you can just grab a leaked dataset?
It would be slightly more interesting to put the canaries in a more commonly scanned location, like a plain-text bio field on a popular site (like this one... or FB... or LinkedIn...).
Since it's far safer to buy email lists from some broker, you tend to get a lot more spam for signing up for things than posting your email publicly.
But he chose to mention that all spam were marked as such and that only 1 or 2 get through. Readers will naturally be curious what methodology and tools are in use.
Or what I'm saying is, if the SMTP server blocks by IP first before determining what mail is being delivered then the actual rate of potential spam to any particular email address is not being discovered.
[1]: https://www.fastmail.help/hc/en-us/articles/360060591393
I honestly get more at my work email, which has never been posted anywhere... I wonder if spammers have started to assume the easy to get email addresses are suspicious or not valuable for various reasons.
Yes. The whole obfuscate my email thing is silly. I have the same email since the mid 90s, I post it without any care to wherever.
/s
Today I basically don't get any spam. I don't keep precise stats on the spam that makes it to my inbox but it's something like 2 or 3 per month. Compared to a few thousand per day twenty+ years ago!
I want to highlight again that it is the exact same email address I use today.
I block misconfigured connections at the SMTP level in postfix (I run my own mail server), that takes care of pretty much all of it.
After that I run bayesian filtering (spamprobe) to score spam, but there isn't much to catch. Over the last 30 days spamprobe has caught 37 messages, all true positives.
Unlike gmail, I also don't get any false positives (zero so far this year).
I know how bad it was couple decades ago and a lot of people are still traumatized by those days and act like it's still 2001, but the reality in 2023 is that spam is a very minor issue.
This is mentioned, but a little hidden (in the description for URL encoding):
> This is based on a small sample size: just six bots that were observed over a one-year period.
I think the number of spam emails would be a bad measure, since a single scan could result in many hundreds of spam emails.