Email obfuscation: What still works in 2023?
spencermortensen.com
spencermortensen.com
Honestly, why would you spider and scrape sites when you can just grab a leaked dataset?
It would be slightly more interesting to put the canaries in a more commonly scanned location, like a plain-text bio field on a popular site (like this one... or FB... or LinkedIn...).
Since it's far safer to buy email lists from some broker, you tend to get a lot more spam for signing up for things than posting your email publicly.
But he chose to mention that all spam were marked as such and that only 1 or 2 get through. Readers will naturally be curious what methodology and tools are in use.
Or what I'm saying is, if the SMTP server blocks by IP first before determining what mail is being delivered then the actual rate of potential spam to any particular email address is not being discovered.
[1]: https://www.fastmail.help/hc/en-us/articles/360060591393
I honestly get more at my work email, which has never been posted anywhere... I wonder if spammers have started to assume the easy to get email addresses are suspicious or not valuable for various reasons.
Yes. The whole obfuscate my email thing is silly. I have the same email since the mid 90s, I post it without any care to wherever.
/s
Today I basically don't get any spam. I don't keep precise stats on the spam that makes it to my inbox but it's something like 2 or 3 per month. Compared to a few thousand per day twenty+ years ago!
I want to highlight again that it is the exact same email address I use today.
I block misconfigured connections at the SMTP level in postfix (I run my own mail server), that takes care of pretty much all of it.
After that I run bayesian filtering (spamprobe) to score spam, but there isn't much to catch. Over the last 30 days spamprobe has caught 37 messages, all true positives.
Unlike gmail, I also don't get any false positives (zero so far this year).
I know how bad it was couple decades ago and a lot of people are still traumatized by those days and act like it's still 2001, but the reality in 2023 is that spam is a very minor issue.
This is mentioned, but a little hidden (in the description for URL encoding):
> This is based on a small sample size: just six bots that were observed over a one-year period.
I think the number of spam emails would be a bad measure, since a single scan could result in many hundreds of spam emails.
If you're really worried about spam, I recommend just keeping a different email address for each separate purpose.
The benefit is that I know where someone got my email from, and I can then try to figure out whether the place has been compromised, or whether they're selling my email, etc. And I can just blacklist that particular address forever as well.
Previously, I just did whatever@mydomain.tld, but I've switched to something similar to blame.email [1].
This makes my emails look a little weirder, but it has stopped the weird looks I'd get when walking into a physical place, like my doctor, and telling them "Yeah, email me at <doctor's name>@<first><last>.com".
It also makes it less obvious that its effectively a throwaway email, particularly combined with my domain; it looks fitting. And since each address is salted and hashed, it pretty much eliminates the risk of someone successfullying trying to phish me by sending me an email to something like `paypal@<first><last>.com`.
Lastly, on my HN profile and elsewhere, I've got my "email", but despite them being unique, I still don't want to have to rotate it if it gets picked up by a spambot, so I've tried to do some plaintext simple "obfuscation" like in the article.
I went for <address> ~АТ~ <domain>.<tld> -- with the "AT" being Cyrillic rather than Latin - I figure at least some will get tripped up by not being able to use purely English regex.
So far, I have yet to receive any spam with that strategy. Maybe I'm lucky or just not getting indexed, or maybe it's working a little.
Still torn about how to handle Git or copyright/license headers, though; those addresses need to last a long time, in case anyone needs to reach out and ask for re-licensing/etc, and I figure it'd be annoying doing different emails for each repo.
[1]: https://news.ycombinator.com/item?id=31820502 / https://blame.email/
Yeah, same. I store all the addresses in KeePass.
The main reason I don't just totally randomize them is just that there have been a few moments where I do have my salt somehow, but for whatever reason, it is either inconvenient or impossible to immediately open up the password manager and add a new entry.
In those moments, being able to deterministically generate the address and then add it at my leisure without having to double-check what I used is nice.
It also likely wouldn't happen to me, but should I ever somehow lose/lose access to both my old emails and my password manager, as long as I have my salt, I can still "remember" my email addresses for important services (e.g., PayPal or whatever) to re-generate the addresses and reset my passwords.
Whatever route you go, be it randomized addresses or hashed addresses, even though I think I am more vigilant and careful than most, it's still nice having an extra-layer to the catch-all that can't easily be targeted by someone malicious without first either somehow obtaining your salt, compromising the service, etc; it's handy being able to immediately filter and flag anything relating to my bank or whatever else if it isn't sent to the right address.
If Spam arrives, I can block that specific address and use different random characters to live in peace again.
Keeping the domain readable makes it easier to explain to people that they must’ve “lost” my email address somehow.
It depends on how off-guard I'm caught and how important it is to me. I usually have my phone, which has my KeePass and email salt inside, and I usually have at least enough battery to last a conversation, so it's rare that I can't generate the proper email address in <30 seconds in most scenarios.
But yes, having like, 5e5ee440@<domain>.<tld>, has definitely resulted in a few "can you repeat that?" or "just to confirm?" moments (especially over the phone since audio quality often sucks). That said, for whatever reason, people are still seemingly less surprised by "5e5ee440@" vs. "<your place of work>@".
On the rarer occasions where I don't at least have my phone or something, if it's something I know I can update later, I'll tell them whatever is easy to input and remember; I separate emails to unknown recipient addresses, but I don't completely reject them outright, so it's not usually an issue pulling out the confirmation email or whatever later, and then updating the address.
However, if I don't have my phone, and I don't know how easily I could update the email, then it depends more. For example, my doctor wants an email on file for whatever record-keeping reason and for sending appointment confirmations and such. In that scenario, I don't know that I'd necessarily be able to easily change it without going in/calling them.
The first time, I did give them <doctor's practice>@domain.tld, because I figure, despite being an important email, it's unlikely that it'd get abused; if someone somehow knows my GP's full name and practice, and is using it maliciously, I've probably got bigger worries than getting a phishing email sent to it or whatever else.
The second time, though, I just asked her to email me the contact update form and told her I'd send it back with the proper email inside.
> I wonder if there could be an easy “word-sounding” generator that could be integrated into something to manage emails?
I figure you could do something similar to like the horse-battery-stapler XKCD meme or bitcoin wallet seed phrases, if you wanted to avoid the "sorry, can you repeat that?" moments.
But it might be slightly more annoying to deterministically generate those, if you care about that aspect, compared to simply salt+hash & truncate. If you find a good method, let me know, though.
I have an unobfuscated mailto: link on my blog and I receive barely any spam to my public email address, ~@eligrey.com
He’s saying this works bc scrapers might not match a left hand side of an @ that contains no alphanumerics
"Surprisingly, the unprotected email address appears to have blocked a spam email. Either that message wasn’t received, or an extra message was sent to one of the protected email addresses."
https://news.ycombinator.com/item?id=38150096
Some comments claim that it can break some of the more complex techniques presented in the article. I've tried it a few times myself with varying results that tend towards not working.
After all of that I learned that privacy protection is not available for .us domains. No wonder they're so cheap.
A long way of saying, if you sign up for a .us domain, definitely do not give them your real information.
<a href="#" class="cryptedmail" data-name="david" data-domain="davidlane" data-tld="io" onclick="window.location.href = 'mailto:' + this.dataset.name + '@' + this.dataset.domain + '.' + this.dataset.tld; return false;"></a>
Paired with this CSS:
.cryptedmail:after { content: attr(data-name) "@" attr(data-domain) "." attr(data-tld); }
Which results in a clickable link that opens whatever is set up to handle mailto: links (1).
I read the explanation, but it just sounds like something is a bit off in the methodology and metrics. Or at least in my understanding of them :)
The problem is that humans who want to cut and paste your email into a different client have to retype it, which is annoying and error prone.
The author has warnings for the last three version, because of usability. I see lots of red flags also for the other versions in terms of accessibility.
Or you could just use a spam filter.
Does not matter, it works then!! But not because of technical cleverness.
No idea if what I am saying is true.
Sites that require email for logins get my common email that I do not check unless required to reset a password or links for authentication.
Long form communication is done via texting or shared chat.
We are living in strange times where the digital stuff just does not exist unless it is consumed in the present.
I have loving emails written to me in pen and email. I read neither yet for some reason the penned documents are held onto with an emotional tie.
My father, who use to painstakingly capture photos, create slide shows using a rotating projector, laments the days where people would review photos as a time of bonding. That no longer exists and now, even with 500x more photos I still long for the days of having the context of a few photos that are forever lost. GPS and timestamps are helpful I guess but I will never know the reason for that captured moment.
That's a nightmare.
In reality, long form just doesn't exist in chat/text. For me, long form is at least 500 words. Years go by before I get "long form" text/chat of that order.
To be fair, most of it is directly from LinkedIn, but still...
in 2023 email spam seems like mostly a solved problem - i very rarely get any that actually makes it through to my inbox. trying to solve the spam problem by protecting your email address from becoming public might have seemed like a valid strategy 20 years ago, but we have better tools now.
I have gmail and O365 and some other web hosting provider. My work email is unknown, probably filtered 5x, plus on O365 so limited to known contacts.
I do this because sometimes in companies, people will put DB dumps in the wrong environment that has an actual SMTP going to WAN, then shit happens. I also make sure the environments have a dummy smtp or mailcatcher, but it's better to be safe than sorry.
See RFC2606: https://www.rfc-editor.org/rfc/rfc2606.html
Cart abandonment, tips on how to stay safe online, requests to go paperless, newsletters, claim your free subscription that came with your purchase, thank you for your purchase (separate to the order confirmation), requests to leave a review, continue your application, offering support, "you haven't logged in for a while", new login from $device, thank you for completing step X ... it's endless.
An endless torrent of excuses to get the their company name in front of your eyeballs.
That's a security precaution, not spam.
I filter these as spam.
I get so many, I tune out. I bet I'm not alone
You really have 2 options from a UX standpoint. Either allow the login and notify you, which gives you less friction in your experience with the application if it really is you.
Or they can stop you right on the login screen and send you an email with a code or a link to click before you go further. It’s more secure, but it adds friction on the more common case that it really is you.
Frustrating! I never opted in!
Also, I removed my public open source projects from GitHub after Microsoft bought them then removed them from GitLab too when the AI warnings about stealing copyrighted code started to sound more real.