Email Validation Logic is Wrong (2021)
netmeister.org
netmeister.org
I am going to put a (sane and reasonable) maximum length for email addresses on my postgres db column so it's not stored in TOAST; I don't care if you're technically allowed 256 bytes before the @ sign: I'm not wasting space in my database so you can feel special.
I do validate that the domain not only resolves but that it also has a valid MX record because this benefits ~100% of my users and protects them from their typos. It's not my problem if you want to register for my service before you set up your nameservers - its yours.
etc, etc, etc.
It's a fun thought experiment but there's no way in hell I'm going to advise any company to actually allow any of these and open themselves up to a can of worms.
A more sane list of rules would be things like "make sure you support - in the prefix or the domain" and others of its ilk. Also, I came across an email address that was created automatically from Active Directory to Azure with the apostrophe preserved (think "Sean O'Henry" turned into "sean.o'henry@example.com" and it blew my mind when that "just worked" in testing. These would be helpful rules because you'll actually encounter them in the real-world and not supporting them will genuinely inconvenience real people and not someone's PhD research.
In absence of an MX record MTA will use an A record and a small but non-zero fraction of domains really has an A record pointing to a mail server.
It's not a huge deal, but I would have preferred to continue without one.
At the time I used "mailcheck": https://github.com/mailcheck/mailcheck
There appears to be a more modern implementation here: https://github.com/ZooTools/email-spell-checker
It reduced the amount of badly entered emails more than any other approach I tried.
/^[a-zA-Z0-9.!#$%&'*+\/=?^_`{|}~-]+@[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?)*$/
If you want to be a little more precise, the '.' characters in the local part cannot be consecutive nor occur at the start or end of the string. Also, the syntax does not account for EAI (add any non-C1 control non-ASCII Unicode character to the list of allowable characters), nor does it account for IDN (complicated, but it's IDN, there's libraries for it).Unfortunately, email is in a weird space where there's a lot of clearly incorrect validators out there that fail on common stuff (e.g., dashes in domain names, weird TLDs, etc.), but trying to find a good one is hard because many people want to "help" by flexing how good they are at covering really, really weird cases that you honestly shouldn't want to support, like relay routing, or RFC 822 comments in email addresses (which aren't actually part of the email address!).
[1] https://html.spec.whatwg.org/multipage/input.html#email-stat...
<local portion>@<destination portion> is how I seem to remember "email address validation".
The '@' is the only reserved visible character.
The local portion is ONLY to be interpreted by the destination host.
Destination portion does not have to follow any pattern except to specify something (such as a path or resource name) findable by the mail server.
Of course, if you specify a destination only understood by your local email server as a destination in your LAN/WAN/whatever when talking to somebody/something outside that, they won't be able to reach you.
Addresses tend to look a lot more orderly than they used to, but I think this would be a "valid" address (within a given system):
12855977.233~foo+esr@bar!baz.crunchly!test%test#test\\splat\3
There is exactly one way to validate an email address, and that is to send it an email.
No matter what regex you use, you will be wrong. Sure, use the 800 line regex to do a pre-validation and ask the user "are you sure this is correct?" if it fails, but let them use it anyway if they click yes, because it might just be valid.
Then send it an email and if it bounces or they don't click the verification link, move on and delete it.
Also, those services don't ding you if you get a single bounce from a new email. They only start penalizing you if you repeatedly send to the same email and it bounces. You could always put checks in place to store previously failed email addresses.
And if your business is generating a bunch of take signups, it sounds like you need to put some rate limits on your sign up page.
And of course that's not deliverable, you can't use a : outside of a quoted part. :) But that is actually a perfect example of why you can't use a regex. Maybe they modified their mail server to accept it.
No one knows that for sure, Amazon doesn't publish how they determine what counts against you.
But from everything I've seen anecdotally, it does not. If you review the bounces and suppress the emails, it works just fine.
Getting booted for bounces really only happens to high volume senders who are using scraped lists to send cold emails. And it's mostly the complaints that ruin your reputation, not the bounces.
12 is almost certainly a typo
13 is probably a spammer
Also <input type=email> will reject several of these as well, so if you have such an edge-case email in your database, you'll run into issues with html forms.
How many legitimate users will use capital letters, which I had to add to an in-production email validator on an app with millions of MAUs?
Just send a confirmation email to whatever string they provide you.
You might want to first check that it is not a local address, and to reject it if it is (unless it is being used for an internal registration service). You should also reject it if the destination is unreachable (in which case it is not possible to send a confirmation email, anyways).
Millions?
If you are in the former category, then yes, follow the spec to the letter. If you're in the latter, then screw the precise guidelines of the spec and reject emails that are very unlikely to be valid: no quoted localparts, no IP address literals. In addition, go ahead and say that email is case-insensitive (more precisely, case-preserving).
As I say in my sibling comment, there are generally two purposes you might have with an email address. If your primary purpose is actually handling email, then that is when you need to be perfectly precise for email. But if your purpose is in using the email address as some sort of "universal internet ID"--this is true for the vast majority of uses of email addresses--then restricting the set of potentially valid email addresses is not only valid but a good idea.
Quoted string local-parts and IP address literals in lieu of domains are things which are generally broken by middleware software that deals with email addresses anyways, to such a degree that there is no way anyone who has such an email address is using it as anything other than a "do you actually support this" email address--it can't be a valid unique-ish identifier for them. Additionally, allowing them to creep into parts of your system may break assumptions of other databases, which increases the chance of weird, deep failures that may cause security vulnerabilities. That alone is a pretty good reason to block them.
Sure there is. It may signal more edge-case behavior to come. If you're building a developer tool, you should probably support it. But if you're building a consumer product, that may not be a customer you want.
I pass no judgment here. I only mean to admit my incorrect assumption.
What?
Some customers cost more to serve than they will ever make you. Still serving them may make sense. But often it doesn’t. If you’re bootstrapping a gummy bear start-up, a customer who pings you weekly with edge-case support tickets is unlikely worth the engineering effort to appease. If you’re building an AWS competitor, on the other hand, you may want to hire them.
For the average company trying to navigate these sorts of decisions, supporting users with emails like "a$123@1-2-3.12397.museum" probably isn't high on their list of priorities, regardless of whether it's "valid" or not.
Or if you REALLY want a regex for an input, /[^\s]+@[^\s]+/ seems as sure as you can get without excluding a valid address somewhere. At least one non-whitespace character followed by an @ followed by at least one non-whitespace character. Any more specific than that and it's dicey.
But also just send an email instead if you need to validate the email address.
I am also now using Fastmail addresses for anything I don't have some trust for and isn't a problem if lost. I no longer use my domain for things outside of this because it seems like an obvious step to simply link an unknown domain together, even if a laughable amount of companies don't.
If that email address validation is "wrong" anywhere in the sense of conflicting with real-world usage, it's probably mostly in not allowing UTF-8 characters in the local or domain parts. I'm guessing that for that I'll need to dig into RFC 6531 and 5890.
[1]: https://docs.racket-lang.org/splitflap/mod-constructs.html#%...
e.g. x@familyname.com is rejected, but x+whydidyourejectthis@familyname.com works