Complexities of e-mail validation logic
netmeister.org
netmeister.org
Developer time is precious at a startup and supporting <RFC>fan 69™@root while still denying b ob@gmailcom is very, very far down the list of things to do.
In summary: I don't suggest doing 'perfect' email validation to RFC spec. You will save money/devtime and make more of your users happy by not doing it.
EDIT: moreover, a service is perfectly within their rights to _internally_ store my email as `myname@gmail.com` if they want - but they should still accept `myname+yoursite@gmail.com` as the identifier used to login with.
That whole setup for tidiness is broken the moment a desired website does not accept an alias in your address, of course.
The '+' alias feature is a fairly common configuration, though, so for source labels it's better to either treat all unlabeled messages as spam or else use a more opaque labeling scheme (unique-hash@example.com) which doesn't hint at an alternative untracked email address.
Why? They don't care about protecting the business interests of wherever they got that address from, and it's not like stripping the plus off will meaningfully increase the success rate.
Worse, I know at least 5 or 6 people personally, which do catch all. It seems like a very poor method to reliably catch spammers.
Additionally, sending a test email like that might also get the sender placed on a black list for triggering a spam trap inadvertently.
Do you know any site blocking domains with a catchall?
I go a little farther. I figure an attentive spammer might figure out that if I use amazon@johnsmith.net to sign up for Amazon, I may have exactly the scheme where *@johnsmith.net will work, so they can just add that to the spam list as a wildcard and pick a new address every time. So instead, I use john101@johnsmith.net, john102, john103, etc, to try and obscure my strategy and prolong the life of the domain forwarding.
At least a few years ago, I noticed a lot of spam to <random first name>@<my domain> -- i.e., completely made-up addresses that I had never used. Since messages sent to those addresses were guaranteed to be spam, I started treating them as free training data for the spam filter.
I don't know if this still happens, though, because I haven't looked.
It seems like an obvious thing to try, but maybe not worth the effort of implementing it, given the high risk of false positives and the low % of people who actually do stuff like this (not to mention they're probably not people who click on ads anyway).
Also unless you're keeping a lookup table you're losing a great benefit of the wildcard. You can, and I have caught a few places, tell when a company sells your email. If I get an email from company XYZ to my email abc@example.com I know exactly who sold my email and to whom.
> unless you're keeping a lookup table you're losing a great benefit of the wildcard
That's true, I don't keep a lookup table per se, though I do have a deleted items folder that I could look back in. I'm not sure what I would do, though, if I knew what particular company sold my email address? Send them a nastygram they will just ignore? I just block the address and move on.
AFAIU, most buld spam is targeted on gullible or vulnerable people. The spam is often terrible on purpose.
Sophisticated or targeted attacks are a different category and they may be a good reason to prefer something non-guessable.
I have found this to work; I hardly receive any spam at all, and do not need any separate spam filter.
I've received spam emails to at least 70 different +addresses. It is absolutely useful for antispam.
Spammers don't care about the reputation of the company they bought or stole the data from.
Not all email providers support the + notion so you'd have to run domain lookup on some hard coded list
Also anyone with gmail address can also place dots almost anywhere into the local part, to create another unique address without using a + sign.
Contrary to popular belief, it is not a gmail feature.
I first heard of the + as destination filtering in the very early 90s at CMU where it was broadly used. Every single email address I've had since then has support the same (and notably, apart from a test account, I've never used gmail much, so that's not including gmail).
I always have to tell people, in real life, "it's .co, not .com," just in case - humans do this too.
For the initial email input, your logic works fine. Once it is applied downstream in a process, it begins to get messy. Someone might do an incorrect email validation that happens to block emails that you have already accepted or which you are importing from a valid source. Someone has already given the example of a login field not allowing them to use the email they signed up with. If such upgrades occur later in a projects life cycle, not only might you have to spend developer's time, you may also have a production outage.
Personally, I suggest using some, even if imperfect, validation when gathering the email initially (for the reasons you point out) and then not validating that information any further.
Sometimes disabling Javascript will fix it, sometimes not. I occasionally have resort to using "I forgot my password" until I figure out what the actual underlying requirements of the passwords are.
I’m always suspicious when sites cap passwords at < 32 characters, that almost always means it’s being stored in a reversesble format someplace - maybe encrypted, maybe obfuscated, or maybe not either (banks).
The sites I really trust don’t care how long your password is because their hash size is fixed. The only real length consideration might be that if a bunch of people send obnoxiously long passwords at the same time and they are using bcryprt or scrypt it might stress the server’s cpu, so they might put an upper limit to prevent that.
With precious dev time, you can do better by doing less.
I just make folks email me first.
If you are in the former category, then yes, follow the spec to the letter. If you're in the latter, then screw the precise guidelines of the spec and reject emails that are very unlikely to be valid: no quoted localparts, no IP address literals. In addition, go ahead and say that email is case-insensitive (more precisely, case-preserving).
The hard part is if you're writing an email client, because you're basically forced to have your hands in both pies.
that works as long as <RFC>fan 69™@root does not write articles for ZDNet
Don't get clever, just follow the spec.
That's surprising to me because there is nothing particularly weird about your email address. What exactly do they complain about?
I'd suggest being clever is wasting countless hours to handle your edge case. Or writing your own email validation in the first place.
Isn't email validation a solved problem in that there are services or ready software which provide RFC-compliant validation? If some company is wasting countless hours to do something because of Not Invented Here syndrome, isn't that the same as some company deciding to write cryptography algorithms on their own and reaping what they sow?
My email is refused by 0% of ecommerce shops... because I just have a normal email.
Don't be clever, pick a better email.
My email is just "me@<my-last-name>.al"[1] which is just a tiny bit "unusual" - and over the years it got refused by a couple stores because of TLD. And Albania is not Cocos Islands, they're surely not popular with spammers.
If a store believes there's only ".com" gTLD and nothing else (this had really happened to me, some galaxy-brain made a form with a hardcoded ".com" suffix; not even ".net" or ".org" were accepted, unfortunately I don't remember the site) - well, fuck that store, their loss not mine. Worst case, if I really want something they sell, I'll give them a throwaway email - which will contribute to their mail bounces after some time.
__________
[1] ".al" is a ccTLD for Albania which is not a country of my citizenship or residence. I've picked the domain name as hack - because my first name is Aleksei and my first and middle names form "A.L." initials as well. That, and because all relevant .name domains were already taken.
Think about it this way: either you can get some big brand .com email with no special username and never have an issue, or you can flail around 5% of the time and yell at the clouds.
Should everyone accept your email? Of course! I'm just saying you live in real life, and in real life people suck at building email forms. The problems you run into are on you.
No, the problems they run into are caused by (at best) mediocre developers. They’re entirely to blame. We have specs and standards for a reason.
Instead you can just get a big name .com email and call it a day. Live your life without trying to make some statement about email standards.
No, because it's not 1993, but I absolutely do use the contact forms or bug reporter for any website that doesn't accept my email. Most of them fix it, because it's objectively a bug caused by their non-compliant code.
However. I completely disagree with the conclusion “The problems you run into are on you.”
I didn’t create the problem by having the audacity to be from a different country.
Get a big brand .com email and you'll never run into an issue.
To adapt from a famous quote: "all email validation logics are wrong, but some of them are useful" ;)
When I was validating myself for Amazon Prime Student, I literally had Amazon refuse to accept my student email in the form first.m.last@myschool.edu because there were two '.'s in the mailbox portion. I had to send an email to support and it was eventually dutifully fixed.
And that's not an uncommon format for, you know, school emails. And that's an Amazon engineer who should have known.
I imagine there's developers who think "domain.tld" is the only thing valid to put in the domain portion, and that's going to fail with "domain.co.uk", or uncommon TLDs, or other perfectly valid constructs. And sure "it's only x% of the users" but it's a pain in the ass if you're that user. You need to be reasonably permissive.
(but on the other hand "myname@..." is not valid either, and that will fail and cost you money as well... hence leading us back to 'just follow the spec')
Furthermore, for a user, it is trivial to get another email address if the one they have causes issues, so it is not really an accessibility issue either.
He made a very convincing argument that while an IP address is technically a valid domain, but how many legitimate users were seriously using an IP address as their email domain? (zero)
One of our testers found XSS with email injection (RFC compkiant validation passed) in our website.
And we are an e-mail company and should now better :D
Never trust user input!
But the way to prevent injection attacks is not to disallow or sanitize input, it is to escape correctly when interpolating strings in other languages.
No matter what you will constantly be getting addresses that conform to the spec but cannot actually receive mail.
e.g. It drives me BONKERS how many systems absolutely reject my single-letter email (~"N@domain.com"), which I created specifically to make it easy and safe to type on mobile devices etc. Others will reject the "+" sign, or underscore, or dot/period, or (brilliantly) two periods or underscors, etc etc etc :=/
If enough people wrote in about not accepting one letter email addresses, then they would likely update the validation.
But if customer service tells users to use another email address in that scenario and the customer does that, then it might not be worth the effort to fix it.
Better to warm an email doesn’t look right, but let them continue if they want to
When I want to check postage or whatever, and they require an email address, a@b.c typically doesn't work, but no@mail.com does.
Writing code that doesn't need to exist that has a possible failure mode of not letting someone sign up at all is just a bad decision. If you're going to write that code, either go through the effort of getting it completely right or soften the failure mode. If you really think the user is somehow mistyping their email with special characters or an unusual TLD, then you could show them a non-blocking warning message.
(Actually RFC5322 already deprecates some syntaxes. For example, "John Hacker"."Ph.D., Esq."@example.com is a deprecated syntax (obs-local-part), because it contains multiple quoted-string components separated by dots.)
> or (brilliantly) two periods or underscors
Two or more consecutive periods is actually disallowed by RFC5322 (unless quoted). foo.bar.baz@example.com is a valid address, foo..bar@example.com is not. ("foo..bar"@example.com is however)
Your other reasons for breaking interop between systems/languages are just whimsical and invalid. :)
Fundamentally, the problem is that if you’re trying to validate an e-mail address as being correct and you’re not sending an actual e-mail message to that address, then you’re doing it wrong.
We learned this lesson back in 1995, people.
And this wasn’t a new lesson then. But at least we were smart enough to listen to the people who had learned that lesson before us.
It is now over 25+years later, and I’m sad to see that many people seem to be bound and determined to force themselves to re-learn that lesson the hard way.
/^[a-zA-Z0-9.!#$%&'*+\/=?^_`{|}~-]+@[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?)*$/
Remember that the web is a platform that lives and breathes this stuff. A lot of thought went into this grammar for valid email addresses. This is a good way of filtering out obviously bad stuff while allowing all realistic and sane inputs.One part of all this that I’m not aware of the situation around is “8. You can put emojis in the local part.” The HTML spec’s validator is all ASCII. It does remind you to punycode the domain labels, but makes no mention of internationalised local parts, and I’ve never learned about non-ASCII local parts or how well they’re supported. I gather they may require the sender to be capable as well as the receiver, whereas internationalised domain names were made compatible with all systems via punycode.
Isn't that a way of saying "while disallowing perfectly valid options"?
There's a distinction to be drawn between the requirements of the actual MTA/MUA/MSA layers and user applications built on top of them. For the latter, considering emails to be invalid if they contain IP literals or quoted localparts is going to be more helpful than harmful (there's less scope for vulnerabilities in doing so). It's just like assuming email addresses are case insensitive: it's inappropriate if you're an MTA, but for everybody else, go ahead and assume they are.
A-ha, but here you're wrong because you've excluded IDNs. This is really why you should not try to be clever.
So what makes it the best?
[Edit: it also assumes you've already parsed out the "real" address from the rest of the text field, which to me makes it a half-validator at most.]
But it would seem to be the best for general-purpose web use, e.g. signing up for a newsletter with an e-mail address that's pretty much guaranteed not to break anything.
Instead of being conservative in output, it's intentionally being conservative in input.
> it also assumes you've already parsed out the "real" address from the rest of the text field, which to me makes it a half-validator at most
I’m confused. The explicit purpose of this stuff is to validate an email address. Not to extract an email address from a freeform text field, which I think is what you’re talking about. Deciding how to do that is a whole ’nother can of worms.
Unless you're developing an app for an intranet, that's not a concern for most people.
This is already a field where there is a lot of misinformation flying around, and a page that merely regurgitates all of that misinformation without the perspicacity to realize that its purported information is internally incoherent is not helpful.
More importantly, what problem is this even trying to solve? Someone accidentally typing a 300 character domain? If they are intentionally feeding you gibberish they’ll just give you more realistic looking gibberish.
I’m absolutely certain that the 63-character limit for domain labels is never going to change, because it’s hardcoded in enormous amounts of software and hardware, and there’s no even vaguely compelling reason to even attempt to change it. But if such a thing did change, then you’d just add this to the extremely long list of things that needed to be updated.
People who thought TLDs would only ever be up to three characters long were simply wrong from the very start because they didn’t understand what they were dealing with. (As a simple example, .arpa was there from the start.) Understand that this wasn’t a matter of anything changing, it was that some people misunderstood and thought that a convention they observed was in fact a rule.
The problem this sort of validation is solving is weeding out things that are definitely not going to work, as soon as possible, because it’s good to point out problems to users as soon as possible, rather than having something silently fail or only notifying the user about it much later. Syntactic validation isn’t the be-all and end-all of accepting email addresses, but it’s definitely still worthwhile, even though you should generally do other validation based on DNS lookups and/or sending actual emails as well.
Seeing as the web has long supported Unicode, where are e-mail addresses currently at in that evolution?
Are full Unicode e-mail addresses something that is decently supported today, or still largely theoretical? Is this regex sufficient? What kind of e-mail addresses do people in China most commonly use, for instance?
Baby shoes because of anglosphere programmers that can't fathom people wanting to use their own alphabets and thus forget to support it.
Clearly "anglosphere programmers" fathom it every day when they use UTF-8 almost universally in webpages. Also, you know, things like emoji are pretty popular in the "anglosphere" as well.
It's obvious that the real reason is an ancient e-mail RFC, and that while upgrading webpages to UTF-8 was relatively easy, in that it only needs 2 parties to support it -- the browser and the server -- upgrading e-mail is almost infinitely more complicated, because you have to wait for virtually all email code in the world to be upgraded, since an e-mail address is pretty useless if it doesn't work everywhere.
It other words, it's a coordination problem. Not an ignorance problem.
And unfortunately, Punycode [1] doesn't seem to be a particularly viable stepping-stone/compatibility solution here. E.g. if a user tries to use ドメイン名例@example.com and it fails, asking them to instead type in a seemingly-gibberish eckwd4c7cu47r2wf@example.com, where that could also conflict with a real e-mail address of that name.
At least four decades of mostly bad internationalization support it's no longer accusatory, it's empirical and quite generously worded.
For the local part, though, it does look like browsers have fallen down, though I’m not particularly familiar with the situation there. Testing it in Firefox to confirm, ascii@υνικοδε validates, but υνικοδε@ascii doesn’t. https://github.com/whatwg/html/issues/4562 seems to be where progress is made from time to time. As usual, it’s not as simple as we might hope.
/^.+@.+\..+$/
That is, "Some characters, an @, some more characters, and a period"I couldn't care less if users want to enter undeliverable email addresses, they won't get emails. All that regex is intended to achieve is ensuring that the user hasn't accidentally filled the wrong field (e.g. tried entering their phone number) or mistyped a punctuation mark (foo#bar.com, foo@bar,com)
Strictly speaking, it won't match some valid email addresses, such as IPV6 domains. But if I receive a support ticket complaining that we don't accept email addresses with IPv6 address domain, I'll reply advising that the customer should purchase a domain name or sign up to one of many free email services.
Unfortunately, many websites are configured to reject email addresses that contain a plus character. I've also encountered websites in the past that did accept the + character when creating the account where the email address serves as the user name, but then could not log in because their log in form rejected the + character in the user name.
> I also get a lot of email from idiots who don't know their own address
Holy crap there are a lot of them. I've got one bank sending me the dude's statements. He's also been on some interesting trips, seen all his hotel stays, etc.
The other finally figured it out but his wife still hasn't after more than a decade. It gets really old receiving reminders to service a vehicle I've never owned from a dealership 2000 miles away among other similar crap.
Sounds like this was in person at store though which is extra weird because seems unlikely that scammers would be trying to sign up en masse at a physical location (unlike if the form is connected to the internet)
Much more reliable than the + -thing, which breaks in the weirdest of places.
I had to switch my hosting provider at one point because they stopped supporting catch-all. I have no idea how many "addresses" I've used, since I don't create a specific email for each, so I had to get new hosting (note: this was over 10 years ago)
my_email+service@foo.bar thus became my_emailservice@foo.bar
I tried logging in, resetting passwords, nothing worked. I had to go to the authorities and make a written request to allow them to interrogate the database by the equivalent of my social security number, and that’s when we realized they just stripped the +.
The one exception is Craigslist; if I email someone with my normal email, I never get a response. I always use gmail for that.
I also use greg-*@domain instead of *@domain, since their docs claim that setting up *@domain tends to attract more spam.
It's not cheap from PM, and there are loads of hosting providers that will provide catch-all email for free with your hosting package (but with some usually pretty poor webmail client) or if you use a mail client it should work too.
I like having good webmail and mail app and other things so I pay, but there are plenty of good options available. Sadly self-hosting email server is not really an option for a variety of reasons, but you should easily be able to use catch-all e-mail addresses.
At some point, I need to migrate away from google and build out my own personal mail server.
It's a freemium model, but I've never needed anything in the paid tier
No idea what sort of security this is supposed to provide.
They probably see it as some sort of security / anti-spam mechanism.
My emails have a tendency to become spam filter bycatch, to the point that when I was job hunting last year I'd have to ring people after I sent them my resumes etc. to confirm they actually received my email.
And when I give people my email address, I usually have to assure them that steve@stevetech.xyz is a legitimate email address and not a joke (it's not actually steve, but you get the point).
Definitely would not recommend using it for your personal address.
If he said he used "joe@company.example.com", then it's possible he has a wildcard MX record for *.example.com, but that's not at all what he said, although perhaps it's what he meant.
Regardless, the question remains unanswered.
Samsung doesn't accept emails with "samsung" as prefix, so I have samsun@mydomain.com for them. I have no idea what's the logic behind.
This allows a person to use any damn thing they want as their email address, provided it works and they can get the email.
Also, humans make mistakes. You should detect spelling errors and typos then suggest corrections. [3]
[1] https://www.mailjet.com/blog/news/3-factors-that-impact-your...]
It's hard to be smart with something like names.
Always send the confirmation "did you sign up?" email. Always.
Obviously, this is a different scenario than your bank not accepting your valid (per RFC) email address. Which is why any sort of blanket advice is pretty dumb. Not that I care to aid spammers...
The other scenario might be a site that puts up a "paywall" type thing, where you are forced to enter an email address to gain quick access to something, but doesn't want to bother you with going and verifying an email (e.g. instant discounts, downloading a PDF, etc.). Or in-person email address collection when you buy something in a store. It's never a good idea to collect email addresses of people that have no desire to subscribe to your marketing.
In practice, you build your UI for the latter. You add captchas or other friction for the former.
- https://blog.jgc.org/2010/06/your-last-name-contains-invalid...
- https://haacked.com/archive/2007/08/21/i-knew-how-to-validat...
- https://fosdem.org/2018/schedule/event/email_address_quiz/
If I can send you an email and you can verify that you have access to that email, your email is "valid enough" for me.
Then, the validation is basically "is there an @ and after a dot in there?". I find that after that, every hour spent on improving the validation will just cause more emails falsely flagged as invalid, more support requests from the people who couldn't sign up with valid emails, it's code we need to maintain, anytime edits the validation logic risks breaking sign ups completely.
So with more "improvements" to the validation, you just cause more problems. Then why do it?
I hear the reputation arguments, but in practice, it never happened to any of the organizations I worked for.
What happens though very often is naive engineers trying to solve problems the business doesn't have with knowledge they lack...
premature implementation is the source of most evil. :-)
Send that address a confirmation email. Now you've got consensual opt-in and you've somewhat protected yourself from adding a wrong address to your recurring mailing list.
Prevent abuse with long (seconds) delays between submissions from the client. If the user thinks they did it right, they're waiting on their email inbox anyway; if they immediately realize they made a typo, it'll take 2-3s to fix.
The RFCs were written when manually (not from cron) sending email to another user on your local system as a thing that actually happened. I'm certain you actively want to avoid that now.
After doing your simple regex, the best move is to just send a verification email and wait for the user to click the link, if you really need to be sure.
"><script>alert("XSS");</script>@example.com
The oppression must end!At this point, not doing email verification should be considered a dark pattern because it causes so much trouble when people's email addresses are used without their permission.
I can confirm that this is a very stupid mistake to make. :-(
Supporting all of the rules outlined in the spec is probably a huge burden for maintainers of mail clients and servers. Obviously some parts of the spec are going to be omitted. It's hard to blame them for it, but the same person that rightfully skipped over implementing the routing thingy might've also wrongfully assumed there won't be a Japanese character in the address. And that's what's so bad.
You might introduce more issues in your system, by taking the full spec into consideration for your validation, instead of using the whatwg regex someone posted here.
Web application validation forms add a different layer to the standard and are sort of hard to tame; anyone can push together a few lines of PHP or Javascript code and conjure their own email address standard out of thin air.
If the standard fails to be used, the standard is defective.
I'd like to see more of "Patterns like [what you entered] are uncommon—are you sure?" instead of "Patterns like [what you entered] are not allowed—change it to proceed."
- An email address ending with ".invalid", unless invalid email addresses are supposed to be allowed (which in some cases is useful, but you can then disable sending email to such an address, using it only for identification). (I do use such an email address for identification on NNTP.)
- Email addresses without at least one at sign.
- Email addresses containing control characters (at least ASCII control characters).
- If the domain name does not resolve or resolves to a loopback address or LAN address (except for some specialized cases where such a thing is desirable). The same is true for literal IP addresses; if it is a loopback or LAN address then it should be disallowed, but otherwise it can be allowed.
The domain in question being david.kitchen, so an email may be email@david.kitchen
The issue I encounter more than any other is trivial: Most sites still have a tld validation that only accepts domains that end in net|com|org and some other small list of accepted suffixes such as co.uk
The list of TLDs is constantly expanding https://newgtlds.icann.org/en/program-status/sunrise-claims-... so even `[a-z0-9.-]+@[a-z0-9.-]+\.[a-z0-9]+` would be better than what I see in the wild.
The amount of times I've tried to sign up with my protonmail account to a service and it doesn't pass validation simply because it's a protonmail account (not a gmail, outlook, hotmail or aol apparently). makes me wish everyone did follow the RFC. I actually emailed a service one time, and they responded that it's due to protonmail usually being associated with shady stuff wtf.
The second. I had to implement an email validator at one of my previous jobs, and fell down the RFC rabbit hole. Not only did I have to follow the RFC as per my bosses request, but I also had make sure that Amazon SES allowed it. Came out of the office wanting to just walk out onto the road. The weird things that not only email servers allow, but also, what do email clients allow.
However, if you still want to validate an email address then use a library. All popular programming languages have email validation libraries. Yes, it's an extra dependency if it's not included in the std lib or the framework you use, but email validation is wrong in 99% of the cases, if you wrote it yourself.
If it is a string that has an @ sign, a dot and is at least six characters long, it's probably a valid email address.
a@b.cc
No need to go further than this. It's not worth the time.
dotless domains are going to be so rare in practice (unless your project has some niche use-case) that you can probably ignore them and call them invalid for the sake of simplicity.
I worked in e-mail security for quite a while. "Write an e-mail address parser" was my go-to technical interview question.
It was pretty easy to see if the candidate had ever given any real thought to e-mail (most had not); and you could also pick up a lot of signals about engineering style, for instance if they started with a regex (fewer did than I expected). And it was trivial to adjust the difficulty: if someone thought the question was easy and had a fast solution, you could just throw them a test-case like the ones in this article.
(Note: the actual title is "Your E-Mail Validation Logic is Wrong" -- and it's only about addresses, the author isn't implying that e-mail systems can't validate messages nor for that matter addresses.)
I think you’re missing some stuff in your regex.
A similar enough issue happens in coding interviews anyway. Sometimes the interviewee is aware of a library that essentially solves the problem for you. In those cases I give them some credit for knowing of it and then ask them to implement it anyway, as if the library didn't exist (because there are a large number of problems out there for which a solution doesn't yet exist, and when hiring a SWE you need to find someone who can write new solutions from scratch for those situations; whether a given toy interview problem is such a situation doesn't matter for the purpose of evaluating said skills).
Digression: I do miss the days when you could assume a candidate for a position at Aquatic Widgets Incorporated would know something about water, or about widgets, or at least would have looked up what an aquatic widget is before bothering to come in for an interview, but those days have long since departed the realm of Software Engineering as far as I can tell. Which may be a good thing from the engineers' point of view, I'm not sure.
IMO well functioning teams do consider the thoughts of their technical members when deciding which problems to solve.
The person "making the call" on whether having perfect email validation is worth solving may not have an appreciation of how difficult it actually is, so having a discussion with engineers on how much work/time it would take should play a big part in prioitising it.
Additionally, things like validating email on signup are mostly solved (albeit imperfectly) so one can and should use existing implementations and focus on building their product.
Password strength requirements and email validation are just like the database examples, and if a company doesn't let these technical questions be answered by the technical people, that's a bad sign.
While maybe the engineer won't actually make the call, the engineer should inform management's understanding of the costs of the approach and the efficacy of alternatives, and management should go along with that recommendation unless they have a good reason not to. Of course tone is important, someone saying "fuck no, I ain't doing that" likely indeed would be unpleasant to work with, but a respectful "I would recommend against doing that" is the sign of a confident and intelligent professional.
True, but as an engineer you do need to provide accurate feedback regarding "Hey, this is gonna work much of the time but email is hard, this is a complex problem. If we do this from scratch we're going to miss a lot of things potentially".
Not to say I wouldn't try just for the sake of working through it as an example / 'where would you start' discussion.
But if we're pretending this is a real world task I'd probably discuss how this is an endless / possibly ultimately futile time sink and there might be better options than starting at point A ;)
In a project we're doing something fancier: we check the result of sending mail and store it in the database record for the account (Mandrill notifies us on a webhook.) Then we might take actions for bouncing addresses. The actual impact on the project has been zero so far.
I'm sorry, but I just can't bring in Clyde's truck for the oil change, cause Clyde ain't me!
I also cannot attend Cassidy's parent teacher conference, apologies, I am not in Ohio.
And yet I've worked multiple places where product people asked for "simple email validation" on user signup. If they insist, I ask them to provide some actual test cases that they care about. Sometimes the product folks can be convinced to drop the validation requirement if they can be shown that anyone who can't sign up because their email address doesn't validate will simply move on and not sign up.
In the case where your product is B2B and all the employees of your customers are users (say an HR product), then the first time a VIP at an important customer complains, that's usually enough to convince your stakeholders to disable the email validation.
I need to use the latter ~5% of the time. Most often I take my business to someone else for the sake of principle.
https://www.youtube.com/watch?v=JENdgiAPD6c
The first 5 minutes are perl specific, but the rest is email and just hilarious.
Yes, I can craft garbage emails that pass this quite easily, but who cares? If I’m crafting fake emails I can make valid ones too. This rule ensures I typed the @ and the dot in my domain (we really don’t need to support dotless email domains and it’s better to catch “foo@gmailcom”) and it won’t reject all the weird random emails people might have.
More generally: If an edge case exists and has nothing to do with accessibility (it was caused by a user having a different workflow like needing a screen reader or being in a less developed part of the world with slow internet) then you should dismiss them and not make your code/life unnecessarily complicated.
Obsessing about regular expressions for these addresses is generally a waste of time except for maybe preventing a lot of failed attempts to send stuff to a clearly invalid email address. A simple string contains '@' is probably good enough for that. Worst case the email address does not work and you discard the entered information after some reasonable time frame. The user has the option to try again and do a better job of typing their email address.
Accept email from user.
Send email to that address with a link to verify.
Go/no-go test if link is clicked.
(If you're doing some fever garbage or otherwise trying to parse it, you're doing it wrong.)
* Validate the string contains "@" and a "." to the right of it.
* Validate common typos
* Validate disposable emails
* Validate MX records
* Validate SMTP server and mailbox
https://github.com/mfbx9da4/deep-email-validator
I don't have the time to keep it maintained but it works for the most part!
My only technical nit would be the statement “if there was an MX record”.
Many systems will fall back to an A record to attempt delivery in absence of a MX record.
If your goal is to catch typos you’re better off with very lax validation plus a library that suggests corrections like “gmail.com” for “gnail.com” (both of which are of course technically valid domains)
Frankly, I don’t give a damn about supporting some l33t h4x0r wanting to be clever about his e-mail address.
So no, my logic is not wrong. It just doesn’t care about all the weird edge cases.
It makes a lot of sense to have a good regex that validates it on the frontend and even tells the user, "hey, this doesn't look valid, are you sure?" but don't ever reject someone for failing your regex.
At the end of the day, you should only validate by having them click on a link in their email. Then you know the whole thing works.
If you're worried about someone hurting your reputation with a bunch of bad email addresses, you can certainly mitigate that by rate limiting your signup API or even deprioritizing sending confirmation emails to email addresses that don't look valid. But you still should eventually send an email to make sure it's valid and not reject it until it bounces back.
Turns out we do minimal validation (make sure there is a local and a domain, that there are not two periods next to each other, and a few other things) but what we really rely on is deliverability.
In other words, if your email needs to be verified, we'll try to send an email to the address you provide. If the link is clicked (or the code entered), that's good enough for us.
Applications using our service (we're an auth provider) can decide for themselves if they need email address validation. It's a boolean flag on the user object. If they do, they can use the functionality we provide to ensure it.
Unfortunately it's done via black box on their server(s), so it's not like I can even dig through the code and figure out what's going wrong.
- nslookup –type=mx email.com
- pick the highest priority MX server
- telnet mx1.email.com 25
- validate SMTP handshake
- Start a connection: EHLO email.com
- mail from:<sender@youremail.com>
- rcpt to:<recipient@email.com>
Obviously, this might be outside the capabilities of some hosts or users. There's a bunch of services that expose this workflow for you as an api. (https://trumail.io/, for example)
"The domain name does not need to resolve"
Also the mail server may be temporarily offline or unreachable.
Edit just to clarify:
KPI's such as bounce rates etc aren't a function of how mail is delivered (RFC5321). These are KPI's collected and collated by non-SMTP applications sitting on top of SMTP infrastructure monitoring bounces.
Nowhere in RFC5321 does it mention that a mail server should or must not delivery mail in respect of bounce rates. These are operator defined metrics outside of the scope of RFC5321, that may be aided by additional software or services such as spam detection.
My original reply arose because there are times when a receiving domain or destination email address can be "temporarily" unavailable. I pointed this out to demonstrate that services that pre-validate recipient addresses upon submission of a form don't take into account transient outages due to any number of valid factors.
SMTP was designed with this in mind, i.e. try to re-deliver up to some acceptable threshold and then at some point give up (the hard bounce which is the thing that should cause the "ding", especially if they keep retrying beyond "soft bounces").
Perhaps re-read "jusssi"'s comment then mine. I didn't assert that bounce tracking was for the convenience of marketers, or suggest it was mentioned in any way in the RFC's, they implicitly did and I wanted to point out the error in their understanding.
> SMTP seems a particularly bad example if you...etc
But the central theme of this whole HN discussion thread is about SMTP.
If you're interested, sections 6 of RFC5321[0] are where bounce messages are mentioned (just three times in the whole RFC - bouncing, bounced and bounce) with no reference to marketers. See also 6.1:
Some delivery failures after the message is accepted by SMTP will be unavoidable. For example, it may be impossible for the receiving SMTP server to validate all the delivery addresses in RCPT command(s) due to a "soft" domain system error, because the target is a mailing list (see earlier discussion of RCPT), or because the server is acting as a relay and has no immediate access to the delivering system.
Which brings us back to my original comment, far above, that services that check once if an email address is "valid" using trumail.io or whatever when upon form filling are flawed solutions.
Although annoyingly in the case of a soft bounce they apparently only retry for up 12 hours, which for a small mail server is very much on the too short side:
In the worst case my mail server/hoster goes down just as I go to sleep, in the morning I either don't notice it or can't do anything about it anyway as I need to go to work, at work I can't do anything about it either, and it's only when I get back home that I can stand up my emergency mail server if the outage hasn't resolved itself by that time, which means > 20 h of down time and therefore far exceeding Amazon's retry window. (The RFC5321 recommendation is to keep retrying for 4 to 5 days, which is much more amenable for that case)
In the twenty-two years that have followed, the only website that has had a problem with my email address is Chapters Indigo, which explicitly rejects it as invalid.
For email validation, keeping it simple is best.
This seems more like a bug than a feature. Maybe in 1983 the average email user knew what DNS was and could be expected to know one part of the email address would be case sensitive and the other not.
But email RFCs are probably like any other RFCs out there and specify existing behavior for the sake of interoperability.
Imagine the average non-technical person talking to some customer service agent on the phone and having to figure out if her email is JANEWATSON@gmail.com, JaneWatson@gmail.com, janewatson@gmail.com, or Janewatson@gmail.com. Could you imagine the horror and complete security nightmare of multiple people running around using the same gmail address with different case. Those Jane email addresses above would be four different people. We'd be receiving mail intended for other people all day long.
At a certain level of complexity it's easier to just let mistakes happen and provide correction tools if & when required
I remember getting our IPs blacklisted trying to programmatically ask email servers if the email address we were provided was real
Doesn't happen very often, though.
American Express and Walgreens don't let you set a whatever@whatever.email address because they check for a TLD known at the time of their app's validation code, or something.
I've run into 2, old-ish institutions that didn't quite work with my whatever@whatever.com and had to modify it slightly for them.
Also, anyone notice the OP posts the same couple of posts constantly?
https://www.youtube.com/watch?v=xxX81WmXjPg
(Loudness warning)
If confirmed, it's valid.
let's say I'm a dev at google, and I'm writing some aspect of gmail. Is !"£$%@gmail.com a valid email? No. So the word valid is clearly not being used correctly here.
At the core of smtp these emails are allowed, but in practice they almost never are, and so the opposite is true. All the cases he described are, in practice, invalid.
"Verify that you have addressed this message correctly. Check your SMTP server settings in Mail preferences and verify any advanced settings with your system administrator.
The server response was: The recipient address <'*+-/=?^_`{|}~#$@[ipv6:2001:470:30:84:e276:63ff:fe72:3900]> is not a valid RFC-5321 address. <...> - gsmtp"
Not even Google engineers can get it right. We are doomed.
If this squicks you for some reason--as maybe that format is non-obvious with respect to the lack of a need to escape @--give the user two boxes with a hardcoded @ between them and have them type the two parts separately: pre-parsed input need not ever be escaped, as you aren't going to parse it at all; no need to implement " dequoting.
All of these escaping rules are then to support embedding this identifier into SMTP. The rules for embedding the same identifier into MIME are different... and even more complex! In MIME they support random stuff like "comments" in the middle of the string... is that part of the email address identifier? No.
An email address simply is not defined by the format you use to send it as part of an SMTP command, nor is it defined by the format you use to send it as part of a MIME message header :/. Into is an identifier that exists separately from either of those two (different) protocols and one would expect any number of ways to escape that content.
To demonstrate how ridiculous this all is, imagine someone comes up with a JSON protocol for mail submission and then documents how email addresses now should use \u encoding and escape quotation marks... does that mean users should type that into your app? No.
Hell: your email address form is taking an email address and then sending it over HTTP... the escaping rules for HTML form fields are different still, yet no one is asking users to type HTML-escaped strings into other applications, right?
The core thing wrong then with your email validation is that you are simply validating the wrong thing: unless you are developing an SMTP server, the rules for how to escape and parse escaped email addresses in RFC5321 are irrelevant; and, likewise, unless you are developing a MIME parser, the rules for how to escape and parse escaped email addresses in RFC5322 are also irrelevant.
The only thing that matters from either of these specifications is the underlying basic rule for what semantically can exist in a hostname and a localpart, and RFC5321 is extremely lax: you can use any "ASCII graphic or space", and so excludes only ASCII control characters and 8-bit characters... and then, as mentioned, another RFC removes the 7-bit limitation and opens up the world of Unicode.
(To push on it even further: it isn't even clear to me that one should consider the ASCII control character limitation to be fundamental to the email address identifier or a weird limitation of the current version of SMTP; and since none of those email addresses are going to work, I think one may as well just consider the local part to be any string of Unicode code points.)
Think about this: it is up to your SMTP library to correctly escape the email address you give it for SMTP, and here's the fun part: if you give it a pre-escaped email address, then clearly it is going to have to double escape it, right? So, semantically, these extended discussions of quoted strings and character limitations are always just so ridiculous :/... you absolutely should not be dealing in SMTP-escaped addresses or asking your user to understand SMTP (and the same goes for MIME).
(BTW, if you want some "real hell", one of these two protocols--I forgot which... I presume SMTP--seriously supports an empty local part. If that doesn't tell you everything you need to know about un-opinionated these RFCs are with respect to "anything goes" then I don't know what will ;P.)
Lookin’ at you, Walgreens.