Stop Validating Email Addresses with Regex (2012)
davidcel.is
davidcel.is
https://html.spec.whatwg.org/multipage/input.html#valid-e-ma...
"This requirement is a willful violation of RFC 5322, which defines a syntax for email addresses that is simultaneously too strict (before the "@" character), too vague (after the "@" character), and too lax (allowing comments, whitespace characters, and quoted strings in manners unfamiliar to most users) to be of practical use here."
The regex is:
/^[a-zA-Z0-9.!#$%&'*+\/=?^_`{|}~-]+@[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?)*$/
Every browser implements this regex for <input type="email">.Chromium:
https://source.chromium.org/chromium/chromium/src/+/main:thi...
WebKit:
https://github.com/WebKit/WebKit/blob/0d7afc5a45c140c44497a8...
Yes, do verify email addresses by sending a confirmation link if you bind users to their email addresses, though. Don't confuse validation with verification.
^[^@\s\x00-\x1f]+@[^@\s\x00-\x1f.]+(:?\.[^@\s\x00-\x1f.]+)*$
It requires exactly one "@", disallows whitespace and control characters, prevents repeated dots in the domain name, and ensures the domain doesn't end with a dot. It catches a few typos and I think it allows every real email address I've heard of.When is this needed in the context of a Rails app? You never want to send emails to local hosts, and emails directly on the TLD is possible, but very little will work with them in practice.
That's not entirely true, it depends on how your mailing infrastructure is set up.
That doesn't mean you have to support their use. I have never seen a company support an email ending in a period and the world seems to continue along.
Are e-mail addresses case-sensitive? I would always lowercase&trim a string meant to represent an e-mail address (or a domain name) before doing anything else with it.
Email addresses aren't case insensitive in nature, but many commercial mail services (and misguided developers) assume they are.
1. https://en.wikipedia.org/wiki/Email_address#Internationaliza...
Or, you know, if you actually want to use an at sign in the manner allowed by the RFCs. That's the point of standards: if something is allowed, then … it is allowed.
Refusing to accept behaviour permitted by the standard is just broken.
By hearsay alone, I say: comments and quoted strings in email addresses are, de facto, more or less dead. Comments especially, which were intended just as a compatibility measure for stuff from over forty years ago (when this was being developed organically, before there was any real spec). Most software whose direct business is email (that is, MUAs and MTAs) will more or less support them for historical reasons, but newer MUAs are likely not to support them in full (again especially comments), and I suspect the considerable majority of other software that deals in email addresses won’t support one or both of them. (This is a deviation from spec in the direction of restriction.)
One of the biggest problems with email input is typos -- and there are some very common typos that could easily be accounted for with code. For example foo@gmail.co, foo@gmial.com, foo@comcast, etc.
It should be common, when these types of typos occur, to prompt the user to fix them. Unfortunately, this is quite rare.
No, I know my email, thx, it's bein prefilled from auto-complete. Don't fucking tell my Im typing it wrong when 1. Its my email, and I'm not even typing
>>>Me wrote: >>Me wrote: >Me wrote: Me wrote:
Otherwise I don't know what common typos are. I would have to do research on that. Then I have popular local email providers where I may come up with typos. but then again, it is only in my region - I would have to get some list of popular email providers in every country. Come up with a way to update that list after few years when new email providers come out. Or maybe develop a solution that matches "close enough" entries in my pre-defined list.
UX effort could be spent in an area which is used more often than only once. If there is a library for that, yeah, then slap it in, UX has been upgraded in no time. And again, only if revenue is on the other side of consideration, the development effort pays off.
human: a customer telling a clerk, "yes, it's really gmial dot com," and the two sharing a laugh and trading stories about funny email addys
human element: the customer curses their phone because the backend folks never tested the edge case of your helpful button for confirming the suspicious case of "gmial.com" and just keep rejecting it. Then, after your helpful chat bot kept autocorrecting their chat input to "gmail.com" they threw their phone on the ground so hard it broke.
Don't be a human element.
A pragmatic, domain-specific spec beats that.
edit: there's more nuance than that. Some RFCs are standards. 2822 is 'Category: Standards Track', so is a standard if referred to by its STD number? I can't find what its STD number is.
Shitty front ends are definitely putting an end to that, though.
(Alas, <input type=email> still doesn’t support non-ASCII in the local part, which isn’t supported everywhere but is, I believe, fairly widely supported now. See https://github.com/whatwg/html/issues/4562 plus https://en.wikipedia.org/wiki/Email_address_internationaliza... for a little more background on what it is.)
I used an-emoji.my.domain for a while until chrome changed it back to punycode
To validate user input, I use that: /^[^@]+@[^.]+\..+$/. It's doesn't tell me if the email is semantically correct per the rfc because I'm not running an rfc correctness validation service. What I want is to make sure that users didn't input their name in the email field. This tells me if it ressembles an email
Most typos, however, are likely to be in the first part, not in the domain. The only real way to validate against those is to try the address and see.
Otherwise, yeah, most people would be better served by a library that detects domain typos like https://github.com/mailcheck/mailcheck than spending time on regexes.
E.g., Type check that the user's name is a String and then let the database/business-logic yell at you when it's not unique.
In this case, checking that an email address input is "shaped" like 'not-empty@not-empty' is the "type check", and then actually sending an email is the only way for your business logic to know if it's actually valid.
As someone who sends a LOT of email for customers every week, a big percentage of our problem is incorrectly formatted emails. So we just don't even let them in these days.
People are forgetting that for services that have to send email it costs dearly to bounce.
Bounces decrease the quality of your list and you can get penalized by your email provider. By letting in stupid sh*t, you just are compromising yourself. So some simple regex that gets rid of the most egregious stuff is worth it.
easrng@easrng-laptop:~$ echo -e "127.0.0.1\tai" | sudo tee -a /etc/hosts
127.0.0.1 ai
easrng@easrng-laptop:~$ ping ai
PING ai (127.0.0.1) 56(84) bytes of data.
^C
easrng@easrng-laptop:~$ ping ai.
PING ai (209.59.119.34) 56(84) bytes of data.
^CIf it’s a specialist email processing tool, then you should probably follow an RFC.
If it’s a dating app, you can probably just use a regex that covers common cases to help users avoid typos.
I think the decision is similar to the one picking how modern are the browsers you are going to support. It’s a trade-off. That’s my take on it, don’t have a cow, man.
Typos will overwhelmingly lead to valid-looking addresses anyway.
You’re right that you can’t prevent someone from typing their address incorrectly, but you can make sure they didn’t put in their phone number or their first name by mistake
It's overkill for an HN registration form, where if I type my email wrong, I can just re-register, but for checkout forms it can be worth putting in a little extra validation.
I needed to do just that once - I was given a flat file with some bulk export data that came from some other system that I had no exposure to and needed to clean up contact data by figuring out what was what, separating names from phone numbers, from postal addresses, from emails, etc. It just needed some rate of success, not to be perfect.
Lets not assume that we always know what people are trying to do and what is the best way to do it. Devil is in the details.
Being too generous in the early input validation regex (e.g just check for non empty and containing one @) risks making registrations fail and customers not returning. This seems like a much bigger risk than locking anyone out who has a weird email Even a character beyond [a-z0-9-_] or domain without tld is probably worth rejecting.
The entire argument I was trying to make was that what constitutes a "legal email address" is not always what you want to validate. That some providers allow using '+' isn't as important as the number of users that actually do. If some nontrivial number of users use + on purpose, then don't reject them.
1. An experienced dev just killed 54k stars on GitHub due to pressing a button in an auto-pilot mode. Do you think a Joe High who wants to give you $100 won't ever type '2' instead of '@'? What about an old lady? Or someone with physical difficulties? Have you personally ever made a typo in an email?
2. That code in the article is not color highlighted (rainbowed for Regex) or formatted properly. If I write something in any language in one line without highlighting — it'd look unreadable as well.
3. A Regex for this specific purpose is write-once-and-forget. You won't need to edit it for 20 years.
4. Regex — for practical tasks — is way easier than it's being painted. Not easy — just not as hard as some suggest.
5. I'm not a Regex fanboy (nobody is).
I do like to check for '@' to enure the user have not entered their name or something by mistake, but beyond that the syntactic validation does not provide any value.
Which is a pet peeve of mine; I've got an older e-mail address that probably ended up on some list, now there's people from Thailand and the UAE registering accounts using that e-mail address. Now while my account is still secure (2FA, long password, the works), it doesn't stop people from using it. Services like this one webshop and Deezer and probably a few others do not wait for e-mail verification before allowing users to place orders or use their service, or at least the free trial part of it.
I am :)
I won't bet my life on it but probably some other sucker will have to fix it. This happened to me where I found a regex that was validating urls was incorrect after urls were allowed to contain unicode stuff (I do not remember the details just that we had an url from a customer that contained arabic looking characters)
IMO the language standard library should include this stuff of validating stuff to avoid developers copy pasting dubious quality regex from Stack Overflow
There simply is no need to check the email addr provided by the user. Send the mail, if it bounces, the user has only himself to blame. What if I don't want them to go through the hassle of an activation link? Then I don't bother with an email account in the sign-up process in the first place. If they want a passwd reset method, they can later provide an email in their settings page, if that isn't valid, well, tough luck.
1) Consider if you really even need to verify that email address. Why are you collecting email in the first place, why do you need it? HN is a great example of this - email totally optional, if you forget your password it's on you. Lot of online services going in the wrong direction with requiring a phone number.
2) Trust the user. Okay to give them nice nudges ("you probably meant gmail.com and not gmail.co") but if I really did mean gmail.co, let me through if I insist.
Don't throw up a garbled mess of a Regex[0] that only serves to frustrate me when I try to sign up with my vanity email. I'll abandon the sign-up entirely.
[0]: https://stackoverflow.com/questions/20771794/mailrfc822addre...
All solved by a password manager.
Password manager is good if it is planned to share all my passwords after my death with my family, for not gifting my funds to some random guys. In every other cases it sucks like Sasha Grey (from my lifestyle's point of view which involves heavy use of random devices most of them even does not support any passwd mngr).
It doesn't for basically everything I do, and I have a lot of systems, subscriptions, and profiles to worry about.
> random devices most of them even does not support any passwd mngr
In the simplest case, 2 things are needed to support a password manager: Network access, and copy-paste.
All of my mobile devices doesn't really support JS so I can not even input my HN's password to there without a PC.
Also I do not believe all of my e-mails which has some accounts with some values on it are still working. At least one time I had to re-register e-mail exactly as previous to withdraw some of my funds which were untouchable for few years :-) So the biggest part of problem is not on user's sides of wire and I have already tired to struggle to formulate such a simple thing using such a lot of sentences.
Nobody is accepting random self-signed certificates, of course, usually they need to be signed by a CA belonging to the party you're authenticating to, but there's no technical reason why you can't use a random certificate to authenticate with a website, or even modify your browser to add a quick and easy button to generate them on the fly.
Browser vendors have stopped caring about this type of auth and are focusing more on webauthn, which stores a cryptographic token in your device's secure storage (if available) or on the file system. When browsing from a phone, this means it's essentially "sign in with your fingerprint" for websites, which is really cool! You can't easily back those tokens up, though, so you still need something like a recovery email if you don't want your users to lose their accounts when they drop their phones too hard.
If there was a way for end users to sync webauthn logins, we'd solve this problem without ever resorting to any kind of blockchain whatsoever.
(which as far as I understand is also to protect the emailing service from getting blocked itself)
I run an online store, people miss entering their email address is one of the largest causes of customers contacting support, and they regularly jump to being angry accusing us of being incompetent or worse.
I would take the 0.001% of people who may have an email incompatible with a regex being frustrated (which will happen to them all the time) over the 10% who screw up entering their address.
We even have code that looks for common typos and prompt the users to double check them. Somehow they still make those mistakes.
Don't put up a web page where any visitor can put in an e-mail address, to which you send something, without any safeguards: like not sending to the same e-mail address more than just several times in a 24 hour period or something.
Have Captches or or something to reduce the bots. Proof of work. Whatever.
It may be wise to validate not for valid e-mail address syntax, but for certain invalid e-mail addresses to which you shouldn't send.
For instance, would any legitimate user be subscribing with an e-mail address of postmaster@example.com? It seems it would be worth it to have a database of patterns of at least some well known mailing list addresses. Certain domains are almost certainly mailing lists; e.g. anything@vger.kernel.org is probably a list; don't send to it.
Process bounces.
It is mostly because proper defenses against such abuse aren't built into software allowing such forms (or cost money). Wordpress form plugins are one such widespread bad example.
Some products don’t want to let users fail so easily. Especially if they spent good money to get you to the point of signing up.
I work at a large web company. We ran a test around removing email validation and we had about 20% of users typo their emails when signing up. Simple things like not putting the period before com like "john@gmailcom". It resulted in customers basically creating accounts they couldn't get back to which was a bad user experience and loss of revenue for us.
Based on our testing, email validation mostly served to prevent these basic typos.
So your system always makes the least effort and pushes the blame to users. Great.
How about just even do a minimal check that is /.@./ so that the basic format is at least there or if you'd take 10 minutes to look around, you'll find the regex browsers are using and just steal it and be done with it, so most of the malformed inputs are warned to the user before the user realizes the confirmation email isn't arriving minutes (or days) later and possibly lose the conversion right there.
https://developer.mozilla.org/en-US/docs/Web/HTML/Element/in...
2) if second level domain not in list of famous second level domains, AND levenschtein distance is small with a domain on the list (typically gmial, oultook, yhaoo,...), display a huge warning. Don't refuse if the user insists it's correct, but show a huge and red warning that blinks. Same with the TLD to detect "cmo", "ogr", "inof" etc.
3) Maybe send a validation email, now that you've ruled out a big number of potential mistakes. You won't get that many bounces. Or don't send the email if you don't need to!
I fail to imagine a scenario that wouldn't be neatly covered by this 3-step process.
I’ve never understood this, but heard it often from developers and product owners in the industry. “I don’t care about the small number of users who X” where accommodating X is essentially free. Or worse: deliberately taking the eng time to reject users X where accepting takes no work!
Especially in a business context where users X are trying to hand the business money.
I had a tech lead once who say we should reject non-ASCII characters in user input “because they are an edge case.” Nothing in the rest of the data flow or database storage required ASCII characters and it took eng effort to filter them out. Boggles the mind sometimes.
That being said, you need to be very careful with what regex validation you are doing. I still use an apple "@me.com" email. Somewhere there is a commonly used library (or commonly used regex copied from stack overflow) that seems to fail because my domain is short. I have had a number of times that I have been unable to get emails because a system flagged it as invalid.
Getting through to support or trying to change my email is always a nightmare in these situations.
So be careful with your assumptions!
I have found two ways of getting around this issue.
1) Continue registration for a new site using a temporary email. Once logged in, I find I am often allowed to change the email to whatever I want within my user settings. 2) Contact support and request they updated it for me manually.
These two work arounds don't always work. But more often than not they do.
I had a big ecommerce site go into their DB the other day and update from my temp email to my real. Now my profile is fucked because it validates the email address on the page load.
How does it work with common e-mail clients? How do people react when you show/tell them your email?
I have a domain that uses non-ascii characters, and while I can receive emails on that domain, hosted by Fastmail, Fastmail clients refuses to _send_ emails to that domain (I can, if I type the domain as Punycode).
You can see the email address on the front page if I paste the punycode web address on here:
While emoji aren't a use case you'll get many managers to care about, there are plenty of unicode characters that can. Email addresses using foreign script, for one, or even just characters like åäáà, not uncommon in European names, might convince people to consider enabling such features.
Sadly, the process of enabling support for such characters is much harder than it should and I've got to admit they my mail infrastructure also can't handle these types of email addresses. Modern MS Exchange servers seem to have finally implemented support, though, so perhaps we may see more support for it in the future!
Though UTF-8 is not the only thing atrocious with e-mail stacks, there's so much maintained-but-not-really non-standard software that a lot of people rely on.
To end on a positive note, e-mail hosts are starting to demand SPF to accept mail.
It makes it easy for me to keep track of who is sending me what + who is sharing my email with third parties, but definitely confuses some people.
I do
> How does it work with common e-mail clients?
Perfectly
> How do people react when you show/tell them your email?
With confusion, so it requires some gentle insistence that I know my own email address.
My primary email addresses are on two TLDs, one of which has been around for 24 years, the other for seven years, but about 10% of sites I try to register on refuse to accept them, saying that they are not valid. They clearly have not updated some internal whitelist since these TLDs were added.
I just had one major ecommerce go into their DB and update my address. This was a bad idea because now my profile page has an error and I can't change anything else.
STOP THIS. JUST STOP.
I've been programming for 25 years. I've been reading this same article repeated over and over again for 25 years too. Clearly it's not working.
It's not the standard's fault that people implementing are not following the old adage of implementing standards: "Be strict in what you give, be lenient in what you accept."
Considering how there's really not a tremendous amount of variety in what e-mail software people use, it boils down to those maintainers' stubbornness. If you dig trough old mailing list threads and issue tracker tickets, the ossification becomes quite visible.
Lots of commenters insisting that they need to validate somehow. If you have a need to validate, use a good email validation library. Better consistency throughout your app, and someone else has figured out the hardest stuff.
Stop validating emails. But if you can’t stop, at least stop rolling your own regex on the fly to do it. Use a good library instead.
That does mean that in practice it is probably best to consider most new TLDs as web-only. Use them in URLs but have @com, @net, or @org email addresses for anything where you want outgoing mail to get through.
When I was running my own mail system, I eventually ended up with all of the following TLDs going straight to a spam folder:
accountant bid christmas click club cricket date download faith gdn gq help info link loan men party press pro racing review science site space stream team top trade uno webcam website win work xyz zone
When you're dealing with signups for a service, there's absolutely no reason to refuse a working email address.
Some negligible fraction of pathological email addresses will get rejected by imperfect regular expressions. I’ve tested multiple email validation regexes against huge databases of actual emails and found they all validate.
Why validate? JusT SpAm!? It's quasi-legal. Bounce rate? Who cares!! Not like we are monitoring ip reputation anyway amirite?
Missing a period can cause problems, because without a period or a TLD at the end, your mail service might try to send email to internal servers in your network. Send mail to a@b from server c.com and you might end up sending email to a@b.c.com instead. Not checking for a period in the domain part can therefore cause some pretty weird behaviour.
If the email ends in an external domain, there's a period. If the email is intended for an internal host without a full domain name, the period should be at the very end of the address, turning it into a proper hostname.
It's easy for people to type gmailcom and if you're using dot less domains in your network infrastructure you probably know about these hacks anyway. I don't think checking for a dot will break anything, even in the spec, except for something@[IPv6 address] but IP address emails are practically unused in real life anyway.
I see complexity as kind of a mass, and our job is in part to reduce it as far as possible.
Like mass, you can shuffle complexity around, or add more, or discover unexpected ways to reduce it, but there is no way to reduce it below the problem's lower bound.
Sometimes, the solution is as simple as it's going to get, and it's still complicated. State management is a good example.
I'm more referring to that situation.
This article's headline is an example of the argument I mean. If not using Regex is the proposal for reducing complexity, it's not going to reduce complexity.
That's fair. There is definitely irreducible, or essential complexity that we have to deal with.
The breakdown of websites this happens most on are large corporate websites or very small online shops using some no-name shopping cart software.
There might be some people who run their own webservers and use more exotic addresses (local parts), but 100% of them also have more normal email address they use for random website signups. If all websites would make sure to accept + in local parts, and - and _ everywhere, there wouldn't be enough people angry at the situation for articles like this to be written and generate any traction at all.
Sending emails might cost a fraction of a cent if you use a 3rd party service. When it's not free, regex validation can save money. There's a potentially valid reason to do it.
Or maybe the people writing the website's javascript are just control freaks. So what? That's not a reason to argue for all websites to stop doing regex email address validation.
Allow + in the local part, along with - and _ everywhere. That's all you have to do. You don't have to give up regex validation and try to send every typo'd email address that a visitor might enter on a form.
I understand the point about not maintaining complex regex, but I just used an npm package that maintains and updates validation regexs with web standards. Besides any remaining left-pad jokes still standing, I don’t see why I would eschew this solution for any future forms?
As others pointed out, for a sanity-check something like /.+@.+/ should be good enough.
If it has an @ sign, split at the @ sign.
If either string has a space or other invalid character, it's invalid. Yes, there are more invalid characters for the domain than the username (I would make the case the username should just be whitespace chars, and the @ sign).
If the second string doesn't have a dot in it, it's invalid.
Split the second string by dot. If the last string of that array is not a valid TLD, it's invalid. (I realize new TLD's are popping up these days, so this step may not be strictly needed)
Not being too strict is in your interest here. You avoid too many email bounces by allowing some things which may be bad but allowed in some email providers but not others. It's a case of "Perfect is the enemy of good enough".
Nope, could have a quoted string which can contain whitespace.
addr-spec = local-part "@" domain
local-part = dot-atom / quoted-string / obs-local-part
qtext = %d33 / ; Printable US-ASCII
%d35-91 / ; characters not including
%d93-126 / ; "\" or the quote character
obs-qtext
qcontent = qtext / quoted-pair
quoted-string = [CFWS]
DQUOTE *([FWS] qcontent) [FWS] DQUOTE
[CFWS]
> If the second string doesn't have a dot in it, it's invalid.Nope, domain names dont require a dot.
<domain> ::= <subdomain> | " "
<subdomain> ::= <label> | <subdomain> "." <label>
> Split the second string by dot. If the last string of that array is not a valid TLD, it's invalid. (I realize new TLD's are popping up these days, so this step may not be strictly needed)Doesnt need to be a "valid TLD" to be a valid domain
The larger point I was trying to make still stands.
Your suggestion is to not use a regexp, but is less valid, and likely less efficient than a regexp that checks for the presence of an @ anywhere than at the start.
This is actually the whole point of what I wrote. Making the maintenance of the code low-effort, and having it be more permissive (i.e. less valid) and relying for the corner cases to be taken care of via email bounce.
See: explainer on "Now you have 2 problems"
https://arstechnica.com/information-technology/2014/05/what-...
If you read this whole comment thread you can see where this leads: increasingly complex regex in order to handle all the corner cases. Forget the corner cases, just make it "good enough" and relegate those corner cases to email bounce.
Now - I don't have a case study to prove that this way of going about it is more valid than the strict regex validation you're suggesting. But I wanted to represent it as a middle ground between "just rely on email bounce" and "write a big long regex".
Some people, when trying to explain something, think "I know, I'll use a Jamie Zawinski quote." Now they have two things to explain.1. The user made a mistake when typing the email.
2. Some internal server failed to send the email.
Without an error, the user will never know what went wrong.
All of the other stuff is important too, but the most common issue is going to be someone tying "gnail" instead of "gmail", not having an obscure domain.
On a related note, I have mylastname@gmail.com and I receive emails from other people with my last name several times every week. I am almost certain they mean to type [firstinitial]mylastname@gmail.com but miss that first character. That's my working theory anyway. I've had to intervene multiple times when getting sent important medical emails for a different person with my same last name.
It’s is one of those “technically” vs “practicality” things, not using a regex will cause you more customer support problems than solve.
I run an online store, typoed email addresses are one of the top causes of customers contacting support. In 10 years we have never had a customer contact us to complain their “valid” email address won’t be excepted on our site. If we were to remove the regex from the validation and “just try to send an email” as suggested it would create so much more work for our support team.
You should also implement some “soft” validation looking for common typos, although people still somehow make those mistakes. I’m convinced that some people have types in their autocomplete address book.
I won’t ask people to type the email twice, that’s just annoying.
.+@.+\..+
Works every time 100% of the time.https://mail.gnome.org/archives/evolution-list/2002-January/...
I wrote a short (but incomplete) research comparing mail servers (MTAs) and their handling of local-part in my ongoing effort to reduce spam and to tackle who is leaking my email address.
‘Local-part’ is the part of the email that is between your account name and the ‘@‘ symbol. For some MTAs, it CAN be the part before your account name.
https://egbert.net/blog/articles/comparison-of-local-part-in...
https://github.com/JoshData/python-email-validator (my project)
The README covers a lot of ground: internationalized domain names, internationalized local parts, SMTPUTF8, Unicode normalization, not performing SMTP checks, not permitting obsolete email syntax, and missing UCS-4 support in Python 2.7.
/[^ ]+@[^ ]+/ is OK except for trailing punctuation, double @'s, and it definitely doesn't work with that quoted example. The first Rails one in this post /\A[^@]+@([^@\.]+\.)+[^@\.]+\z/ is at least better than what I came up with - I'll take a battle-tested regex if it's on offer.
But instead of showing an error you can just kinda say, hey are you sure you want that? and let them do it anyway. this is more of a validation suggestion.
Also, never check file existence.
These things just take up time, introduce race conditions, and can't be trusted anyway.
That said, there are reasons to want to check some things with UI-level validation because calls to a slow back end are slow. Thus the name. I get that.
So if you're doing UI validation, don't get it right. Just get it mostly right. And do it fast. Cheat!
This caught my eye and I’m dying to know more - could you elaborate or point me to a good resource on this? My team has been dealing with some issues related to this recently.
[0] https://en.wikipedia.org/wiki/Time-of-check_to_time-of-use
https://www.youtube.com/watch?v=xxX81WmXjPg
It's more a joke than edifying, but it was very fun for me (and, I think, the audience), and illustrates the difficulty of email validation well.
There is not reason whatsoever to validate to a standard just because it's a standard. Validate to what you will support.
For example, control characters, line breaks, shell escape characters, SQL injections, or simply uploading an ISO image into the E-mail field.
There is the RFC, and there is what we would nowadays consider a sane E-mail address. Nobody has addresses with spaces, nor does anyone have an address with an IP literal in it. Why? Because no other system will accept it. Think bank, etc.
TL;DR at the very least you need to validate field length, control characters, quote signs, and backslashes. Those have no business being in am e-mail address.
(?:(?:\r\n)?[ \t])*(?:(?:(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t]
)+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:
\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(
?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[
\t]))*"(?:(?:\r\n)?[ \t])*))*@(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\0
31]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\
](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+
(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:
(?:\r\n)?[ \t])*))*|(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z
|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)
?[ \t])*)*\<(?:(?:\r\n)?[ \t])*(?:@(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\
r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[
\t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)
?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t]
)*))*(?:,@(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[
\t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*
)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t]
)+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*)
*:(?:(?:\r\n)?[ \t])*)?(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+
|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r
\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:
\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t
]))*"(?:(?:\r\n)?[ \t])*))*@(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031
]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](
?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?
:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?
:\r\n)?[ \t])*))*\>(?:(?:\r\n)?[ \t])*)|(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?
:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?
[ \t]))*"(?:(?:\r\n)?[ \t])*)*:(?:(?:\r\n)?[ \t])*(?:(?:(?:[^()<>@,;:\\".\[\]
\000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|
\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>
@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"
(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*))*@(?:(?:\r\n)?[ \t]
)*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\
".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?
:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[
\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*|(?:[^()<>@,;:\\".\[\] \000-
\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(
?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*)*\<(?:(?:\r\n)?[ \t])*(?:@(?:[^()<>@,;
:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([
^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\"
.\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\
]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*(?:,@(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\
[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\
r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\]
\000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]
|\\.)*\](?:(?:\r\n)?[ \t])*))*)*:(?:(?:\r\n)?[ \t])*)?(?:[^()<>@,;:\\".\[\] \0
00-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\
.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,
;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?
:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*))*@(?:(?:\r\n)?[ \t])*
(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".
\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[
^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]
]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*\>(?:(?:\r\n)?[ \t])*)(?:,\s*(
?:(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\
".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*)(?:\.(?:(
?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[
\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t
])*))*@(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t
])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?
:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|
\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*|(?:
[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\
]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*)*\<(?:(?:\r\n)
?[ \t])*(?:@(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["
()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)
?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>
@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*(?:,@(?:(?:\r\n)?[
\t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,
;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t]
)*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\
".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*)*:(?:(?:\r\n)?[ \t])*)?
(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".
\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*)(?:\.(?:(?:
\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\[
"()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])
*))*@(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])
+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\
.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z
|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*\>(?:(
?:\r\n)?[ \t])*))*)?;\s*)
Furthermore: "The regular expression does not cope with comments in email addresses. The RFC allows comments to be arbitrarily nested." Thus, clearly RFC 822 addresses cannot be described by any regular language. Ooof. I've been internettin' for a few years now and I had no idea that the email address format specifies comments.(Personal experience, yes.)
daydreams about strong cryptography
The original standard for email addresses seems to be so bad that it's just being scrapped and ignored. I think I'm OK with this.
```
//pseudo Go
import (
"strings"
"internal/inputs"
)
func emailValid(input inputs.ProfileInput) bool {
if len(input.Email < 3) {
return false
}
return strings.Contains(input.Email, "@")
}
```
An email should be at least 3 chars long and contain a @ sign to be valid.a@a can be a valid email address if your hostname is a and your mta accepts the email address. A user a may or may not exist (virtual).
And as the author wrote 10 years ago, if the mail doesn't arrive, no validation can help if the email address doesn't exist. But len>=3 and @ sign presence is enough of a sanity check.