Stop Validating Email Addresses With Your Complex Regex
davidcel.is
davidcel.is
"We see that bob@localhost doesn't look like a email address are you sure it's right?"
That way you can help users that messed up their email but not prevent all the corner cases. The idea is that most email addresses fall in a very narrow subset of the RFC: user@domain.tld and most people would have entered their email wrong if it didn't match that pattern.
From a previous startup we saw a ton of signups like, "john@gmail" and the like. Obviously this person will not get a validation email -- and in all likelihood will not be able to log in to his account when he returns. It's best to catch him when he's entering the information.
Anyone who says "let the confirmation email handle it" clearly only has HN readers as customers. The type silly stuff, then email support to complain they never received the confirmation.
At Ventata, we used to have the same issues you've all described with people forgetting things like ".com" and "gmial" vs "gmail". Once we started using mailcheck our bounce rate went way down. Now we only get bounces when people deliberately give us faulty addresses, but we aren't really concerned with trying to convert them. They are checking us out and want to stay anonymous, I don't mind that.
That condition aside, Mailcheck is pretty much all you'll ever need.
For example, it catches foo@gmail, but it doesn't catch foo@outlook. Of course, gmail is going to be the more common case. Still, it looks good overall and an improvement over no client side validation.
Edit: Actually, I see you can define a set of domains to check against. Very nice.
I've tried to help this situation by creating an API for you guys: https://www.emailitin.com/email_validator
If you want to prevent "john@gmail", then use a real RFC-compliant e-mail address parser, and attempt to resolve the domain component MX/A records (and remember, it might be an IP address).
If that fails (or your regex fails, or whatever validation you use), ____SUGGEST____ to the user that the address appears to be invalid. There's no reason for an overzealous registration form to refuse to accept the user's actual e-mail address.
Bingo. Help people, don't hinder them.
I don't see why it has any material affect on someone requesting e-mail addresses: if it's valid, then it's valid.
This seems to be an example of the misplaced sense of propriety with which people approach validating e-mail addresses -- that somehow, your job isn't just to help the user enter an address, but also to define what is and is not 'reasonable'.
Imagine if companies refused to accept "1 Infinite Loop" as a street address because it's clearly ridiculous. Except that it's Apple's actual street address.
Overzealously rejecting valid addresses is an application of subjective and inaccurate ideas about what addresses 'should' look like, and ignores the simple fact we've already mutually and formally defined valid address formats via the IETF RFCs.
Yes! Great Comic Book Guy impression.
Why it matters is that for most smallish companies, you want to get something up that helps your users not do stupid stuff (†), but due to time and resource constraints, you're likely to end up with some kind of 80/20 solution. It'll work well in most cases, and fall down in some others. I would certainly agree with the idea that you not force people, but a nudge is probably going to save you money in increased user retention and fewer support hassles.
† - I once had a person ask why their emails to http://example.com were failing.
In that case, great junior engineer impression on your part.
Pedantry matters in complex interoperable systems, because otherwise they're not interoperable. This is why we have detailed standards documents on e-mail address formats.
> I would certainly agree with the idea that you not force people, but a nudge is probably going to save you money in increased user retention and fewer support hassles.
A 'nudge' isn't going to come from yet-another-broken-email-validation regexp. There's no need for an 80/20 solution; this just isn't that hard.
> † - I once had a person ask why their emails to http://somesite.com were failing.
That's not a valid e-mail address (as per RFC822).
During all those new account registrations they're constantly making?
Does this feel slow to you? https://www.emailitin.com/email_validator
> In this case worse is better.
No, it's not. Every time you exclude a valid e-mail address, you lose or annoy a customer for no reason other than your own lack of understanding of the RFCs.
There are a number of steps that can be taken that don't involve broken regexps that filter out valid addresses. Please stop trying to justify doing it completely incorrectly.
Yes, it did.
For example: requiring shirt and shoes in a restaurant will block some customers from dining there. But it improves the experience for all the other diners, so it's a net win.
I imagine I would have trouble renting a car with cash, even though it would allow more customers.
In what world is not validating an email a good thing? It's not like emails vary after a certain complexity is reached. A better article would have been someone documenting a validation regex that approaches perfect without exceeding insane complexity.
Next we'll see articles to not run the Luhn algorithm on credit cards =/.
Validating an email address is important. The way you do that is send an email to that address. You can't do it with regex, and attempting to do so leaves you open to a variety of flaws.
Like I said, when you accept a credit card, normally you validate the formatting of the card before trying to charge money using the info. Emails are very similar.
You appear (but perhaps I'm misunderstanding you) to be asking a user to enter their email address; checking that against a regex; storing it; and sending email to it. That's bad, don't do that.
> You potentially just lost a user and/or customer (or made them unhappy because now they have to register again)
You're potentially losing customers because their valid email addresses are not validating through your broken regex; or their incorrect email is validating through your regex.
> when you accept a credit card, normally you validate the formatting of the card
Credit cards are trivially easy to check for formatting. You use the Luhn algorithm which tells you if it's possible for that number to be valid or not. This is because there's a strict format for credit cards. There is no such format for email addresses. That's why the only sensible way to check a user's email address is to send a confirmation email to them.
The problem is that people are naive and write incorrect regexes. Also, don't attribute bad programming to me in your comments when you have no idea what regex I use - that's just rude and belligerent.
Emails are trivially easy to check for basic user errors - such as leaving off the TLD or not even providing the domain. Saying otherwise is just being naive again.
* to ensure it is deliverable? Well, then you better send them an email.
* to let people know when they misread the labels and put something that was clearly not an email in the email field? A simple check for an at-sign is usually sufficient.
* because some tester opens a ticket saying you can enter an invalid email in the email field? Yeah, that's where most of the complicated regexps come from.
I actually think that this library functions as a really great client-side validation that won't get you tripped up in trying to be RFC compliant. There's really not anything more that I'd do aside from sending that blessed confirmation email.
Is he really though?
We need a much more restrictive standard for emails, but until we have that we have to accept that each site will have a competing not-completely-overlapping set of standards. So make sure your email doesn't attempt to do anything too funny.
So while I agree with you in principle, in practice it seems infeasible. The differences bite in practice, unfortunately.
"My email address is joe.aol or was it aol.com@joe? Wait joeaol@com?"
It's also worth considering that one of the biggest generational markets (baby boomers) include a lot of those people you're telling us to ignore.
to ensure it is deliverable? Well, then you better send them an email.
I deal with user support for a site and I'd estimate at least 2% of our new users (>50 people PER DAY) enter wrong email addresses. Not "I forgot to put .com at the end" but "I thought my email was john.doe@gmail.com when it's actually john.doe@yahoo.com" which would pass validation with flying colours. The only real "solution" is to tell a user if the validation email has been sent yet (to deal with "well maybe I should wait 5 more minutes") and if it has and they don't have it allow them to change their email to their real email address. So many sites (incl. the one I manage) do not allow this, it's crazy.Couldn't an unscrupulous individual use that feature to take over non-activated accounts? An immediate use for that exploit doesn't spring to mind but this makes my spidey sense tingle. What sites allow this?
This is the source of 80% of all "bugs" I've fixed over the years.
Another personal favorite: If you enter WWWWWWWWWWWWWWWWWWWW W WWWWWWWWWWWWWWWW for name, it messes up the layout on the display screen.
If the issue only appeared on all W's I'd guess you could set the bug as minor and discuss if it's worth fixing. If it costs an incredible amount of time to fix it, the problem relies more on defining priorities than on having too much granularity on the testing side.
To go back on your parent post, I'd say validating funky emails is of the same level. The test team should bring up the edge cases, fixing them or not is a matter of priorities.
So we did lots of multivariate testing, with permutations of regex patterns, MX record validation, sending a confirmation email, client-side only, server-side only, etc.
What we found was that a decent regex pattern (we used \b[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,4}\b, taken from http://www.regular-expressions.info/email.html), with MX record validation and common junk domains blocked (e.g. mailinator.com) produced the largest conversion rates in the follow-up email.
In other words, less validation produced more entries, but they would have been lower quality, which affected our sender reputation and cost more. The confirmation email was awful for conversion. YMMV.
/.+@.+\..+/i
A trailing dot is valid at the end of a domain name. These domains are said to be fully qualified.http://www.dns-sd.org/TrailingDotsInDomainNames.html
This also potentially eschews internal deployment.
His other suggesiton, /@/, at least, is not harmful, but there are validators based on the RFC that do the job for most platforms.
Validating the address may provide useful feedback to users who accidentally mistyped something, invalidating the permise of his last paragraph.
.
/someone is wrong on the internet.
It still amazes me that 70% of the places I attempt using foo+bar@gmail.com call it invalid. And that does not even begin to touch the myriad valid permutations that are "invalid" out there.
You can then block any address with a click if it's being abused.
If you're feeling generous, enough people using this link will earn me premium features: http://www.33mail.com/rj37w3
I think there is going to be a very low percentage of users who will register again if they don't receive a confirmation email, unless you're giving away free iPhones. Granted regex can not catch most of the user-generated errors, but it can catch a few which could still increase you registered users.
To that point a better UI/UX (font size, spacing, etc.) might do a better job in lowering typos in email.
Incase I've typed it wrong, that should basically work for anything that contains at least one @ and one dot, in that order, as well as at least one character at beginning, middle and end. It's served me well thusfar.
Edit for clarification: The reason I prefer this over just checking for an @ is that if you're just checking for @ a common mistake like "me@hotmail,com" will be considered valid.
Somewhat artificial, yes.
I wonder what the security implications of this is.
I wouldn't add IP address support to an email regex because I'd rather turn away such perverse data anyway. Nobody uses IP-based email addresses.
And I just tried sending an email to me@[my.ip.addr.[0]]. Postfix somehow recognised it was for this local host, but failed because it wasn’t part of the virtual domains I had put into it. ‘perverse’ seems to be somewhat appropriate.
[0] Sorry, too lazy to check the appropriate documentation IP address ranges. 20db?
Why not just check for x@x?
^(.+)@(.+)$
max length is 254 according to the RFC I believe, so you can check for that too.
http://stackoverflow.com/questions/386294/what-is-the-maximu...
Turns out the 320 value was actually incorrect, its 256 but the mailbox is wrapped in square brackets, so 256 - 2 = 254.
My guess is that nobody will use just the TLD because of the confusion. And because a lot of software does not support fully qualified names.
It should be similar to your version, but only matches just enough parts that require for email validation (i.e. "o@example.c" part of foo@example.com).
It has regularly come up on HN, and pretty much any programming related forum I've used since the mid-90's.
As an industry at the heart of the information society you have to wonder what the hell we are doing wrong if we cannot stop this constant regression into well known bad practices.
What I don't understand with this ever-repeating discussion is why the complexity has to be visible. e.g.
> <LARGE REGEX>
> Yeesh. Is something that complex really necessary?
Many functions are complex - we put those in libraries, pushing them under the hood, and move on.What is so special about parsing email addresses that makes everyone invent their own solution - regex or otherwise?
Also it's bothering for the user, if you need mail confirmation then do it, but otherwise it should be a RULE OF THUMB to always avoid annoying user. Thus avoid mail confirmation.
This article is actually a really bad advice. I don't know why it's upvoted so much.
A valid email address can contain almost anything; this makes validation via a standard parser mostly useless. As such, devlopers reach for stricter parsers out of a combination of a not comprehending the standards, feeling vague discomfort about letting 'just anything' past data validation, and misplaced concern for users that they believe can't type their own e-mail address.
Add to that the occasional business complaint from the marketing arm about bogus e-mail addresses, and you have people repeatedly solving the problem in slightly different ways, justifying their own divergences from the standard by applying the justification that nobody will use a 'weird' address anyway, and they're actually being helpful.
How is this misplaced? People screw up even the most basic of computer tasks all the time.
2) There are so many ways to get the e-mail address wrong that it's almost not worth bothering validating the few things that you can validate.
Now, here's what would be an interesting validation method that doesn't actually require sending an e-mail. It requires an RFC-compliant e-mail parser, not a regexp:
- Perform A/MX lookups on the domain part. The domain part can be an IP address, so those get a free pass.
- Connect to the returned MX, issue a MAIL FROM+RCPT TO:
c> MAIL FROM: test@example.org
s> 250 2.1.0 Ok
c> RCPT TO: is_address_valid@example.com
s> 554 5.7.1 <is_address_valid@example.com>: Relay access denied
c> RSET [reset the transaction, no e-mail is sent]
- If you get back a permanent 5xx error, the address is invalid. If you get back a 250 Ok, the address is probably valid (it could still be a relay that allows backscatter, in which case it will allow any address on one of its configured domains). If you receive a 4xx, the address may or may not be valid -- graylisters will send 4xx, as will servers that can't currently accept e-mail, etc.This gives you definitive failure (5xx) and almost-definitive success (250 Ok). It's a cheap DNS lookup + TCP connection that you can begin performing immediately and asynchronously when a user enters their address in a form.
... or just send the user an activation e-mail.
* There's no space between FROM: and the address in SMTP
* Email addresses must come between angle brackets
I'd reject (give you a 5xx) that from my mail server for those reasons alone.I typed it out live. I'm not an SMTP client and I don't have the RFCs memorized.
> I'd reject (give you a 5xx) that from my mail server for those reasons alone.
Postfix accepts it. I haven't checked the RFC to verify your concerns, but assuming they're correct, then my expectation is that postfix is liberal in what it accepts because A) it's a good idea, and B) a real mail transfer agent probably ignored those two minimal rules at some point in the past.
No, it might be worth scoring the e-mail with a spam filter, but the MTA shouldn't be overzealously throwing away e-mail.
Whether you think you're 'Throwing away' is just semantics. From a user's perspective that's exactly what you're doing.
Want to know why it's not more common than the regex "method"? His method has its own host of problems - what if your mail server is down for six hours - will people come back to your site six hours later when they get the email? What flags will get set on your sender account when Gmail gets 100,000 bogus email sends? Do you force your users to "Look in your inbox and click the activation link" for every email address change also? There are others but I've made my point. There's a finite amount of "stuff like this" that users will put up with - you can either put the onus to "get it right" on the user (regex validation for emails), or you can put that onus on your system.
An argument for another is always, "If a user can't get their email address entered correctly, I don't want them as a customer". And you can take that multiple ways - technical difficult entering emails, "challenging" email addresses, etc.
We actually had an email list of ~50k people that had been validated within nothing other than "check there are at least 3 characters in the string" and when we looked at which addresses were bouncing when we sent to them there were approximately zero that failed because they had ommited the @ or because they were using some weird invalid unicode.
Even the spam bots were submitting valid email addresses.
This is going to your dominant type of failure.
Kinda harsh sentiment considering that we all mistype stuff, especially on site that disable auto complete.
filter_var('bob@example.com', FILTER_VALIDATE_EMAIL)
more info here : http://php.net/manual/en/function.filter-var.php
Saying the user will just come back and register again is not good. That's like saying if your page is very slow to load, users will just wait for it to load. They don't. They leave and never come back most of the time.
If you can't afford to write the code to help the users fill out the form in a way that will work, fine, that's something you didn't have time for considering the percentage of users it will help/retain. But don't claim it is useless.
No. This puts the burden of checking email validity on every user, even perfectly capable valid users. If you're validating for edge cases (mistakes or otherwise invalid addresses), treat it as an edge case and don't annoy users who can type.
? Whenever I hit a form which wants me to retype my address, I just triple-click to select the entire address, then middle-click to paste it into the confirmation field.
:(
(Looking up and implementing a regex) * 1 + (running the regex) * (every email) + (sending email) * (every valid email) < (sending email * every email)
Also, this post only considers the signup/activation use case. If you're getting an email for ecommerce to send an order confirmation, you want to know if the email might be invalid before the user completes the order and you try to send it.
After some very basic checks, e.g. "contains at at least 3 chars, one of which is an @", you should Just. Send. The. Email.
Who bothers to type in a complex but invalid email address? The overwhelmingly common failure modes are:
1) Nothing entered at all. The basic check catches this.
2) Deliberate invalid email address. e.g. homer.j.simpson@springfieldnuclear.com - a regex will not catch this.
3) Typo in email address. e.g. john.smith@gmial.com - a regex will not catch this either.
The regex has downsides and complexity, but essentially no benefit.
A regexp for validating RFC-2821 email addresses is actually fairly simple.
You are only trying to catch email addresses that are entered in error at account creation time so that a user will actually get the confirmation email.
The actual problem is that if they enter an email address incorrectly they will crate a dead account that they can never log into again. In addition if they used their favorite user name, or a referral code or any other important consumable when creating the account then you've effectively blocked that user from even creating a second account.
The real solution is to use validation email to confirm an email address, but to allow them to login to the account even if the email is not yet validated. You won't even have to make them type it in twice. Simply limit them to only being able to edit account information and settings.
Email is validated, users have a window to correct any issues and you've eliminated the unrecoverable error altogether... oh and no regex.
<input type="email">
Of course, if your user is not using an HTML5 compliant browser, then this will be ignored.You want an RFC821 (or more specifically RFC5321 now) email address regexp. See my post here about the email validator I wrote: https://www.emailitin.com/email_validator
Edit: Oh wait, on your own sites :-) I'm just slightly annoyed that I have to add a fqdn when doing local development with some apps...
However, they can, if implemented correctly, tell you if the email address is syntactically valid.
Of course there's plenty of ways around that, but this seems to be the most common pattern.
I've included both jQuery and server side example code on the site.
Thus, you must send a confirmation email, with a "click to confirm" link in it.
This keeps your email address list clean; it also validates all the email addresses.
1) use LPeg or something similar to validate the actual text of the email (here's some LPeg that parses the headers of an email, certain one can pull out the email address portion: https://github.com/spc476/LPeg-Parsers/blob/master/email.lua).
2) Take the domain part and do a DNS MX lookup on it (to be pedantic, if that fails, then one should do a DNS A lookup). That will check if the domain is at least valid.
While I definitely enjoyed how Friedl's book (http://regex.info/book.html) builds over several chapters to an ever more complex solution, maybe a page long, my takeaway was: don't bother. A friendly UI will help users avoid an obvious mistake, but as other posters have pointed out, the only real validation is, does an email get there?
As a side note, be cautious of using such a tactic. I have recieved their logins, CC and Physcial Address information because of this.
However, I think most people know how to at least access their email (always logged in), so provided you could get them into their client, with a token, quickly, with a small number of clicks, might be interesting. Of course, it could be spam central.
Depending on how you're sending the mail, it may be possible to insert arbitrary headers and body after a \r\n in the email address field. I know I've built at least one system that is vulnerable to this. Then you can put the body after your special headers and hide the rest of the message (either as an attachment or an HTML comment).
This then makes your signup form into what is effectively an open relay.
It's not 100% idiot-proof, but I'd imagine it would be pretty effective for laypeople and hackers alike.
We use this clever (and well-explained) solution from http://my.rails-royce.org/2010/07/21/email-validation-in-rub...
First of all - less incorrect emails sent - less chances to get marked as spam host. Second - it is very easy to catch obvious errors user can do on front end and ask user to correct it. These two is big ones imho
Some people use many email addresses and so could create many accounts. With one email address they can still use the '+blahblah' method to sign up unlimited times, unless you prevent that which would annoy people who use it legitimately for filtering.
Some people have a garbage or throwaway email account that they sign up for everything with, and only ever look at to find the confirmation emails.
If people don't want to give you a valid email then there's no reason to be sending them anything.
Regex from the source: https://github.com/php/php-src/blob/master/ext/filter/logica...
http://fightingforalostcause.net/misc/2006/compare-email-reg...
I personally prefer: /[^@]+@[^@]+\.[^@]+/
Basically the same except that it will throw an error if someone enters an extra at sign.
/^[^@]+@[^@]+\.[^@.]+$/
"a# b.@c"@[IPv6:2001:4dd0:fc8c::1]
a#\ b.\@c@[IPv6:2001:4dd0:fc8c::1]
now :-)That's just me.