Stop validating email addresses with your complex regex
davidcelis.com
davidcelis.com
I don't know that anyone uses it, or would even want to use it. It was a fun project, but I certainly wouldn't use it in an app (unless it was an MTA or MUA). What comprises a valid email address is way too broad.
I think there are reasonable regexes for email addresses FWIW. The RFC(s) are stupid and broken, and you can hit 99.99999% of addrs and provide useful errors back to the users with a regex.
http://github.com/sixarm/sixarm_ruby_email_address_validatio...
We use a combination of client-side JavaScript validation and server-side validation in Rails. Typical server-side validation is for REST JSON API calls by third-party apps, and also for parsing freeform text fields like "tell your friends about us".
Original code is by Tim Fletcher & Cal Henderson.
Don't test for an RFC-compliant address if you don't want to accept all RFC-compliant addresses. Being able to send an email is a much better test because it matches what you're going to use the email address for in your app.
It wouldn't be RFC-compliant, but it would catch 99% of typos.
Instead of being an error when the e-mail fails validation though, it would say something like: "your e-mail does not appear valid; please double check your entry. You will be sent an activation e-mail; click [Continue] if you're sure the address is valid."
Basically if it fails the "99%" test, then if that fails, let the user decide if their e-mail is in the 1% or not.
However, that specification, and the implementation you provide, is really designed for parsing e-mail address headers (as you say: for an MTA/MUA), and so contains a bunch of properties specific to structured MIME fields that really has nothing to do with e-mail addresses.
Instead, if you are verifying an e-mail address that someone types into a form, you probably are looking for "the kind of e-mail address that SMTP would accept for delivery", and that is covered by a different standard with a different and unrelated grammar.
Specifically, you implemented RFC 2822, the successor to RFC 822 that has now been obsoleted by RFC 5322, the standard on "Internet Message Format" (in essence, MIME). The related RFC 2821, the successor to RFC 821 that is now obsoleted by RFC 5321, is for SMTP.
For an example of the kinds of differences this would cause, RFC 5322 (with errata) believes that ""@example.com is invalid (by errata), but hello(ignore)@example.com is (MIME comment); RFC 5321, on the other hand, believes the exact opposite validity.
(edit: When I realized that I should probably write a blog post about this, given how much time I've put into implementing this stuff recently, I realized that there was more to say on this general subject, and I'm including it below.)
That said, I will go even further: these formats are designed for escaping e-mail addresses in the context of a larger standard and protocol, one that might already have special characters. This is why they contain so much quoting support.
This is then why the grammer is often so highly restricted for things that don't need to be quoted: given that an @ cannot be found in a domain name, you really shouldn't need to quote anything to the left of the @ to get a valid e-mail address.
However, "(" is a special character in a MIME field (begins a comment), and thereby if you want to include it in the local part of an e-mail address, you will need to escape it somehow; the same is true of things like whitespace, commas, or angle brackets.
The user typing the e-mail address into the form, however, isn't dealing with these restrictions: asking him to escape special characters in his e-mail address seems silly: one might as well be asking people to HTML escape their username in the username field.
That said, there is then a separate RFC 3696 which talks about the semantics of contemporary e-mail addresses and how one might go about validating them, and it includes the idea of quoting in its implementation (so maybe it believes that RFC 5321 is king).
If there are legitimate emails that don't follow RFC, you should absolutely allow users to enter non-RFC-compliant email addresses.
[submit]
does email address have a '@' and a '.' after that?
yes -> send a validation email.
no -> Hey there, [email] seems like an odd address, are you sure?
yes -> send a validation email.
no -> restart
I think this does a reasonable job of staying out of everyone's way, and it catches a large percentage of the actual typographical/user error email entries.[formatting edit]
;; ANSWER SECTION:
ai. 14400 IN MX 10 mail.offshore.ai.
;; ADDITIONAL SECTION:
mail.offshore.ai. 14400 IN A 209.59.119.34I'm unaware of any common MTA software which wouldn't handle it.
I have a cow-orker who's first name is "G" - he's very used to poorly-validating websites claiming that his firstname is "wrong".
I've got another friend/acquaintance who's got a fully legal (self chosen) single name. No first/middle/last name, just a "mononym". He has the expected hilarious outcomes with things like Google's "real name policy".
Edit: sorry for the echo, bigiain - I saw your post right after adding the comment.
A simple check for an "@" sign would go a long way to avoiding bounced mail notifications from usernames being entered in email address fields.
But not all the way. And so a simple "you might have got this wrong" flag would be helpful, no?
In our case, accounts may be created by third-party apps using REST JSON APIs, so we want to let the third-party app know that the email address isn't RFC-valid.
But yeah, that's effectively dead and you probably don't want to try sending mail that way.
Edit: punctuation.
Of course, whether any email server actually accepts that is a different story.
Anyway, the regexp I use is /.+@.+\..+/ which supports the address you describe, but (usually) catches the relatively common mistake of user@yahoo,com
We have seen a real email address without any dot, and it routes successfully to a TLD MX exactly like you describe so yes, it does happen. :)
There's very good odds that the email they send will have their "From:" (or "Reply-To:") address correctly set. Then just have an email autoresponder which emails them back a link with a token in it, when they click on that it'll take them to a page to create their account, with their email address already filled in by the token.
Ideally, all this would be done away with by OpenID or client-side certificates or something along those lines. The whole password-and-email-in-case-you-forget-it paradigm is broken, really.
So on our production servers, I need the @, but on our dev/text servers, I don't. And no, the domain is not appended to the address before sending. That's the part I actually know about.
I wish web forms had a "Tell us we are wrong" button next to validated boxes.
I wish web forms had a "Tell us we are wrong" button next to validated boxes.
You know... that's a very good idea actually.What the heck, how can I take MS Word online?
https://wikipedia.org/wiki/Postcodes_in_the_United_Kingdom#F...
Also, something not many people know is that any postcode in the UK has a maximum of 100 addresses in it.
You can tell I've been writing a search algorithm for UK address data recently ;)
real email: name@hushmail.com
servicexyz: name.servicexyz@nym.hushmail.com
serviceabc: whatever-isnt-already-taken@nym.hushmail.com
https://www.hushmail.com/ (no affiliation)Even worse than excessive validation is when they make you change your password often, for no apparent reason.
On some sites that do both (ahem Apple), the only way I can login if I haven't been there in awhile is the security questions or the password reset mechanism.
Weak password: correcthorsebatterystaplefoobarbaz
Strong password: password1
The worst is when they do that and enforce a maximum password of 8 or 12 characters (I'm looking at you, every bank in the US ever).
And I've have my problems with short addresses before with Microsoft. http://answers.microsoft.com/en-us/windowslive/forum/liveid-...
If you have a Rails app you’re most likely using the Mail gem to send mail, so that’s why I wrote this: https://github.com/codyrobbins/active-model-email-validator. It lets the gem worry about whether an address is valid. Since the mail library is actively maintained, in my opinion it’s a safe bet to trust that it is properly parsing and validating addresses insofar as is possible.
http://stackoverflow.com/questions/201323/using-a-regular-ex...
and that demo email address may not be as valid as you think
http://isemail.info/%22Look%20at%20all%20these%20spaces%20%E...
http://www.webdigi.co.uk/blog/2009/how-to-check-if-an-email-...
Or use a service that will do it for you:
You can let the user immediately know that the email was sent, but then you can also push updates to the user whilst they're still on the site if there is a rejection/bounce.
I wrote https://emailprivacytester.com/, which does the DNS checks that I mentioned when you enter your email address, and then keeps you informed of the status of the message delivery as it happens, including any SMTP rejection messages.
Also none of the e-mail systems I've operated in the last 15 years or so will let on whether or not the user actually exists until at earliest when you have committed to sending a message, and many of them wouldn't even then (instead accepting the message and sending a bounce) to reduce the spam harvesting.
This is horrible practice from the perspective of a mail server. Too many illegitimate email addresses and you will start getting more "soft bounces" (a polite way to say, "We're not delivering this") when attempting to send mail. ISP's keep score for how accurate a mail server is when delivering mail, attempting to sift out spammers. When your score dips to a certain level, ISP's stop cooperating with you and basically label you as a "dubious/bad actor".
For a small-scale operation, perhaps this is okay. For anything larger or more vital, I wouldn't trust this method of operation, as it will at some point result in emails to legitimate addresses not being delivered. That's simply unacceptable for many applications.
Anyone who deals with forms of this nature will have seen this firsthand, and with enough frequency to cause trouble with mail relay as the person above has described. It's a real problem.
What's more common, someone accidentally typing "foo:ar@gmail.com" or "foonar@gmail.com"? Both would bounce and cause mails server issues.
If someone was being malicious, they can do it with an unregistered address that passes any validation you throw at it.
Still, sending an email beats regex validation any day for determining it's a real, working address.
For example, I regularly receive email for an individual who has a nearly identical email address to my own, but at ymail.com instead of gmail.com. Any system that tries to guess at domain misspellings is going to catch ymail.com and think "Ah ha, they meant to type gmail.com, I'll correct that for them!" Viola, their email is sent to me. Again, this is not a theory, it happens to this poor guy all the time.
user@ua (.ua = Ukraine)
user@km (.km = Comoros)
user@as (.as = American Samoa)
(and many more)
Because these ccTLDs have MX or A records at the top level, pointing to real MTAs. (RFCs say you should not have MX records at the top level, but many ccTLDs do it.)No, you're excluding the set of emails that are entered incorrectly and thus are not valid. The result for those is not the same as if the UI included a simple test (such as your "ambitious" example).
1) Without UI validation:
- 1.1) Email address entered correctly -> activation email sent
- 1.2) Email address entered incorrectly but forms a valid address -> activation email sent to wrong address
- 1.3) Email address entered incorrectly but doesn't form a valid address -> activation email not sent
2) With UI validation:
- 2.1) Email address entered correctly -> activation email sent
- 2.2) Email address entered incorrectly but forms a valid address -> activation email sent to wrong address
- 2.3) Email address entered incorrectly but doesn't form a valid address -> user warned
-- 2.3.1) Email address re-entered correctly -> activation email sent
-- 2.3.2) other states
In the 1.3 case all of the activation emails fail to be sent. In the 2.3.1 case activation emails are sent that otherwise wouldn't be.
2.1a) Email address entered correctly -> validation fails, no mail sent
This prevents activation mails that would be sent without validation (or validation against a regex like [^@]+@[^@]+).
There is no technical way to verify that an email address works other than sending it a message.
And if someone doesn't want to give you a real email address, they're just going to enter bogus@fake.com to get past the validation.
1. Use regexes for client-side validation to catch typos and warn the user against potential problems without having to round-trip to the server
2. Check DNS records on the server side and send a confirmation mail
The client-side regular expression can be as simple as /@/, but something more complex like
/^("(\\"|[^"])*"|[^@\s]+)@([A-Za-z0-9-]+\.)*[A-Za-z0-9-]+$/
is fine - even if you mess up the regex, that's not a big deal as long as you allow the user to send the form anyway, probably after asking if he really knows what he's doing...This actually strengthens one of the points I was trying to make: the need to fail gracefully. The application I took the snippet from (which has been retired some years ago) would have accepted IDNs after asking the user for confirmation.
Internationalised Domain Names are going to REALLY screw over some regexes, given how poorly understood Unicode is. Does Ruby even have Unicode regex support yet? I don't program in it so I'm unfamiliar with the state of the art.
On my pet topic of Unicode, I especially enjoy the use of hidden form fields to reverse-engineer the character encoding certain browsers ACTUALLY send on submission rather that what your code hoped for...
Good luck convincing management to Do The Right Thing here.
Alternate option: Validate it with a simple regex. If it fails ask the user "Are you sure this email is correct?" If they say yes, then allow it even if it fails validation.
When you just want to store the email address without acting on it (e.g., think a landing page for an app that hasn't been released), the best you can do is validate using a regex since sending them a confirmation email would provide no value to the user.
How do you know you can trust a random email validator you found via Google? Especially if apparently the rationale is to use a googled one because they are so complicated nobody can really understand them?
That advice seems bad to me. Perhaps it is not necessary to validate and just emailing is sufficient (as the article advises). In that case perhaps downgrading the validation to a warning might be helpful, though ("the email you entered looks weird, please double check").
Oddball email addresses probably don't last long anyway as their owners quickly realize they can't use them to sign up for things and then switch to something more conventional.
It was a dumb thing to say when he said it, and it still is. Granted, REs aren't appropriate for everything, but for some problems they're the right solution.
>So eschew your fancy regular expressions already. If you really want to do checking of email addresses right on the signup page, include a confirmation field so they have to type it twice.
No. Just no. I hate you and everyone else who thinks this is a good idea. Don't make me type things twice - the point of computers is to take some of the grunt work out of life, not to add more.
There's nothing wrong with checking an email address for validity, and there's nothing wrong with using a RE provided it's correct. You'll miss a whole bunch of typos, but since you're going to send out a "click to activate" mail anyway it doesn't matter, and the ones you catch will save the user a bit of time.
cat file|program1|program2
instead of program1 file|program2
It is futile. There are so many examples of programmers just doing mindlessly stupid things, often because "everyone else is doing it" or they read some "howto" they found somewhere, or they are using some library written by someone else.How many times do you think people use extended regular expressions and backtracking when it's entirely unnecessary? They often have no idea that there is even a simpler way that will work (in some cases it might be faster). They think more complexity is actually making things "easier". Must have PCRE. Why? "Because I can't get basic regex to do what I want."
Let 'em enjoy their complex regex. Until there's a problem and they have to try to decipher what they heck it's actually doing.
PEG is fine. Lua has a good PEG library.
Still, a good handle on basic regex will take you a long way.
It's also easier to change the first command in the pipeline without having to step over the input argument, again more consistent with the rest of the pipeline.
To avoid the cat, one can write
< file program1 | program2
but in practice, cat adds no noticeable overhead.As for regexes, I personally find using POSIX regular expressions to be a bit like using vi after becoming familiar with vim. You can get by, but it's crap and there's a reason why people came up with something better. Of course, using complicated features of any language or toolkit without understanding how they work is dumb, but that's not a reason to go back to the 1980s.
http://www.useit.com/alertbox/designmistakes.html
Stylish rewrite FTFW.
Personally I test for @ and . with any characters surrounding.
Also, I learned that the local part of the address (the name) can contain pretty much anything, including '@'
So, how would your validation handle my hypothetical, valid email address "@foo"@bar?
As I said, we are not looking for RFC compliance, but rather user error. Missing a dot is user error in 100% of cases in a web application, unless you are installed in and sending mail in an intranet.
As unlikely as an email with @ in the username is, the regex would still match (something like /.+@.+\..+/.
I'm still not sure which approach I prefer, but having been thwarted by zealous validation in the past I lean towards this double-check-on-weird-shit-then-send-mail system.
I will never have enough domain specific knowledge to reject a given email address with absolute certainty. That is how much fun that RFC is.
OTOH, most languages have proven, stable libraries for validating e-mail addresses (e.g. Mail::RFC822::Address for Perl).
one implementation is in lelp - http://www.acooke.org/lepl/rfc3696.html - but that package is no longer maintained (i know, because i wrote it).
i don't know of any other implementation. but that's the right way to do it. imho.
Also, no one has mentioned using DNS. For example, extract stuff before and after the @. Check the domain looks like one and does it have an MX record? Is the local name malicious? Send activation email. Large services should use some machine learning for common mistakes and warn the user. (grandma@aol may be one common error.)
Some services have two input fields for an e-mail address. The second is to verify for typos. After that, just send the e-mail, already. If it fails you can delete the user entry from your database and print out something in the likes of "Who types their e-mail wrong two times?".
I can see both sides though, on one hand you don't want someone accidently entering an invalid email and then never getting their email confirmation...
However you still shouldn't be writing your own unless you're writing an email validation module of some kind. Laziness is a virtue after all.
Try http://news.ycombinator.com./ ;-)
(^[-!#$%&'*+/=?^_`{}|~0-9A-Z]+(\.[-!#$%&'*+/=?^_`{}|~0-9A-Z]+)*|^""([\001-\010\013\014\016-\037!#-\[\]-\177]|\\[\001-011\013\014\016-\177])*"")@(?:[A-Z0-9-]+\.)+[A-Z]{2,6}$"my@scary$doublequoted.address"@example.com
is a valid address.
http://fightingforalostcause.net/misc/2006/compare-email-reg...
http://isemail.info/_system/is_email/test/?all
For example: !#$%&`*+/=?^`{|}~@iana.org is valid.
Here's an essay on the subject: