Perfect email regex finally found
fightingforalostcause.net
fightingforalostcause.net
If they get it wrong, intentionally or not, then they don't get their receipt, confirmation, validation link, etc. and I believe in most cases the incentive is there for them to get it right.
In the rare case where there's some incentive to circumvent the system and this has some measurable impact on a site, then more validation may be warranted. Otherwise, why worry about it?
Also: HTML5. ;)
http://tools.ietf.org/html/rfc5321#section-2.3.5
RFC 5321 section 2.3.5 specifically prohibits TLDs from receiving email. Other RFCs back that up, actually including RFC5322 (not directly, but in wire format).
That said, we see some TLDs run MX's. I think at least some portion of them sell the mail received there to spammers. Seriously.
You obviously know a lot more on this topic... but where exactly does it say in that section that TLDs are prohibited from receiving emails?
From my reading, it seems to suggest the opposite, e.g. "In the case of a top-level domain used by itself in an email address, a single string is used without any dots."
I've re-read that section many times and forgive me if I missed it — it's late! =)
People have been posting conflicting responses throughout this thread suggesting that TLDs can/cannot receive emails, and I'd really like to know which should be valid... thanks!
"@"@example.com
?
Retyping only makes sense for password field, which is obfuscated and doesn't allow copy&paste.
Multiply that by say, 130,000 people, and you are dealing with 325 people who don't receive their download, etc. and are not happy!
I think what would be really awesome is a regex that catches these common typos and warns the user immediately.
(The point being that for most users the faster approach that requires less thinking is to type it twice. Sometimes I do things the "slow" or "long" way when coding because it doesn't require a mental shift from the task at hand.)
Have you ever abandoned a sign-up because of the e-mail confirmation?
Spoken like a touch-typist with a short email address. I feel confident in claiming that annabelle.t.johnson@woodandplaster.co.uk will be looking for that copypaste button. It's not like copypaste is a new technological innovation - it's practically the most used feature of personal computing, after backspace.
http://blog.mozilla.com/faaborg/2010/03/23/visualizing-usage...
Display the email address on the submit button: http://www.userglue.com/blog/2009/09/09/solving-the-repeat-e...
Also try out the live examples: http://infinityplusone.com/experiments/email-repeat/version5
Regexes take care of the syntax. The semantics still have to be checked by a human.
Not a lot of chance of mistyping, there;
1) Send an email to the address.
2) If it succeeds, the address was valid.
BTW, the email address that doesn't work is "jon-whatever@jrock.us". The .us confuses people and the - confuses people. WTF!?
There are regexes out there that catch all valid addresses that people reasonably use (including those with dashes and ending in .us) and the fact that all kinds of incompetent developers use a homebrewn regex is no reason to lower the best case scenario to "just don't validate". Just do it right.
Oh, and don't forget to make sure that one component in your spam^W email processing chain correctly encodes unicode charaters in the domain part into punycode.
If they get it wrong, intentionally or not, then they
don't get their receipt, confirmation, validation link,
etc. and I believe in most cases the incentive is there
for them to get it right.
Well, I hope you have a large enough support team, because with any reasonable amount of customer growth, you'll soon be swamped with support emails that say "I didn't get my ..., where is it? You suck!". Validating the email address (which includes checking for common typos in domains) is a service that reduces frustration and the amount of customer support needed.You will still need a support mechanism (doesn't have to get all the way to a person) to sort out undelivered verification mails. It's just how it goes.
In particular, one of the evaluation tests used here is wrong: it requires failure-to-match on a TLD with a digit in it:
numbersInTLD@domain.c0m
In fact, IDN TLDs will have digits in them. An internet-draft is in the works to replace RFC1123's IDN-unfriendly implication that digits in TLDs are illegal:
The existing specs are in conflict, with the more recent ones (such as IDN) allowing digits. Internet authorities, including ICANN, have enabled domains with digits in TLDs; software which is far more foundational than any web-app's email validation regex has been updated.
Registration of names in some of these digited-TLDs has begun; you can visit these TLDs with your browser; your users can have functioning email addresses on these TLDs.
If your app rejects such email addresses because of slavish compliance with imprecise language in a 31-year-old RFC, you'd be the one violating prevailing standards, which are a function of more than just formal IETF RFCs.
That the Internet-Draft I referenced may soon become an RFC is just cleaning up loose ends on a change that's already happened. This final step isn't even strictly necessary for the de facto standard to have changed by consensus among practitioners. Plenty of vibrant well-understood standards never reach RFC status, nor pass through any formal standards body. The standard is ultimately what people do, not what someone once-upon-a-time decreed.
Users of http://موقع.وزارة-الاتصالات.مصر won't be pleased :)
In fact, I often wonder if punycode is a prank that got out of hand.
UTF-8, on the other hand, would have been excellent for this purpose.
I fully understand that these are not in common use, but they are part of the RFC and may be in use somewhere.
Granted, that was 17 years ago, but who's to say it's not in use somewhere?
I don't get this preoccupation with making sure addresses look valid. The ONLY way to validate an email address is to send it a message.
I find these large catch-all email regexps silly for two reasons:
1. They are hard to write, hard to understand, and hard to maintain.
2. Most importantly, they are difficult to understand for users. "You entered an invalid email". Now what? the user asks.
This is why e-mail validation should be done in steps. Here is some Rails pseudocode:
validates_format_of :email,
:with => /@/,
:message => "Needs to contain an @."
validates_format_of :email,
:with => /\.[^\.]+$/,
:message => "Has to end with .com, .org, .net, etc."
validates_format_of :email,
:with => /^.+@/,
:message => "Must have an address before the @"
validates_format_of :email,
:with => /^[^@]+@[^@]+$/,
:message => "Must be of the format 'something@something.xxx'"
Much easier to write, much easier to maintain, and much better error messages to the users.Sometimes, you aren't validating a whole string, you are searching for email addresses in a sea of text, or an arbitrarily delimited, user-entered list of contacts.
Support usernames with alphanumerics, dashes, underscores, periods, and plus signs; Require a single @; Support domains with alphanumerics, dashes, and at least one period. Screw anyone with something more complex than that. Done deal.
That is part of Lepl - http://www.acooke.org/lepl/ - and although it's implemented in a recursive decent parser, much is compiled to regular expressions for efficiency. So you get the best of all worlds: regexp efficiency; parser accuracy; standards based.
A blog post on the compilation to regexps is here - http://www.acooke.org/cute/LEPLOptimi0.html
qtext = '[^\\x0d\\x22\\x5c\\x80-\\xff]'
dtext = '[^\\x0d\\x5b-\\x5d\\x80-\\xff]'
atom = '[^\\x00-\\x20\\x22\\x28\\x29\\x2c\\x2e\\x3a-' +
'\\x3c\\x3e\\x40\\x5b-\\x5d\\x7f-\\xff]+'
quoted_pair = '\\x5c[\\x00-\\x7f]'
domain_literal = "\\x5b(?:#{dtext}|#{quoted_pair})*\\x5d"
quoted_string = "\\x22(?:#{qtext}|#{quoted_pair})*\\x22"
domain_ref = atom
sub_domain = "(?:#{domain_ref}|#{domain_literal})"
word = "(?:#{atom}|#{quoted_string})"
domain = "#{sub_domain}(?:\\x2e#{sub_domain})*"
local_part = "#{word}(?:\\x2e#{word})*"
addr_spec = "#{local_part}\\x40#{domain}"
pattern = Regexp.new "\\A#{addr_spec}\\z", nil, 'n'Seriously, I can't think of a single good reason why you would want to check whether an email address is "valid". What you should be concerned about is whether or not the address works (and usually, can/does the person who just signed up actually read and reply to email to that address).
Hypothetically, if an invalid address works (due to bugs in mail systems) -- then it works, and the only problem with accepting such an address is that the bugs might get fixed. If an address is valid, there's no reason to assume that it will work or that it belongs to the person who signed up. It isn't even a good way of detecting typos; transpose two characters in an email address and it will most likely still pass your validation.
Because it's fun.
After this I almost stopped paying attention: "It's my philosophy that it's better to accept a few invalid addresses than reject any valid ones, so I'm shooting for 0 false-positives and as few false-negatives as possible."
But then I looked at the regexps and they miss an absolutely trivial fact: valid email addresses can end in a dot. "jemfinch@supybot.com." is just as valid (more so, in fact) than "jemfinch@supybot.com".
Things don't have to be well done, nor do you have to agree with them for them to be worth consideration/stimulating.
Good luck trying to register on any site with it though :)
http://www.messagingnews.com/onmessage/ben-gross/validating-...
There's a fine line between clever and practical. Why do I have a gut feeling that this approach is way over that line?
/^[-a-z0-9~!$%^&*_=+}{\'?]+(\.[-a-z0-9~!$%^&*_=+}{\'?
]+)*@([a-z0-9_][-a-z0-9_]*(\.[-a-z0-9_]+)*\.(aero|arpa|
biz|com|coop|edu|gov|info|int|mil|museum|name|net|org|
pro|travel|mobi|[a-z][a-z])|([0-9]{1,3}\.[0-9]{1,3}\.
[0-9]{1,3}\.[0-9]{1,3}))(:[0-9]{1,5})?$/i
I mean... Why do this? Just, why? It's almost unreadable.Just write a short 20-line function which validates an email address. Use if statements. Write comments. Then verify that your algorithm does in fact handle all corner cases, just like the regexp does. (Your verifications will be in the form of short, simple unit test case functions.)
To encourage regexp abuse like that is to encourage bad programming.
Frequently because Perl's regex library is insanely fast. Faster than if statements + smaller regex / roll-your-own.
(as long as there's no look-ahead / look-behinds. It's still fast then, but custom functions can sometimes do better.)
It's 6,598 bytes long.
You could almost make a game out of finding valid addresses that are not matched, or invalid addresses that are false positives, or optimizing candidate regexps. The list could be continuously growing as new variations are discovered, and tested against previously submitted candidate regexps.
But we know a valid email address when we see it, right?
--
Edit: Nicely done, HTML5.
Another know-it-all web developer guy who thinks he got his regex right... These are valid addresses that he rejects: user@ua (.ua = Ukraine) user@km (.km = Comoros) user@ne (.ne = Niger) Many ccTLDs have MX or A records pointing to real MTAs.
Based on the HN title I thought it was going to be an article describing a post-it found on Fermat's dressing table mirror.
Instead it's a list of mostly correct regexes. As Miracle Max might have observed, mostly correct is partly imperfect, and partly imperfect is Not Perfect.
Now someone please make the perfect JSON regex decoder :-)
Regex can be easier to read if you have something do a graphical expansion for you. Otherwise, it's write-once, read-never.
What would be interesting, but I can't find with some googling: Has someone implemented a parser-generator based on the spec? The ideal would be that the parser specification looks a lot like the RFC, since then you'd have more confidence it was actually correct (and it'd be easier to maintain for future changes).
While you do that, I'll do something that doesn't involve tweezers and code.
There are plenty of ‘better’ (in the sense of ‘more powerful’) string-validation techniques. For example, lots of grammars are expressed in BNF; the languages that can be so expressed are (if I remember my Chomsky hierarchy correctly) the context-free grammars, a strictly larger class than the regular languages. The extra power comes from the fact that they have the expressive power of a finite-state automaton augmented by an (infinite) stack. (It's fair to argue that it's not ‘really’ infinite, since a computer's memory is finite; but, in that sense, real-life computers will never be Turing complete.)
(Of course, common ‘regular-expression’ libraries aren't actually regular any more, because of added features like capture groups. I don't know if they recognise all CFG's, though; I suspect not.)
Given this, why would we use regular expressions? Well, by intentionally sacrificing power, we can achieve faster matching (http://swtch.com/~rsc/regexp/regexp1.html) and, probably, lower memory usage. Sometimes this trade-off is worth it, even if it means that the match must be somewhat fuzzy; but sometimes one needs a precise match, and regexes just aren't up to the job.
Isn't that what I said?
Possibly this approach is too slow to use by itself without a regex. Also maybe there are other problems with this method I'm not aware of?
The best way to validate an email remains to be a test email apparently.
Merely looking up MX records is something you do for a _domain_ and it's no different than the DNS requests your computer makes when you go to a website.
(I know.)