Email Validation Rules
rumkin.com
rumkin.com
Additionally, lots of 3rd party mail systems have some mechanism to notify you of bounced emails, so you'll figure out which are invalid at your first attempt anyway.
/^[a-zA-Z0-9.!#$%&’*+/=?^_`{|}~-]+@[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)*$/
So if you're going to use <input type=email>, you might as well use the same check on the server side.We (Mailgun) have a free service which does all of these things: http://blog.mailgun.com/post/free-email-validation-api-for-w...
1: Does the domain part have either an MX or A record? 2: Does the address in that record accept mail? 3: Does the server that accepts mail at that address respond favorably to the local part as a RCPT TO argument?
If yes, it's a valid email address.
In particular this article is wrong that emails are always 7-bit. With SMTPUTF8 you can put whatever you want in there.
Practically speaking, the following regex -
/^[_\.0-9a-z-\+#]+@([0-9a-z][0-9a-z-]*\.)+[a-z]{2,6}$/i
coupled with a DNS check for domain name is yet to generate a single false negative in years that we've been using it across multiple projects.(edit) Do read through the thread below before getting all downvote happy.
hello@visit.melbourne.
Yes, that's a valid TLD now.If you are referring to {2,6}, then it's easy to adjust once longer domains are put to actual use. If you are referring to the trailing dot, then that's exactly what I was saying - practically, email formats that are used by real people is a much simpler subset of what's permitted by the spec.
http://newgtlds.icann.org/sites/default/files/ier/hu62oef5uc...
gTLD can also contain punnycode/unicode now, which additionally breaks your regex in another way.
--
But more generally, you just don't want to hear what I'm saying.Right now, people do NOT use email addresses with internationalized domains, they do NOT put trailing periods in domain names, etc. Once they start doing that, the regex can be adjusted to accommodate for that on first false rejection. As it exists now, the regex is tailored exactly to how emails are used now, by real humans. A negative is an indication of a bot activity or some other oddity that needs attention. Feel free to support that transparently, but the question is why would you want to do that.
This assumes that you manually verify every rejection, or that they contact you. In other words: Either you're not using this anywhere with volume that matters, or you don't know that it hasn't caught false negatives.
> A negative is an indication of a bot activity or some other oddity that needs attention
High volume of negatives is an indication of bot activity or some other oddity. A single or handful of negatives is just humans going about their business as horrible typists.
If your volume is small enough that you can treat every negative as something to investigate, you really have no basis for saying anything about how well your regex works.
Why?:
It's very bad UX to say to a human "we don't support you because we think you just might be a bot".
You should better lean on the permissive side; that would be one less thing to think about.
That kind of "validation" does not check for existing email addresses anyway, so there is not much value in trying to be strict in this area.
Add to that the fact it would fail on a legitimate top-level domain like .рф or .中国.
I'd think a grammar (http://www.ietf.org/rfc/rfc2822.txt) would work just great.
Max is 254: http://stackoverflow.com/questions/386294/what-is-the-maximu...