[1] https://stackoverflow.com/questions/201323/using-a-regular-e...
[1] https://stackoverflow.com/questions/201323/using-a-regular-e...
RFC 822 describes how messages are encoded when email servers talk to each other. It isn't really about email address validation and is not intended to be used to validate a form field on some registration page.
Unless you're writing an MTA or similar piece of infrastructure there is no reason you should be using the RFC grammar. Even if implementation were easy, it probably isn't what you want. For example, the spec permits inline comments but that's a nonsensical thing to have in the middle of an address you typed into an HTML form. Email addresses entered on a web form should be rejected if they contain comments, IMHO.
I think what most developers really want to know is something like: Can this given email address receive messages? Or: Does this given address actually belong to this user? Well, the only way to test that is to send it a message. At best, regex validation might warn you earlier that a given address couldn't possibly work because it's so obviously malformed. But you can't validate your way into getting people to enter their real email address if they don't want to or if they don't know what it is. If your intent is really just to help catch typos and mistakes, you'd be much better off looking to something like mailcheck [0] which will flag common typos like "foo@hotnail.com" even if they result in valid looking addresses.
The actual standard used by e-mail servers talking to each other is 5321 (which you might still know 821 ;P), the standard for SMTP: this protocol actually has a different way to write comments and escape characters, as it is embedded into a different structure (which I think decimates the arguments people tend to make that you should validate comments).
Years ago I was working quite in earnest on an e-mail server suite, and at the time I was extremely deep in the various standards, and wrote a comment that goes into somewhat more depth on the semantics of e-mail verification. The example I was really happy with is the notion that you would never ask your user to HTML escape their username or password ;P.
Our company won't accept weird emails for free trial sign ups. We should be nudging users towards good behavior.
Of course, if your email address is provided at the corp level, then you don't have as many options.
Non-standard emails can be used as tools for phishing is another reason why we should not validate them.
For example: example@localhost is valid but is refused by most validators.
Why? I use several addresses that start something like myfirstname.mylastname@... . A comment at the start could make it much easier for me to use browser autofill. (Arguably that's "really" a browser UX issue, but we work with the tools we have)
/^[a-zA-Z0-9.!#$%&’*+/=?^_`{|}~-]+@[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)*$/
Since this is an official W3C doc, I see no reason why people shouldn't use this.Edit: There is also a version by WHATWG[2] here:
/^[a-zA-Z0-9.!#$%&'*+\/=?^_`{|}~-]+@[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?)*$/
It apparently does a more thorough validation than the W3C one in the domain part, but the difference between the two is not apparent in practice.[1]: http://www.w3.org/TR/html-markup/input.email.html [2]: https://html.spec.whatwg.org/multipage/forms.html#e-mail-sta...
[1] http://referencesource.microsoft.com/#System.ComponentModel....
I would hate to be the programmer who had to debug that regex.
Behold: http://www.ex-parrot.com/pdw/Mail-RFC822-Address.html
Well, as far as authority/canon goes, it's typically dictated by the IETF and not W3. And, sure, they could cooperate with one another but, really -- if there's an authority on the protocols that describe email (SMTP, POP, IMAP, etc) -- it shouldn't be the World Wide Web Consortium.
That said, the IETF does tend to draft RFCs that reflect actual implementations (at least their intended design), but since they often bias towards interoperability, it's unlikely they'd narrow the scope of the email address grammar.
Ditto for internationalized TLDs like something@taiwan.台灣).
I'm not sure if this is really that bad, but I don't like it.
Otherwise you're just asking for trouble when the same user who signed up with an email address at maré-design.fr later tries to reset their password with an email address at xn--mar-design-d7a.fr. Sorry, you don't seem to have an account with us.
The same thing happens with URLs. Some browsers send the Host: header in punycode but send the referer in UTF-8. Who knows how they encode CORS headers and all the other newfangled stuff that contains bits and pieces of URLs. You have to consistently convert one to the other before using any of them.
IDNs are a mess.
I don't think <input type="email"> does that though. Maybe it should, but users might be surprised to see their address automatically turn into some ugly unreadable xn--whatver-d7a domain. After form submission then yes of course any sane process will convert them to a canonical form.
I know there were good reasons to use punycode instead of UTF8 for IDNs, but it sure is a mess.
Email addresses do not need escape sequences.
That way if you don't mind letting through double periods or domain segments having more than 63 characters, you can validate an email with two character classes separated by an @. The basic regex looks like [xx]+@[yy]+
You know that's exactly what "valid" means, when talking about something standardized, right?
In HTML there's no reason to use " outside of an attribute value. Should browsers reject it?
Also W3C produced HTML specs are hardly gospel about email-related things (but I don't know whether there's anything wrong in this case).