This list may help to level the playfield between domain grabbers and legitimate domain users.
If you've got a script searching for available domains you don't really want it wasting time searching dictionary words on *.accident-investigation.aero as they'll probably all flag legit, meaning if you use the DNS lookup then whois method you've wasted a lot of time searching for domains you can't have.
> Please don't.
He is correct. Seriously. Don't.
The right way to validate an email address, if that's something you need to do, is by sending it a message with a link and a response code, and seeing whether the user clicks the link or enters the code. This is the only right way. There are many wrong ways. They make people unhappy. They will make you unhappy. Do not use them.
You cannot validate email addresses for shape and form. The sole invariant is that there be at least three characters one of which is @. Everything beyond that is in the hands of the DNS and the domain's MTA. You cannot predict every fashion in which they will behave. You should not try to.
Have you seen the regex that correctly matches every variant of email address form described in RFC 822? It is five kilobytes long. "Ah," you may now think, "I can use that!" You should not. There are new standards with new variations. The regex is incomplete. It will be incomplete forever. It is a five-kilobyte Perl regex. No one will ever understand it well enough to extend it.
Just send a link and a code. Ask the user to click on the link or give you the code. When you have received the click or the code, you know the email is valid. When you have not, assume it is not. This is the method that works. It is the only method that works. Use this method and be happy. Use any other and be sad. Which you prefer is up to you.
I used to agree with this viewpoint, but now I'm having some doubts.
Suppose I enter the email address "geofft@example.net@example.com". It validates according to your rule. Where does it go?
Suppose, furthermore, that I upgrade my servers and they start parsing it differently. They're allowed to do that, right? Maybe my old MTA sent a confirmation email to example.net, and the new one is now sending emails to a user named "geofft@example.net" at example.com. Haven't I just done something very wrong by allowing a user to trick me into sending emails to an unvalidated address? Isn't this both a violation of the Postel principle ("conservative in what you send to others") and a security hole?
It seems like it would be better for my application to parse the email address and validate what it's doing, and store only unambiguously-parseable addresses.
Or maybe they're not allowed to parse it differently. Maybe there's a single consistent way to parse geofft@example.net@example.com, and every single email application I might use will get it right. What is this esoteric lore that email applications know that my application cannot? If they can parse it, can't I?
Slight word charge on the rule. "..one and only one of which is @..." with characters within quotations ("") not being counted. Quotations need to be considered because this is a valid email address: [0]
`"very.(),:;<>[]\".VERY.\"very@\\ \"very\".unusual"@strange.example.com`
>Suppose, furthermore, that I upgrade my servers and they start parsing it differently. They're allowed to do that, right? Maybe my old MTA sent a confirmation email to example.net, and the new one is now sending emails to a user named "geofft@example.net" at example.com.
If you upgrade your servers and they parse incorrectly by inserting quotations where there were no quotations entered, then the parsing is bugged. Although the validation was bugged to begin with for allowing two delimiters in an email address.
Furthermore, if anyone uses an email that is so heavily eccentric just because it is "technically valid" I'm sure they don't expect their email to work most of the time.
I will fully agree with the claim that a regex is the wrong way to validity-check an email, and I will easily believe that implementing the check you describe takes 5 kilobytes of regex. But a rule like "Exactly one un-quoted @ sign" or "Exactly one un-quoted @ sign, and the string on the right needs to be a well-formed domain name" or something is pretty simple in normal code.
"Well-formed domain name" is an existing concept: one or more labels separated by dots, each of which contains only ASCII letters, digits, or hyphens, cannot start or end with a hyphen, and cannot be more than 63 characters long, and no more than 253 characters total. Again, not something I'd do with a regex, but a small number of lines of code.
Although you have me curious how you would check with code without using regex.
No. Read the rule again.
That's not the rule that 'Freak_NL posted, which is why I read your rule as "at least one of which is @".
And, besides, as others have pointed out, "exactly one of which is @" is incorrect. Valid email addresses can have multiple @ characters.
Your opinion might have a lot of reasons to support it, but it does not seem to be consistent with the position you were advocating. I am interested in not validating email addresses, because of your convincing argument that I should not. And now you want me to validate them.
The most sensible approach (in my opinion) is to validate with this minimal regex:
/.+@.+/
Essentially just requiring an at-mark between two other bits of text. After that send a confirmation email to see if it works. That is pretty much fool-proof and low maintenance.If you absolutely must help the user in preventing any common mistakes, use a library that provides hints in the UI for those presumed errors without actually invalidating the input. There are JavaScript libraries that do this fairly well.
I've been using this regex to make sure there actually is text before and after the @
/.+@.+/ /^[^@\s]+@[^@\s]+$/
Filtering out emails with multiple @'s and whitespace seems like an obvious win to help prevent typos and copy-paste errors.Once you get that into a regex format, remember that a quoted dot is different from an unqoted dot, and don't forget that you handle comments within email addesses correctly.
There are plenty of reasons why people say not to use regex to validate email addresses.
Wikipedia <https://en.wikipedia.org/wiki/Email_address#Valid_email_addr... > gives the following example:
"very.(),:;<>[]\".VERY.\"very@\\ \"very\".unusual"@strange.example.comThere are regexes out there that capture all the complexity of valid email addresses today (but who knows if they'll work with, say, next year's additions to the top-level domains?), and you can copy them from StackOverflow if you really want them; but why bother?
Nevermind. `/.+@.+/` it is.