How not to validate email addresses
mdswanson.com
mdswanson.com
The first two can be done without requiring any additional work for the user, but people are so used to clicking verification links that they don't really mind that either.
(I do understand that it is mailchimp and not the author)
I like the "push the boundaries" thinking though! There is no reason why I couldn't add the JS library to the form at the bottom.
Minor nit: this is not anything Gmail invented. This is RFC 5233 -- subaddressing:
https://github.com/dominicsayers/isemail
The regex's floating around out there are horrible.
Validating email addresses doesn't necessarily mean that you affect the user's experience. I think of it as an opportunity to avoid losing a potential customer due to a silly mistake. One such example would be a one page sign-up site where you are trying to collect the email addresses of those interested in your offering. In this context it is important to try and catch errors. You have a visitor who wants to keep in touch with you. He or she mistypes the email address. If you don't detect it you might lose them forever.
Granted, all errors are not detectable. If someone types jeo@example.com vs. joe@example.com there's precious little you can do about it in terms of automated detection.
You can accept obviously bad email addresses, store them in your database and simply tag them as such. This is where ML or human intervention might be able to fix the problem or choose to discard it. Email list pollution can be dealt with in other ways, for example, if you use this list to reach out to prospective customers bad emails will simply bounce.
In the end what is important is to avoid losing real potential customers as much as possible. I think a little software-based verification along with giving the user the opportunity to catch the mistake is enough. All the junk easily falls though the cracks of a multi-stage filter after the fact.
Swanson: I need you to take a look at address forms. I don't want to enter my city and state any more after I've given my zip code.
PS Yes, I know that not everyone lives in the US.
PPS Yes, I've heard about some places where a single zip code serves two cities. Edge cases, there will always be one or two.
On the other hand I'd be quite happy for sites to calculate shipping based off this since the reported location is the capital city.
I've seen this handled quite well by a number of sites - you enter your post code and they'll just drop down a list of all the places it matches. Choose your town/suburb and you're done!
The ridiculous part is that they'll often ask you to enter your state as well, which you can derive from the first digit of the post code :/
https://en.wikipedia.org/wiki/Postcodes_in_Australia#Austral...
Singapore has 6.
Australia's post codes, on the other hand, are really wide. The entirety of central Sydney is all "2000" and mine covers around 6 city suburbs.
And there's a special place in hell for webdevs who force me to use their fancy javascript date picker rather than typing in a date.
If you are saying positioning the zip field first actually wastes more user time because it's so jarring, there are so many possible improvements:
- Don't change layout but populate city & state if zip is entered first. I know, hard to justify the effort.
- Populate zip from address.
- Offer city, state completions when you start typing city.
- Offer full-address completions from street address & geoip.
- Just statically populate city, state (and possibly zip?) from geoip.
"https://example.com/unsubscribe?email=me+example@example.com"
vs. "https://example.com/unsubscribe?email=me%2Bexample@example.com"
[ On the plus side, Zappos was really responsive, and fixed the issue when I reported it. ]I originally thought this was a bug. But if you think about it, MediaWiki is capable of being deployed on an internal network. An internal wiki actually could actually interact with email addresses only available in a given intranet server, and not reachable from a given TLD.
http://www.washingtonpost.com/world/national-security/nsa-co...
In the case of validation I tend to look at it as "What is the minimum I can check for to ensure that I can get the data that I need out of this form?"
Too much and it becomes arduous to sign up, too little and the app ends up trying to send an order confirmation to "Matt Swanson" instead of "matt@mdswanson.com"
Going through the error logs on our mail server, there are a lot of people out there that get their email address wrong even if you have them type it last. .cmo instead of .com, transposing letters in their name, spelling their company name wrong...
"hotnail.con" -> "hotmail.com"
user@myhostname, is a valid email address, and yet it's rejected by a lot of libraries.
My aunt swears blind that an email address without the name in double quotes and the domainy bit is not a correct email address. She types the lot out.
http://blog.mailgun.com/post/free-email-validation-api-for-w...
The best part of that was that the MMO account didn't allow me to change my email address.
I've had the netflix account for years.
[1].
(?:[a-z0-9!#$%&'+/=?^_`{|}~-]+(?:\.[a-z0-9!#$%&'+/=?^_`{|}~-]+)* | "(?:[\x01-\x08\x0b\x0c\x0e-\x1f\x21\x23-\x5b\x5d-\x7f] | \\[\x01-\x09\x0b\x0c\x0e-\x7f])") @ (?:(?:[a-z0-9](?:[a-z0-9-][a-z0-9])?\.)+[a-z0-9](?:[a-z0-9-][a-z0-9])? | \[(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3} (?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?|[a-z0-9-][a-z0-9]: (?:[\x01-\x08\x0b\x0c\x0e-\x1f\x21-\x5a\x53-\x7f] | \\[\x01-\x09\x0b\x0c\x0e-\x7f])+) \])
Python/PHP Code and explanation is here http://www.webdigi.co.uk/blog/2009/how-to-check-if-an-email-....
It was eye opening to understand the underlying SMTP protocol. There are some pitfalls too as mentioned in the article.
The RFC specs emails as a CF grammar, not a regular grammar, which is why validating all possible emails with regexes is so hard. Use a parser and call it done.
If your goal is "don't prevent a user from signing up" then why validate emails at all? Why not just accept anything, and whinge at them after you've already captured them in your system?
Like others have said most popular web app languages have some email validation built in (like php has the FILTER_VALIDATE_EMAIL filter).
Another one, using the stupid asterisk character hiding password input field. It's user-hostile, especially on mobile. No one is looking over my shoulder, and if they are I'll take care of it myself, thank you very much.
That being said, I completely agree with you – I don't think there's much validity in masking the password field except maybe when it's auto-filled by the browser.
We have tested turning off the masking on various sites we've developed and in general users tend to freak out and think the site is insecure as a result.
Sure there's better ways to build a regex then to hard code it into each method, but nevermind that, lets just accept whatever comes through the pipe into our barely tested (in production) and highly insecure frameworks, as long as it contains an '@'.
No matter what solution I propose, It's better than this "nonsolution" because it's a solution.
Not all validation needs to happen at the time of data entry.
Your "hard" regex may reject international email addresses. In fact, if your regex's input isn't converted to Punycode first, you are a fool for even attempting to use regex, because now your regex will likely fail on all IDNA inputs.
And what is your test suite going to be?
And what did you "validate" exactly? That you matched your regex? What if the e-mail isn't active, or the mailbox is full? Outlook 2013 actually has this really cool feature called MailTips that provides more advanced mailing list and e-mail address validation and warnings: http://blogs.technet.com/b/exchange/archive/2009/04/28/34073...
Suppose when you first signed up the user, they validated their email address, but now the account seems to be inactive. How do you handle that scenario? Continuous validation.
And how generally useful is your regex? What are you going to do if the email came from OCR software output, or screen scraping output? Your ERP may have the original document it was scanned from. Are you going to not store the bad e-mail address simply because you wrote some "hard" regex that rejected it? Not a straight forward question to answer, as it depends on your data model for storing addresses. You might have a column IsConfirmed.
Here is another example of "continuous validation". Validating mailing addresses. Most major e-commerce sites allow very liberal input, but scrub the data in real time or near real time, because the postal service gives discounts to companies that print "correct" address labels. "Correct" here could mean "One Post Office Square, Boston, MA, 02109" instead of "1 Post Office Sq, Boston, MA, 02109". This process is called Address Standardization, and in areas of the world with rapidly growing economies, often times Address Standardization vendors are behind, because some "streets" don't have addresses yet and aren't known to exist in any GPS system. This is common in many parts of China.
Here is another example of "continuous validation". How Google does spell checking, as compared to the "fixed validation" in Microsoft Word's spell checker.
The article argument in a nutshell is that validating email is hard, so don't bother, in fact, let users submit whatever they want including javascript. Then just check for @ and send it off to your next parser in the chain, in fact get lots of 3rd party parsers for misc features and send data to them first. spend effort fixing autocomplete so users can enter data easier that you will automatically accept. I'm sure this can only improve data quality...
I can imagine that wanting to know all the stupid shit your users submit as an email is the correct solution in certain contexts, but for a majority of cases, this article is wrong in everything that it suggests. Admittedly, there is very little context given.
Perhaps the context is "I don't care about security of my users or my services, and I will run whatever 3P code on my backend that appears to do the job of making a webpage look spiffy and easy to use. Once I have 10 Million (unverified) users, you sell your spaghetti factory and it's no longer your problem."
After all that, he recommends not letting people use software without a validated email address. Too bad he never bothers saying how he would get to that point, only how he would avoid doing to work.