2. Make sure there is at least one '.'
3. Make sure the entire thing is at least 4 characters long (@, . and two other characters)
4. Resist the temptation for something smarter
5. Send an email with unique link to verify
2. Make sure there is at least one '.'
3. Make sure the entire thing is at least 4 characters long (@, . and two other characters)
4. Resist the temptation for something smarter
5. Send an email with unique link to verify
I worked on SaaS product with a largely non-technical audience, and we had a frequent issue with people mistyping their email addresses.
We tried several things. Turned out that both confirmation email and asking to type email twice hurt conversion rates badly (in our case - all audiences are different).
However, checking email for potential typos worked really well. We had a small set of rules:
1. Domain is very close to a popular email provider (@gnail.com, or @yaho.com, etc.).
2. Email contains a fragment very close to user name: pol@rodgers.tld for Poul Rodgers.
3. We had universities as customers, and a lot of students would enter "name@university-domain.com" instead of "name@university-domain.edu". We had a special check for it.
Overall, a couple lines of JavaScript helped us to get rid of 97% of mistyped email addresses.
This will fail on sending an email to an IPv6 address, which has no '.'.
I know, it is nitpicking and especially non-tech would never send an email to an IP address. ;o)
A few TLDs have MX records. There's no reason to reject an address like, say, "postmaster@ws" - it's a perfectly valid address that could be actually working.
And why bother validating anything beyond the fact string's non-empty and there's "@" character? Shoot an email, if they receive it — it's a valid address (no matter how weird it may look), if they don't — well, it's not like something bad happened.
I have built quite a lot of data capture forms, and while no or limited regex increased number of entries, a more complicated regex combined with checking MX records improved conversion rate, because there were fewer junk entries. I used `\b[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,4}\b` (taken from regular-expressions.info).
It has a few false positives and a few false negatives, but overall, it optimises conversion rates, which is what I was being paid for.
YMMV.
That should definitely be {2,}:
* there are a bunch of gTLD with more than 4 characters: http://en.wikipedia.org/wiki/List_of_Internet_top-level_doma... even ignoring geoTLD and brandTLD
* ccIDN are all more than 4 characters since the ACE prefix ("xn--" for IDNA) is already 4 characters all on its own