The “Chad” bug
plus.google.com
plus.google.com
It reminds me of a hackathon I attended where a food ordering startup (I forget the name, but they were chosen to feed us dinner that night) had a similar bug, which baffled me beyond belief. Without going into crazy detail about my password, it typically follows a certain pattern but is never the same across websites. For some reason, the website kept saying my password was invalid. It met all the password requirements that the website asked for (length, capital letter, etc.).
I forget the exact details, but it ended up being the exact location of a capital letter, the location of a number, or some combination of both. I could never figure out how a bug like that could even be coded up. My best guess is that it was some poorly-formed regex.
> Some people, when confronted with a problem, think "I know, I'll use regular expressions." Now they have two problems.
If you're doing regex or any other text manipulation on user input when you ask them to set a password, you're doing it wrong.
I'm not entirely opposed to requiring a minimum length, but imposing max lengths / character class rules / etc ends up hurting people who want to pick strong passwords more than it helps people who would pick weak ones (enforcing character classes just gets us lots of password1A! and similar)
I think that this is a good balance between security for short passwords, while still allowing ridiculously long ones (pass-phrases).
[1] http://arstechnica.com/security/2014/04/stanfords-password-p...
8 random upper/lower/digit/symbol characters are equal to 9 random mixed-case letters. Not 16.
Their cutoffs for different mixes are 8, 12, 16, 20. Realistic cutoffs would look more like 10, 11, 11, 14.
Even worse, they encourage counting the individual letters in words. Never do that. Random words are only as good as two random characters.
$ dd if=/dev/urandom bs=32 count=1 | xxd -p
That nearly every site on the Internet will refer to that output as "not strong enough" and instead suggest P@ssword1 as a better alternative definitely speaks to the issue.I have implemented that on a site. It wouldn't accept the top 10,000 most common passwords. (or was it 1k).
Otherwise ~1% of your users will pick "password" as a password.
It's now in django[0]. And it's a hoot to read...
[0] https://github.com/django/django/blob/master/django/contrib/...
I would however suggest trimming trailing spaces.
Also, for a nicer user experience try the password twice: As is, and with case reversed. This lets people login even with capslock on and has little impact on security (i.e. don't be case insensitive! just case reversed.)
I do like the idea of warning the user upon creation.
For instance, the German sharp s (ß) has an asymmetrical casemapping.
From the Unicode standard[0]:
>The German sharp s character has several complications in case mapping. Not only does its uppercase mapping expand in length, but its default case-pairings are asymmetrical. The default case mapping operations follow standard German orthography, which uses the string “SS” as the regular uppercase mapping for U+00DF ß latin small letter sharp s. In contrast, the alternate, single character uppercase form, U+1E9E latin capital letter sharp s, is intended for typographical representations of signage and uppercase titles, and in other environments where users require the sharp s to be preserved in uppercase. Overall, such usage is uncommon. Thus, when using the default Unicode casing operations, capital sharp s will lowercase to small sharp s, but not vice versa: small sharp s uppercases to “SS”, as shown in Figure 5-16. A tailored casing operation is needed in circumstances requiring small sharp s to uppercase to capital sharp s.
[0] http://www.unicode.org/versions/Unicode7.0.0/ch05.pdf#G21180
As they say "so don't do that".
This is about capslock, not lettercase in general. Only switch characters that change with capslock on the keyboard.
And even if the mapping is not perfect - so what? The worst that will happen is nothing.
In that case you'll want to pre-hash your BCrypt-encoded passwords, otherwise you'll be imposing a silent 72 character limit.
Remember to encode the hashes, since BCrypt also silently truncates after NULL bytes.
Or just use PBKDF2 or scrypt, neither of which imposes an artificial length limitation on passwords.
Not that it really matters, since a 72-character password will have hundreds of bits of entropy as long as the alphabet is more than two characters; it's overkill.
Or XKCD style passphrases. Using the Oxford 3000 dictionary gets you about 1 bit of entropy per character. If I have my password manager generate 128 bit passphrases using them for ease of use in the real world, plain bcrypt will quietly reduce their strength by about 17 orders of magnitude. If nothing else that's obnoxious.
And that's ignoring the soft 55 byte limit both the bcrypt spec and the original scrypt paper mention. If we're looking at less than one bit per byte that's starting to get into worryingly weak territory.
Either way, it's all solved by stuffing a (e.g.) base64 SHA384 in there. Bam, every input bit equally effects the output hash regardless of length and position, and the entire thing even becomes NULL byte safe.
[1] https://stackoverflow.com/questions/201323/using-a-regular-e...
/^[a-zA-Z0-9.!#$%&’*+/=?^_`{|}~-]+@[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)*$/
Since this is an official W3C doc, I see no reason why people shouldn't use this.Edit: There is also a version by WHATWG[2] here:
/^[a-zA-Z0-9.!#$%&'*+\/=?^_`{|}~-]+@[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?)*$/
It apparently does a more thorough validation than the W3C one in the domain part, but the difference between the two is not apparent in practice.[1]: http://www.w3.org/TR/html-markup/input.email.html [2]: https://html.spec.whatwg.org/multipage/forms.html#e-mail-sta...
Well, as far as authority/canon goes, it's typically dictated by the IETF and not W3. And, sure, they could cooperate with one another but, really -- if there's an authority on the protocols that describe email (SMTP, POP, IMAP, etc) -- it shouldn't be the World Wide Web Consortium.
That said, the IETF does tend to draft RFCs that reflect actual implementations (at least their intended design), but since they often bias towards interoperability, it's unlikely they'd narrow the scope of the email address grammar.
Email addresses do not need escape sequences.
That way if you don't mind letting through double periods or domain segments having more than 63 characters, you can validate an email with two character classes separated by an @. The basic regex looks like [xx]+@[yy]+
You know that's exactly what "valid" means, when talking about something standardized, right?
In HTML there's no reason to use " outside of an attribute value. Should browsers reject it?
Also W3C produced HTML specs are hardly gospel about email-related things (but I don't know whether there's anything wrong in this case).
[1] http://referencesource.microsoft.com/#System.ComponentModel....
I would hate to be the programmer who had to debug that regex.
Behold: http://www.ex-parrot.com/pdw/Mail-RFC822-Address.html
Ditto for internationalized TLDs like something@taiwan.台灣).
I'm not sure if this is really that bad, but I don't like it.
Otherwise you're just asking for trouble when the same user who signed up with an email address at maré-design.fr later tries to reset their password with an email address at xn--mar-design-d7a.fr. Sorry, you don't seem to have an account with us.
The same thing happens with URLs. Some browsers send the Host: header in punycode but send the referer in UTF-8. Who knows how they encode CORS headers and all the other newfangled stuff that contains bits and pieces of URLs. You have to consistently convert one to the other before using any of them.
IDNs are a mess.
I don't think <input type="email"> does that though. Maybe it should, but users might be surprised to see their address automatically turn into some ugly unreadable xn--whatver-d7a domain. After form submission then yes of course any sane process will convert them to a canonical form.
I know there were good reasons to use punycode instead of UTF8 for IDNs, but it sure is a mess.
RFC 822 describes how messages are encoded when email servers talk to each other. It isn't really about email address validation and is not intended to be used to validate a form field on some registration page.
Unless you're writing an MTA or similar piece of infrastructure there is no reason you should be using the RFC grammar. Even if implementation were easy, it probably isn't what you want. For example, the spec permits inline comments but that's a nonsensical thing to have in the middle of an address you typed into an HTML form. Email addresses entered on a web form should be rejected if they contain comments, IMHO.
I think what most developers really want to know is something like: Can this given email address receive messages? Or: Does this given address actually belong to this user? Well, the only way to test that is to send it a message. At best, regex validation might warn you earlier that a given address couldn't possibly work because it's so obviously malformed. But you can't validate your way into getting people to enter their real email address if they don't want to or if they don't know what it is. If your intent is really just to help catch typos and mistakes, you'd be much better off looking to something like mailcheck [0] which will flag common typos like "foo@hotnail.com" even if they result in valid looking addresses.
The actual standard used by e-mail servers talking to each other is 5321 (which you might still know 821 ;P), the standard for SMTP: this protocol actually has a different way to write comments and escape characters, as it is embedded into a different structure (which I think decimates the arguments people tend to make that you should validate comments).
Years ago I was working quite in earnest on an e-mail server suite, and at the time I was extremely deep in the various standards, and wrote a comment that goes into somewhat more depth on the semantics of e-mail verification. The example I was really happy with is the notion that you would never ask your user to HTML escape their username or password ;P.
For example: example@localhost is valid but is refused by most validators.
Why? I use several addresses that start something like myfirstname.mylastname@... . A comment at the start could make it much easier for me to use browser autofill. (Arguably that's "really" a browser UX issue, but we work with the tools we have)
Our company won't accept weird emails for free trial sign ups. We should be nudging users towards good behavior.
Of course, if your email address is provided at the corp level, then you don't have as many options.
Non-standard emails can be used as tools for phishing is another reason why we should not validate them.
Yes. My favorite was a Coyote Point load balancer bug. If the last character of the HTTP options is "m", the connection will not get past the load balancer.[1] I found this because a web crawler was having trouble with one site. Fortunately, I knew someone with their own Coyote Point load balancer, and was able to establish that the connection went into the load balancer and never came out.
The load balancer has a big file of rules which contain regular expressions. Somewhere, I think there's a "\m" where they meant "\n". Reporting this to the vendor, along with a Python program to demonstrate the problem, was of course futile; they suggested "upgrading the software". I demonstrated that the bug existed on their own load balancer on their own site. I finally added a completely useless field to the HTTP header so that the last character was not "m".
[1] https://www.webmasterworld.com/webmaster_hardware/3312997.ht...
(Current web crawler problem: sites that won't let you read their robots.txt file if they don't like your user-agent string.)
That's hilarious. So do you use borrow a browser's user-agent or do you ignore the robots.txt?
That would get them to fix the issue pretty quickly :-)
More likely the site is trying to serve custom versions of `robots.txt` to different bots, with good intentions, and the code is buggy.
they told you this because they probably fixed the bug in a newer release. what response were you expecting, exactly? to dive in with a hex editor?
for what it's worth we run into the same issues with our network devices, it's a pain in the ass, and that's why we're shifting to SDN.
"You've obviously put a lot of time in to this. Thanks. We actually had found that bug 4 months ago in R1892, and it's patch in R1899. Please upgrade, or let us know that you're, in fact, using R1899 and still experiencing the issue (it's issue #847 on our bug tracker at foobugz.xyz)"
Or maybe...
"You've obviously put a lot of time in to this. Thanks. We actually had found that bug 4 months ago in R1892, and we're planning on a maintenance release next quarter. If you'd like to test out our upgrade version, please contact jaz@balanceco.com and let them know you want to be in the testing group - he'll get you the appropriate documentation and code. Thanks for helping fix and test these issues!"
"Update your software" is often a BS answer that can introduce many other issues or breakages in production systems. Without a clean "downgrade" path (which many companies don't provide) you run a risk of introducing more problems.
in other words, what i'm saying is to expect that kind of response is naive, and that you should seek alternatives.
And then putting it into the program without testing it properly
So, sorry, the issue is not regexes, but people just going for it at an "trial and error" fashion (and sometimes just trial)
That, and for the love of God please comment any non-trivial regexp. Either like this, with 'x' option:
preg_match('/^
.* # Match any number of characters...
(?=.{6,}) # ... AND match at least 6 characters (lookahead) ...
(?=.*\d) # ... AND match one digit after any number of characters (lookahead) ...
(?=.*[a-zA-Z]) # ... AND match one letter after any number of characters (lookahead) ...
.* # ... AND allow any number of characters later.
$/x', $password);
... or just with normal programming language comments and stitching regexp from multiple strings in multiple lines. Also give some semantic meaning to groups if you use them, e.g. tag them with constants so that your code isn't full of stuff "result.get(3)", which makes you waste time on trying to recall what was that group 3 in the code from last month.I know it's pretty much software engineering 101. It's the basics of basics. But from my experience, even the brightest of engineers in most serious projects suddenly forget how to write code when they touch regular expressions.
[0] - https://www.debuggex.com
If you're forbidding items beginning with numbers, just have a test try passing '1a' and failing the match
Great quote. However, I think regexes got a bad reputation just because of the way people use them. In essence they are a pretty reliable way of parsing because the parsing engine is well tested. But the expression should be kept as simple as possible and developer should avoid using any nonstandard / nonexplicit extensions. I even avoid using \w because, well, what IS a word character? I am sure it is defined somewhere... but I'll always use explicit form (like "[a-zA-Z]" when I want ASCII chars) instead.
Anyway, if you use the form as used in the regex puzzle [0], you'll be fine. As long as you use regex only for what it was meant for, of course [1]...
[0] https://news.ycombinator.com/item?id=10787509
[1] http://blog.codinghorror.com/parsing-html-the-cthulhu-way/
Yeah sure. I think I heard that quote before... just about a million times?
People expect regex to be an easy-to-use tool. Well it's not, and it's a foot gun if you don't take your time to learn it right. But no, people hack up some expressions, hit their feet and blame... the tool of course, not themselves.
Just learn it right, it's a great tool if you know how to make it work for you :)
As a workaround, I bet you can use CHAD@... or chad+blah@...
(Hangouts Dialer still does not see him if saved as CHAD@. It sees him when saved as chad+bla@ but it's annoying because then his email is wrong in my contact list as his email provider does not support + aliases.)
Think about this: one would hope that if you use characters in that field which would normally need to be escaped if used in a MIME message, and that e-mail address were to end up in a MIME message, that the e-mail client would get the unescaped e-mail address from the database and would then escape it correctly--and by "correctly", we mean a different way of escaping it for MIME vs. escaping it for SMTP--the same as we would expect its usage in an HTML page, an argument to a shell script, or a value in an SQL statement, to also be escaped for each specific purpose.
Thanks for the information!
The only way I've found of getting anything resolved is to forward issues to a friend inside the company, or hope that you can write a blog post which gets enough attention.
I get that filtering and testing millions of random bug reports from all corners of the Internet is hard - but it's a problem which Google desperately needs to solve if it wants to retain the trust of its users.
> It's a pity there's no way to report bugs like this to Google.
This is a popular myth.
General instructions are here: https://www.google.com/tools/feedback/intl/en/
In this particular case, it's an android app, so what you do is tap on the hamburger menu, hit "help and feedback", then "send feedback".
Looking at Android (OS) issues - https://code.google.com/p/android/issues/list?can=2&q=&sort=... - it's clear that the majority of bugs are ignored. Even when they're well described and affect multiple users / devices.
> Looking at Android (OS) issues - https://code.google.com/p/android/issues/list?can=2&q=&sort=.... - it's clear that the majority of bugs are ignored.
A quick glance at that page appears to disprove this claim. If you flip it to display "all issues", there are 206689 bugs in that tracker at the moment, of which 44583 are open. That tells you that 79% of all bugs filed have been closed - so, at least 79% of bugs were not ignored.
Note that this tracker is for the operating system only, and does not include any of the Google apps that the feedback system covers.
Which goes back to my original point about customer trust. If you know I've reported a bug, why would you deliberately not tell me that it has been fixed?
> so, at least 79% of bugs were not ignored
Well, take a look at some of the ones which have been closed - https://www.reddit.com/r/androiddev/comments/2on1fe/google_c... and https://news.ycombinator.com/item?id=8803118
I know a good many people who work in Google - they're all smart and dedicated. But there's something about the corporate culture which imposes a "don't listen to external feedback" mindset.
It's your OS and they're your apps - you can do what you like with them. But don't be surprised when users stop trusting you to listen to their concerns.
Sounds like a typical fallacy of end-users looking at software. Assuming that the developer deliberately is denying you a feature, rather than having simply not spent the engineering time to make it possible.
In this case, for instance, it could be that the pipeline to get from external feedback channels to Google's internal bug-trackers very one-way and its hard to go back. Or, there's a disconnect between when the fix ships and when the ticket is closed, and keeping track of the entire chain of data requires some work. Or, there's no easy distinguishing features on bugs that came externally to make them easily identifiable once fixed. Or, there's no process yet for an automated response that says the bug is fixed (should it provide more detail). Etc.
> I get that filtering and testing millions of random bug reports from all corners of the Internet is hard - but it's a problem which Google desperately needs to solve if it wants to retain the trust of its users.
> Which goes back to my original point about customer trust. If you know I've reported a bug, why would you deliberately not tell me that it has been fixed?
Another example of http://danluu.com/wat/ (discussion:https://news.ycombinator.com/item?id=10811822).
My guess is the dialler hashes some parts of the contact to get a UUID, but for this contact it happens to be outside the range the dialler can look at - perhaps off-by-one, where the dialler looks for UUIDs of 1 and above and this happens to hash to 0.
India
Georgia
Jordan
I'd guess they are all as common or more common than Chad.Continents too: I met a girl called "Africa", and "Asia" is certainly used as a name.
However I don't ever expect to meet someone called "Democratic People's Republic of Korea"
This could have simply been a race case unrelated to the filename, but it's much more amusing to speculate that it was due to hack introduced during development. I now regret not trying to reproduce it, but I was pretty frustrated after I found my document again. I did contact support, but didn't hear anything back.
This is kind of horrifying. Google being tripped up by trailing whitespace?
Took me quite awhile and many failed attempts to find that Comcast will throw an error when your requested username contains "comcast".
Moneygram will freeze payments with the word "moneygram" in the associated email. Which is great for those of us that use catch-all emails and use the email address to discern what to do with an email...
My current favorite two include when the change password form permitted longer passwords than the login page, and one where the change password form happily allowed special characters, but if there was e.g. a semicolon in the password, submitting it from the login page would throw a SQL error.
jwz's wit was a lot of fun in the early 2000s but that quote is too often used wrong.
This quote is not about regexps, it's about using wrong tools for the job. Using it without context makes it sound dumb. "Well, what if regexp is exactly the right solution for that problem?".
That's quite a sweeping statement.
Sure, complex regexps can be hard to read but state machines are just hard to read for humans in general, whatever the form.
What do you recommend for parsing simple text entries, then?
First of all, "more readable" is extremely subjective.
Second, there are a lot of different parser combinators, all with very different syntaxes.
Finally, parser combinators are readable by people familiar with them and regexps are readable by people familiar with them. Regexps are also much more widespread and approachable. And very often, writing a parser combinator to parse a simple text entry is way overkill.
There are many good reasons why regexps are so popular.
You don't have to be familiar with them to find something like:
def emailAddress = userPart ~ "@" ~ hostnamePart ^^
{(username, at, hostname) => EmailAddress(username, hostname)}
clearer than any regex. Named capture groups help a little bit but I've never seen people using them (and they don't have a consistent syntax across regex implementations either).> And very often, writing a parser combinator to parse a simple text entry is way overkill.
Disagree. They can be very much a one-liner.
First of all, what language is this? Well, I know, but you seem to forget that the parser combinator syntax varies per language. What's the equivalent syntax for Java? Or for Python? What about a language that doesn't have a parser combinator library? Or one that has several ones, all slightly different?
I think you're falling prey to the specialist fallacy: you're obviously very comfortable with parser combinators but you've forgotten how long it took you to get there and you now see them as an ultimate solution to all problems without realizing their downsides.
I very deliberately didn't mention the language or the library, because I think the snippet is readable without knowing that. Minor syntax differences between libraries matter when writing, but not when reading, and reading is more important. (And it's not like there aren't several slightly different implementations of regexes)
I'm not that committed to parser combinators - I'd be happy to consider alternatives - but anything where you a) name the things you're capturing b) can easily combine several small parsers to make a bigger parser will have a huge readability, testability and maintainability advantage over regexes.
I think many regex libraries have named captures and in most languages you can either concatenate regexes or build regexes from strings which in turn can be built from concatenation. Furthermore, I thought building regexes by concatenation of smaller pieces was a commonly recommended technique for improving readability (I have no data on whether the recommendation is commonly followed or not, of course).
And about captures inside repetition: I don't see any reason they couldn't capture a list of strings instead of a string in dynamically typed languages but in all regex libraries I'm aware of they do something that seems useless to me: they only capture the last occurrence! (They probably just overwrite the capture on each repetition.)
Best image to explain a hanging "chad": http://images1.fanpop.com/images/photos/1400000/Halloween-ho...
Example: a site let me register and login with the plus. But resetting the password was hard, I had to escape the plus to get it working.
"chad"@example.com
It seems to be supported by Exchange, gmail, and a few other MTA's I tested, and gets routed to the right place.