Creative usernames and Spotify account hijacking
labs.spotify.com
labs.spotify.com
When a new user is created, it then gets assigned a unique user ID. The email address is assigned to that user ID.
Then, if they do a password reset on user = "bigbird", it should do the exact same lookup to find the email address.
The security bug was not really about having an improper function to do unicode translation. It was more about having different functions for the same check, simply because they were in different parts of the code.
Modular code is just so much better on all fronts, including security.
However, when the link was used, canonical_username was once again applied
So after they sent the password reset link, they called "fetchUserIdByName" again, but they passed in a username that had already been canonicalized once. Because of this bug, I wonder if password resets worked at all for users with unicode characters in their names.Another possible solution would be to assume the username given to the password reset form is already canonicalized (which would be necessarily true, as far as I understand).
EDIT: However this wouldn't solve the other bug that's been discovered here, which is that "ᴮᴵᴳᴮᴵᴿᴰ" is canonicalized differently than "BIGBIRD", thus defeating the purpose of canonicalization (for that particular case) in the first place.
DRY tells us the code for looking up a user id by user name should be written in only one place, not that the code should be called from only one place.
The mistake was assuming the name->name function was idempotent, because it wasn't.
You are right to suggest using a name->id function instead. It would not suffer the same problem because the canonical name should not be stored... it's an implementation detail!
They having nothing to do with each other.
DRY tells us the code for looking up a user id by user
name should be written in only one place, not that the
code should be called from only one place.
If you write 'canonicalize(username)' in eight different places, you are not being DRY. If you need to write 'canonicalize(username)' in eight different places, your code probably doesn't separate responsibilities properly across separate modules.As such, they have a lot to do with each other. After all, if you call code from more than one place, you are writing the calling code in more than one place. Lack of DRYness is about the fact that you are doing so. Lack af modularity is about why you need to do so.
Could the method for computing canonical usernames based on nodeprep.prepare() be salvaged? If not we would be in trouble since we use canonical usernames in various databases so that changing how to derive them in a non-backwards compatible way would be quite costly.
How else do you map username to user id?
VirtualAlloc: http://msdn.microsoft.com/en-us/library/windows/desktop/aa36...
VirtualAllocEx: http://msdn.microsoft.com/en-us/library/windows/desktop/aa36...
VirtualAllocExNuma: http://msdn.microsoft.com/en-us/library/windows/desktop/aa36...
It all started out nice and clean I'm sure, but within a few years you start to see many more than one basic interface to some things.
"In this case the two users who posted to the forum where actually rewarded with some Spotify premium months."
I'd say: Premium lifetime memberships would be better :)
They could simply store two names: One is provided by the user (verbatim), and the second is its reduction to lowercase letters and digits (canonical). For all internal logic, they could use only the canonical name, and use the verbatim name in the front-end to make the user happy.
> Lower casing has the key property of being idempotent, i.e., that applying it more than once has no effect: x.lower() == x.lower().lower(). So if a username gets passed from service to service and you want to make sure it is in canonical form you can safely apply .lower() and if it was already in canonical form there is no harm done, and it is easy to stay safe.
Apparently, they thought that it's ok to use verbatim and canonical names interchangeably, relying on idenpotence property of the XMPP function.
It looks like a huge design error.
Presumably you mean in the database? I don't see a reason to keep a copy of the lower() transformation of a string when it is incredibly cheap to transform a small string to lowercase.
What exactly is the point of that? I would just call lower() as needed, personally.
With lower(), we can expect we'll get the right transformation of string A each time. If instead, we store string A, and then store string B as A.lower() and copy it... A.lower() will always be A.lower, but it's much easier for someone to come along, screw with the database, and change B.
They need to store the verbatim username in order to know how to display the username in the UI.
They need to store the canonical username in order to efficiently know whether a given canonical username is in use.
But - when the issue here is the question of the reliability of the implementation of the canonicalisation function, having it done once in python, and then again by PG is going to be a huge issue.
Well, I still think that it is better to have two (hopefully) correct fields in the database, rather than only one. (Consistency of the two fields can be checked once in a while).
Their canocialisation function, which is a standard one, was broken by subtle changes in Python 2.5, but worked previously.
Well, in general, modular code should not make assumptions about other parts of the program (when possible).
You know, if the function is idempotent by definiiton, it does not mean that its implementation is. Unicode is changing too, new symbols are added.
I wonder if Unicode UTR#30 would be an alternative (more reliable?) method of idempotent normalization? http://www.unicode.org/reports/tr30/tr30-4.html
Or maybe the twisted algorithm really is UTR#30, just not labelled that?
As far as I can tell, UTR#30 did not make it to formally being part of the unicode spec, for reasons I'm not entirely clear on -- it is nonetheless quite useful, and this case is an example. Solr for instance still uses it. (http://wiki.apache.org/solr/AnalyzersTokenizersTokenFilters#...).
It might be a pain to find code implementing UTR#30 in your language of choice though (I am not sure if it's part of current ICU libraries or not).
It's also worth pointing out, that in addition to this kind of 'folding' of different-but-look-the-same graphemes, in this sort of use case you ABSOLUTELY need to do byte normalization as per UAX#15 http://unicode.org/reports/tr15/ . Probably NFKC for this sort of use case.
Why?
Suppose you only pass the original name around instead. Then you don't require your canonicalization function to be idempotent, which might be good in your case since it wasn't.
http://en.wikipedia.org/wiki/Ohm#Ohm_symbol
"Unicode encodes the symbol as U+2126 Ω ohm sign, distinct from Greek omega among letterlike symbols, but it is only included for backwards compatibility and the Greek uppercase omega character U+03A9 Ω greek capital letter omega (HTML: Ω Ω) is preferred."
And from the Unicode Standards doc that is the source for that section:
"Greek Letters as Symbols: The use of Greek letters for mathematical variables and operators is well established. Characters from the Greek block may be used for these symbols.
For compatibility purposes, a few Greek letters are separately encoded as symbols in other character blocks. Examples include U+00B5 µ n the Latin-1 Supplement character block and U+2126 Ω in the Letterlike Symbols character block. The ohm sign is canonically equivalent to the capital omega, and normalization would remove any distinction. Its use is therefore discouraged in favor of capital omega. The same equivalence does not exist between micro sign and mu, and use of either character as micro sign is common; for Greek text, only the mu should be used."
Even English speakers will easily confuse 1,I, and l, depending on how they're represented by the browser. And 0/O.
For more fun, try drawing any shape on http://shapecatcher.com/ and see all the similar-looking Unicode characters.
Note that omega was probably used so that the 'O' wouldn't be confused with '0', e.g. 4O would be confusing, but 4Ω is not.
[1]The tesla is abbreviated 'T', joule is 'J', etc. etc.
I have seen also much worse solutions. Where actually giving username, logs you in (sets logged in session cookie) and then prompts for password. When you enter invalid password you're logged out. If you give username, and then change url, you're in. Business as usual. When you test it, it works. Username + right password = ok, Username + wrong password != ok. Tests passed, and that's it.
You may not create as much havoc as in the original post, but some level of confusion at least.
canName=canonical_username(name); for(canName1=canName,i=0;canName1!=canName;i++){ canName1=canonical_username(canName); if(i>=treshold){ stop_registration(); break; } }
Erm...
Similarly, do you really want to do tech support when someone forgets that their username is "nodata" instead of "Nodata", since their phone auto-capitalized things when they signed up? It happens. And how do you know they're "Nodata" and not "nOdata" or "nOdAtA"?
It's true that many sites don't care about this, but I don't fault Spotify for trying to prevent it.
The letter À (A grave) can be written as the UTF-8 bytestream 0xC3 0x80 (i.e. a single "character"), or as À - i.e. a letter A, then a combining grave character i.e. 0x41 0xCC 0x80.
The two are identical. Except they have different byte representations. If you don't normalize your unicode you will run into major problems.
The one you are talking about is the one unicode actually calls normalization, and is dealt with in UTR#15. http://unicode.org/reports/tr15/
You are absolutely right that, in almost any situation taking unicode input where you're ever going to need to compare strings (and in most where you're ever going to need to display them), you are going to need to apply one of the UTR#15 normalization forms. UTR#15 normalizes different byte representations of what, in ALL circumstances are indeed identical characters/graphemes. A lot of people don't take account of this.
Then there's the kind of canonicalization that OP talks about, which Unicode actually calls 'folding', and is about characters/graphemes which really ARE different characters but which, for _some_ but not all contexts may be treated as 'equivalent' (if not neccesarily identical). The simplest example is case insensitivity, but there are other trickier ones in the complete repertoire, like those discussed in the OP.
This second kind of 'folding' canonicalization is a lot trickier, because it is contextual, not absolute. Which is maybe why Unicode started out trying to make an algorithm for it in UTR#30 but then abandoned it. Nonetheless, despite it's trickiness and contextuality, you often still really do need to do it, as in OP.
I suppose I feel that I was probably right to call myself arrogant then, to find out what you noted about Spotify's creators/creation. Interesting, thanks for the perspective check
Current work is on the PRECIS framework [2] which uses the metadata for Unicode code points to determine how to handle them during canonicalization instead of relying on a hard coded set of mapping tables. There's still a lot of work to be done, mainly to review that the process works reliably and doesn't introduce subtle new issues. Peter Saint-Andre (one of the authors of PRECIS) has just started on a Python tool for testing how a given version of Unicode is handled by PRECIS (https://github.com/stpeter/PrecisMaker).
[1] https://www.ietf.org/rfc/rfc3454.txt
[2] https://tools.ietf.org/html/draft-ietf-precis-framework-08
> As it turns out, even though Gmail will act as if there is no period in a username when delivering mail, it will not permit users to register accounts whose only difference to other account is a period. That is, if bob.jones is registered, bobjones cannot get an account (he could get bobjones2, obviously).
So, it's not a bug but rather a pretty nice feature: ever seen a handwritten email address with a period squeezed in there?
I've always liked that they do this on my accounts.