Proper encoding of your data is key indeed.
I mean, user data should be treated as binary anyway. And then hashing it to whatever makes sense for performance.
But still binary comparing for accuracy.
I mean, user data should be treated as binary anyway. And then hashing it to whatever makes sense for performance.
But still binary comparing for accuracy.
Also good moment to point out the you can't do a binary comparison of unicode strings, you need a unicode library that handles normalization.
Names exist at a different level of abstraction to binary and it’s a mistake to treat them as binary.
I don't know about that, there still needs to be some validation. Is your legal name really "Bobby $@134 Tables"? In this case validation comes in the form of "is every character EBCDIC-compatible?"