Someone who decides to identify themselves with Klingon (whether their parents gave them that name or not) should expect to have an alias ready...
Someone who decides to identify themselves with Klingon (whether their parents gave them that name or not) should expect to have an alias ready...
Even if it wasn't for historical characters that aren't part of Unicode, this will probably stay that way because of the inefficiencies of encoding Asian text in e.g. UTF8.
That's part of the reason the ruby programming language didn't have proper Unicode support for a long time (and now supports arbitrary encodings for its strings, not just Unicode ones)
That's true but doesn't really answer GP's question. Shift-jis is an encoding for one of the JIS X character sets. Unicode includes all the characters defined in JIS.
http://unicode.org/faq/han_cjk.html#8
Though if we enter into details it gets a bit messy due to the complications due to different simplifications occured in (mainland) China and Japan, questions about different glyphs for the "same" character, etc.
I think it's more about business process than technology limitations.
So my point is, maybe then it's better not to implicitly promise that your system has no limitations at all, but make these limitations more obvious instead. This'll help to avoid surprise errors and ensure consistent user experience by setting correct expectations.
I admit that most users today expect systems to accept their names in their language though, and that it might be an edge case depending on particular occasion, but then I can see why most payment systems apparently still accept only latin letters (cardholder's name is an example).
Actually I think the image field is a good idea, just have everyone use a typed alias and then draw/render what they want to be called. Maybe an image and/or a sound, to be more inclusive.
There's a lot of code out there that works fine with one-char code points but breaks on two-char code points.
UTF-8 is inherently variable-width, but it's a flat scheme in that one universal mechanism suffices to represent everything you can represent.
UTF-16 appears fixed-width until you realize it isn't, at which point you discover it's segmented and the segmentation scheme is both somewhat complex and likely to be missed entirely in naïve testing.
What does match is the fact that you can use a huge number of combining characters [1] to form a single glyph; each combining character and the base are a code point, so in order to figure out how many glyphs there are you have to iterate with the knowledge of what code points are combining characters.