So the idea of ever having a fixed-length encoding for Unicode is basically impossible now. Best to just use UTF-8 for everything and logic to group it in to code points, grapheme clusters, or whatever other granularity is needed.
So the idea of ever having a fixed-length encoding for Unicode is basically impossible now. Best to just use UTF-8 for everything and logic to group it in to code points, grapheme clusters, or whatever other granularity is needed.
I've heard more than one person tell me they don't need to worry about text shaping since they are using UTF-8. (That statement doesn't make any sense) There is a lot of confusion with Unicode text rendering stack.
We do it all the time. Splitting a string at U+002C or U+003B or U+000A, looking for U+007B or U+007D, etc. In my experience, it's really common.
That means that what Unicode today considers "one character" can be infinite in size.
Typing in many Latin languages generates diacritics using ‘dead keys’, which are prefix operators. (Originally, on a physical typewriter, these were keys that simply struck the paper without triggering the mechanism to advance the carriage.) If Unicode had taken this hint, life would be easier.
(ISO/IEC 6937 had prefix diacritics, but it was too late.)
Process whole strings, the meaning of a partial string may not be what you hope, even without fancy writing systems and Unicode encoding.
"Give the money to Steph" - OK, will do
"...enson" - Crap, OK, somebody chase Steph and get our money back, meanwhile here's James Stephenson's money
"... once he gives you the key" - Aargh. Somebody chase down James and get the key off him.
Either way there's already a recommendation in this area. I'm not sure if it strictly fits this particular scenario but the "UAX15-D3. Stream-Safe Text Format" suggests that you should support sequences at least 31 code points long.
Well never mind that then; how do you like U+41,U+300 aka U+C0 latin capital letter a with grave accent. Not just in the BMP but in the part of the BMP that's only 2 bytes of UTF-8.
And there's a whole bunch of similar characters that don't have NFC encodings.