UTF-8 bit by bit (2001)
wiki.tcl-lang.org
wiki.tcl-lang.org
Since that limit is artificial it can safely be ignored. UTF-8 was designed as a 32b encoding. I don't think it is arbitrary-precision either though, since it was also designed to be self-synchronising.
> glypheme
Grapheme. A grapheme is more or less a user-perceived character, a "grapheme cluster" is a group (cluster) of codepoints roughly corresponding to a grapheme.
> A sequence of code points can be different kinds of normal forms, for example.
That has limited relevance, since not all clusters exist in precomposed forms they have to be handled regardless of normalisation, there's just some redundancies.
I can see an extension to 7 bytes (by making the leading byte 11111110 which I think would still be unambiguous) going up to 37 bits, but at 8 bytes you'd start getting collisions in the byte patterns and would lose the self-synchronising properties.
0xxxxxxx
110xxxxx 10xxxxxx
1110xxxx 10xxxxxx 10xxxxxx
11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
111110xx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx
1111110x 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx
11111110 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx
11111111 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxxPossibly because I assumed you'd need a special continuation byte and all of them would be unavailable, but the existing continuation bytes work fine so that's just not correct. FF could even lead to special patterns of continuation bytes allowing for smuggling more data in there I guess.
The original UTF-8 was at most 31 bits long (one bit in the leading byte, 6*5 = 30 bits in six subsequent bytes).
> you'd start getting collisions in the byte patterns and would lose the self-synchronising properties
I'm not sure why. The main requirement for self-synchronization in this case is that the leading byte doesn't appear anywhere else, which would be the case for the leading byte FF. It would require substantial modifications to go beyond 7 bytes (say, it would have alternative means to denote the length of following data bytes) but it is surely possible.
Although the term "character" is ambiguous so Unicode use more specific terms. An encoded integer represent a "code point" which usually corresponds to a character, but in some cases characters are represented by multiple code points. For example "â" might be represented by the code point for "^" following the code point for "a". This in turn opens a whole can of worms since the same character can be represented as code points in multiple ways.
“Since RFC 3629 (November 2003), the high and low surrogate halves used by UTF-16 (U+D800 through U+DFFF) and code points not encodable by UTF-16 (those after U+10FFFF) are not legal Unicode values, and their UTF-8 encoding must be treated as an invalid byte sequence.”
The article is unfortunate because it's actually handling a common non-UTF-8 encoding in which U+0000 is encoded differently so as to keep C-style null-terminated string semantics while allowing the NUL character. You should probably avoid the necessity for such semantics.
(110)00000 (10)000000 => C0 80
EDIT: OK I asssume he's referring to C and octal here and not a hexadecimal and decimal numbers?
I think he might have been clearer if he'd said
"To represent a NUL byte without any physical NUL bytes, we don't discard the indicators and instead treat it like a character above ASCII, which must be a minimum two bytes long 11000000 10000000 => hexadecimal C0 80."