Everything else has.
And then is UTF-16 which has all the pains of UTF-8 with none of the advantages of UTF-32
Everything else has.
And then is UTF-16 which has all the pains of UTF-8 with none of the advantages of UTF-32
Officially, it's at most four bytes, of which 21 bits are usable for encoding codepoints - so that's an upper limit of 2^21 codepoints.
There is an initial byte encoding the length as a series of ones, so if you went ahead and extended the standard to simply allow more bytes, you could get up to 8 bytes, of which 48 bits would be usable.
I can see that a six-byte version with 31 data bits was previously standardised before they settled on four.
I guess you could extend it further by allowing more than one initial byte encoding the length, then it would be arbitrary length. But at that point I'm not sure if it loses its self-synchronising ability, and in any case it would be a different standard at that point.
I think you'd only be able to go up to 7, since 10xxxxxx is still reserved for trailing octets. And even with 7, the entire first octet is consumed by the length indicator alone.
So you get 0xxxxxxx, 110xxxxx, 1110xxxx, 11110xxx, 111110xx, 1111110x, and 11111110 as the 7 different length-indicating head octets. In the last case, you'd have 36 usable bits for encoding a codepoint.
Also note that if you did add 11111111 as a valid head octet representing an 8 octet long encoding, you'd still only have 42 usable bits (since the first byte is still entirely consumed by the length indicator)