- 0xxxxxxx -> 7 bits, ASCII compatible (same as UTF-8)
- 10xxxxxx -> 6 bits, more bits to come
- 11xxxxxx -> final 6 bits.
It has multiple benefits: - It encodes more bits per octet: 7, 12, 18, 24 vs 7, 11, 16, 21 for UTF-8
- It is easily extensible for more bits.
- Such extra bits extension is backward compatible for reasonable implementations.
The last point is key: UTF-8 would need to invent a new prefix to go beyond 21 bits. Old software would not know the new prefix and what to do with it. With the simpler scheme, they could potentially work out of the box up to at least 30 bits (that's a billion code points, much more than the mere million of 21 bits).The