AÉBéC would give
0x41FF42FF42
where FF would mean, refer to another table with the index of the UTF8 char like the following:
---------------
|------|------|
|0 |1 |
|------|------|
|0xC389|0xC3A9|
|-------------|
This would make random access fast however will also increase overhead at other places.
Also IIRC, UTF8 doesn't use values 0x80 <= unused_utf8_values <= 0xFF. Value between 0x80 and 0xFF could be used to refer to an index of the table. eg 0x80 = index 0, 0x81 = index 1 ... 0xFE = index 126 and 0xFF = refer to other table.
Regardless of values over 0x80. The index in the special table will still have to be kept when iterating over it for find and stuff.
EDIT: table formatting