> Sorry, you can’t squeeze Unicode in that way, even unsynchronised: you fall just over 2½ bits short. No, reserving the 95 printable ASCII characters scuttles your hopes, meaning you get less than 1½ bits of that byte to use for the rest of Unicode, when you needed a touch over 4 bits.
So it depends on exactly what we mean by compatible, and that's why I removed it.
But in particular, the encoding I had in mind when I wrote that was one that allows most ASCII characters as the middle byte of a triplet. That way whether you truncate from beginning or end you'll never have a rogue byte successfully decode.
So something like: Lead bytes from C0-FF, middle bytes from 00-BF, final bytes from 00-1F or 7F-BF.
If you also exclude null, tab, newline, and line feed, that gives you 64 x 188 x 93, which is just enough for every code point.
But arguably that's not compatible enough because you couldn't search for single ASCII characters with a dumb byte-wise algorithm.
> Not quite. If you discarded ASCII compatibility, then you’d have almost four bits to spare, which is enough for a self-synchronising fixed-width 3-byte encoding (reserving the first bit of each byte to signal if it’s a start byte) or for an unsynchronised variable-width encoding of some kind; but to make it even 2-or-3 or 1-or-3 (rather than 1–3) bytes wide variable-width, you’d need four bits (3 + log₂ 2) for self-synchronisation, and you’re about 0.085 bits short. (1–3 bytes variable-width would require 3 + log₂ 3 ≈ 4.58 bits.)
I don't know where you got that equation but it's not the right one for the situation.
Here's a really simple encoding just to disprove it: Leading bytes encode the plane (17 options) plus two more bits into decimal values 0 through 67. Continuation bytes encode 7 bits into decimal values 128 through 255. 17 planes + 2 bits + 7 bits + 7 bits, perfect fit. The remaining values, 68-127, can be used for single byte (and two byte) encodings.
But also the definition I was using for self-synchronizing doesn't seem to match the one on wikipedia. My intent was a code where you can guarantee sync after a specific small number of bytes (unlike UTF-1). In particular, if you let standalone bytes and continuation bytes overlap each other, you can still guarantee sync within 3 bytes. This gives you lots of space to play with, and you could make a very useful encoding. You could have >150 single byte codepoints and thousands of double byte codepoints and fit everything else into three bytes.