This change takes advantage of the fact that most JS strings fit into an 8-bit charspace, so for those that do, it uses a more compact representation internally.
This optimization is simply: if we have a string and we know that all of the uint16_ts in the string are <= 255, then just store it as a sequence of uint8_ts.
1. Opaque cursors pointing somewhere in the conceptual sequence of code points, with constant-time dereferencing,
2. Ranges, defined by starting and ending cursors, and
3. The ability to move cursors forward or backward by either code points or composed grapheme clusters.
This would be a saner interface than any other I've seen, and it puts very few constraints on the underlying encoding.
3. Forward and backward are likewise language and tailoring dependent because they depend on graphemes. There may also be application-specific tailoring such as the handling of combining marks, in some scripts "forward" and "backward" are not clearly defined.
For example which of the four Unicode character normalization interests you most? Or you need grapheme clusters? Or you need code points? Or byte values?