But you can index UTF-8 strings with precomputed bookmark indices (byte offsets) just fine. Point is, to precompute them, you would surely have iterated the string before.
What I mean is that "extract the 42th codepoint/glyph/whatever from this UTF-8 string" is a pointless operation for free-form strings because in a free-form string the character position is meaningless.
(Basically all non-ASCII UTF-8 is pretty much a black box. You can't do serious computation with general Unicode because it's so complex and, as a consequence, ill-defined in practice).