From PyPy's Twitter feed:
"unicode-utf8 just got merged! Unicode strings are now internally represented as utf-8 in PyPy, with an optional extra index data structure to make indexing O(1). We'll write a blog post about it eventually."
Python strings are random access indexable. UTF-8 has variable length characters, so it is not directly randomly accessable. So the Python 3 representation inflates them to 1, 2 or 4 bytes per character, depending on the longest character in the string.
In practice, most strings are not random accessed. Most subscripts are +-1 from a previous subscript. So most actions can be optimized into uses of "advance one UTF-8 char" or "back up one UTF-8 char". Hard cases require generating an index array of the entire string. That takes a full string scan. Apparently that's what PyPy is doing.
This raises the cost of functions like "find()", which search the string for something and return an integer which is a position in the string suitable for random access use. I'd once suggested that "find", etc. return an opaque object which was really a byte index into the string. If you used that in a subscript, the right thing would happen. If you converted it to an integer, the index array would have to be computed. Is PyPy doing that, or does the use of "find" force index generation?