* Number of native (preferably UTF-8) code units
* Number of extended grapheme clusters
The first is useful for basic string operations. The second is good for telling you what a user would consider a "character".
* Number of native (preferably UTF-8) code units
* Number of extended grapheme clusters
The first is useful for basic string operations. The second is good for telling you what a user would consider a "character".
Reversing a string is what I would consider basic string operations, but I also expect it not to break emoji and other grapheme clusters.
Nothing is easy.
Tbh I can't think of a time when I've actually needed to do that.
find ... | rev | sort | rev > sorted-filelist
I had several directories and I wanted to pull out the unique list of certain files (TrueType fonts) across all of them, regardless of which subdirs they were in. (I'm omitting the find CLI args for clarity; the command just finds all the *.ttf (case-insensitive) files in the dirs.)By reversing the lines before sorting them (and un-reversing them) they come out sorted and grouped by filename.
Recently I rewrote the codepoint reverse function to make it faster. The trick is to not write codepoints, but copy the bytes.
One the first attempt I introduced two bugs.
I wonder: several comments are saying it hardly makes any sense to reverse a string but... Certainly there are useful algorithms out there which do work by, at some point, reversing strings no!? I mean: not just for the sake of reversing it but for lookup/parsing or I don't know what.
I have never once wanted to actually do this in real code, with real strings.
Furthermore, the few times I've tried to do something like this by being cute and encoding data in a string, I never get outside the ASCII character set.
* Passing strings around
* Reading and writing strings from files / sockets
* Concatenation
Anything else should reckon with extended grapheme clusters, whether it does so or not. Even proper upcasing is impossible without knowing, for one example, whether or not the string is in Turkish.
I'd quibble it's not the basic unit of measure so much as how changesets are represented. The user edits based on grapheme clusters. The final edit is then encoded using codepoints, which makes sense because a changeset amounts to a collection of basic string operations (splitting, concatenating, etc). As you note, it would be undesirable for changesets to be aware of higher level string representation details.
I can see why it would happen to be a codepoint, this might be ergonomic for the language, but it seems to me that, like clustering codepoints together in graphemes, clustering bytes into codepoints is something the runtime takes care of, such that a changeset will be a valid example of all three.
Using byte offsets also makes it possible to express a change which corrupts the encoding - like inserting in the middle of a multi byte codepoint. That goes against the principle of “make invalid data unrepresentable”. Your code is simpler if you don’t have to guard against this sort of thing. And you don’t have to worry about that if these invalid changes are impossible to represent in the patch format.
> The first is useful for basic string operations
Can you expand on this? I don't see why knowing the number of code units would be useful except when calculating the total size of the string to allocate memory. Basic string operations, such as converting to uppercase, would operate on codepoints, regardless of how many code units are used to encode that codepoint.
Converting 'Á' to 'á', for example, is an operation on one codepoint but multiple code units.
In some languages this is 90% of everything you do with strings. In other languages it's still 90% of everything done to strings, but done automatically.
I’ve used this for collaborative editing. If you want to send a change saying “insert A at position 10”, the question is: what units should you use for “position 10”?
- If you use byte offsets then you have to enforce an encoding on all machines, even when that doesn’t make sense. And you’re allowing the encoding to become corrupted by edits in invalid locations. (Which goes against the principle of making invalid state impossible to represent).
- If you use grapheme clusters, the positions aren’t portable between systems or library versions. What today is position 10 in a string might tomorrow be position 9 due to new additions to the Unicode spec.
The cleanest answer I’ve found is to count using Unicode codepoints. This approach is encoding-agnostic, portable, simple, well defined and stable across time and between platforms.
dz - \u0064\u007a, 2 basic latin block codepoints
DZ - \u0044\u005a
Dz - \u0044\u007a
dz - \u01f3, lowercase, single codepoint
DZ - \u01f1, uppercase
Dz - \u01f2, TITLECASE!
What happens if you try to express dż or dź from polish orthography?You can use
dż - \u0064\u017c - d followed by 'LATIN SMALL LETTER Z WITH DOT ABOVE'
dż - \u0064\u007a\u0307 - d followed by z, followed by combining diacritical dot above
dż - \u01f3\u0307 - dz with combining diacritical dot above
multiplied by uppercase and titlecase forms
In polish orthography dz digraph is considered 2 letters, despite being only one sound (głoska). I'm not so sure about macedonian orthography, they might count it as one thing.Medieval ß is a letter/ligature that was created from ſʒ - that is a long s and a tailed z. In other words it is a form of 'sz' digraph. Contemporarily it is used only in german orthography.
How long is ß?
By some rules uppercasing ß yields SS or SZ. Should uppercasing or titlecasing operations change length of a string?
The only thing it's useful for is sizing up storage. It does nothing for "basic string operations" unless "basic string operations" are solely 7-bit ascii manipulations.
Yes you locate the specific indexes by using extended grapheme clusters, but you use the retrieved byte indexes to actually perform the basic operations. These indexes can also be cached so you don't have to recalculate their byte position every time (so long as the string isn't modified).
That has little to do with the string length, those indices can be completely opaque for all you care.
> Yes you locate the specific indexes by using extended grapheme clusters
Do you? I'd expect that the indices are mostly located using some sort of pattern matching.
~Also, DOM Strings are not UTF-16, they're UCS-16.~
EDIT: UCS-2, not UCS-16. Also, I'm confusing the DOM with EcmaScript, and even that hasn't been true in a while.
Hm, according to the spec they should be interpreted as UTF-16 but this isn't enforced by the language so it can contain unpaired surrogates:
From https://heycam.github.io/webidl/#idl-DOMString
> Such sequences are commonly interpreted as UTF-16 encoded strings [RFC2781] although this is not required... Nothing in this specification requires a DOMString value to be a valid UTF-16 string.
From https://262.ecma-international.org/11.0/#sec-ecmascript-lang...
> The String type is the set of all ordered sequences of zero or more 16-bit unsigned integer values (“elements”) up to a maximum length of 253 - 1 elements. The String type is generally used to represent textual data in a running ECMAScript program, in which case each element in the String is treated as a UTF-16 code unit value... Operations that do interpret String values treat each element as a single UTF-16 code unit. However, ECMAScript does not restrict the value of or relationships between these code units, so operations that further interpret String contents as sequences of Unicode code points encoded in UTF-16 must account for ill-formed subsequences.
RE the UTF-16 vs UCS-2 stuff, that’s probably a distinction which has technical meaning but will collapse at some point because no one actually cares, much like the distinction between URI and URL.
Meaning the software will deal much less well when it's wrong.
> and provides a nice intuitive estimation of how long text actually is
Not really due to combining codepoints, which make it not useful.
> as well as providing a fair encoding for most commonly used languages.
Which we know effectively doesn't matter: it's essentially only a gain for pure CJK text being stored at rest, because otherwise the waste on ASCII will more than compensate for the gain.
UTF-8 is the encoding you should generally always reach for when designing new systems, or when implementing a network protocol. Rust, Go and other newer languages all use UTF-8 internally because it’s better. Well, and in Go’s case because it’s author, Rob Pike also had a hand in inventing UTF-8.
Ironically C and UNIX, which (mostly) stubbornly stuck with single byte character encodings generally works better with UTF-8 than a lot of newer languages.
To summarize: the "codepoint" is a broken metric for what a grapheme "is" in basically any context. Edge cases would be truncating with a guarantee of a valid encoding on the substring? But really you want to truncate at extended grapheme cluster boundaries. Truncating a country flag between two of the regional indicator symbols might not throw errors in your code, but no user is going to consider that a valid string. The same is true of all manner of composed characters, and there are a lot of them.
So the only advantage of using the very-often-longer UTF-16 encoding is that it's an attractive nuisance! This makes it easier to write code which will do the wrong thing, constantly, but at a low enough rate that developers will put off fixing it.
Unicode is variable width, and what width you need is application-specific. That's the whole point of the article! UTF-8 doesn't try to hide any of this from you.