The benefit with UTF-16 is that you can't accidentally pass a string to, say, strlen(). But, yes, people will forget that not all Unicode code points fit in 16 bits, and won't test with the right kind of input to find out that they've done it wrong. So errors can still creep in; but I will continue to argue that those errors are less common because the fact that you're dealing with Unicode (and not ASCII or something else) is more obvious and you're more likely to need library calls (that will do the right thing) to do anything useful.
I'm not sure why you'd prefer UTF-32, it doesn't make correct text manipulations any easier, but — much like UTF-16 — it does make incorrect assumptions and text manipulations much easier.
UTF-8 for strings on disk, yes. But UTF-8 for in-memory strings has a long track record of being much harder to actually do correctly.
UTF32 will not help you much with that, save that you've got a separate type for "strings" and "bunch of bytes"
Which you can have anyway, so do that, it's a good idea which doesn't require using UTF32.
> But UTF-8 for in-memory strings has a long track record of being much harder to actually do correctly.
As if other in-memory encodings had a better track record.
Yes, people forget to normalize their UTF-32, UTF-16 and UTF-8 strings before comparing for equality. Yes, people forget that whether a code point is a letter depends on who's asking the question (I usually don't consider Greek letter pi a letter, but Greeks do). Yes, it's true that reversing a Unicode string in any encoding requires more than simply reversing the individual elements (because of combining characters).
So, yes, it's possible to get things wrong in any encoding. But UTF-8 has more ways to screw up than the alternatives. And UTF-16 has a few more than UTF-32; but I can accept UTF-16 if there are external reasons to (e.g., working on Windows).
But you can still copy only parts of a grapheme cluster. If you want to do Unicode right, you have to treat even UTF-32 as a variable-length coding.
I realize I won't convince you. That's what I meant in the original comment that I know many people disagree with me on this.
Glib seems to.
https://developer.gnome.org/glib/2.37/glib-Strings.html
Though in complete fairness the type appears to be just a "bags of bytes" type and so could hold anything, it's really meant to hold UTF-8 (as you can tell from the functions that append/prepend Unicode chars).
Which is relevant… how?
> Can you point me to any projects that manipulate UTF-8 encoded in-memory strings and actually use a different type for it them?
Rust does. Python 3.3 does something similar but slightly different (it switches internal representation between iso-8859-1, UCS2 and UCS4 depending on the string's codepoints).
> I realize I won't convince you. That's what I meant in the original comment that I know many people disagree with me on this.
Of course you won't convince me, your original comment is based on inane premises.
Note to put aside, since this was all about C++, the toupper() in <locale> is a template type that takes an arbitrary character type and a locale. There are also functions for locale sensitive collation and comparison. So, modulo bad implementation, you should be able to do basic unicode string handling in ISO C++.
The GNU stdlibc++ manual goes in to quite a lot of detail:
https://gcc.gnu.org/onlinedocs/libstdc++/manual/localization...
The locale stuff has been intentionally left vague. So:
(1) assuming that your implementation supports passing a char16_t or char32_t to the functions in <locale>, you're guaranteed it will do the right thing as far as Unicode and the C++ standard (although, there are some well-known problems with what the C++ standard defines the "right thing" top be for lowercasing epsilon -- since in Greek, there are two lowercase forms of epsilon, and the C++ standard doesn't provide the function enough information to decide between those two forms -- and for uppercasing ẞ (LATIN CAPITAL LETTER SHARP S) -- because in German the uppercase form takes two letters, and the C++ standard assumes doesn't allow for that).
The fun part is that plain char's can be encoded in several different ways, and not all have anything to do with Unicode. So you need to have a Unicode-aware locale for the functions in <locale> to do anything sensible with Unicode. Of course, that's something of a tautology, but I'm not sure what the standard requires for locales. So, yes, if your implementation is Unicode aware, you can call a standard C++ function to get what you want, but be sure to call the right one. There's another, with the same name, that isn't guaranteed to do what you want.
A language where the "String" data type is as follows:
A rope of "logical characters" (One or more code points, such that they are logically one character. So an accent is combined with the previous character, that sort of thing.)
With the additional "restriction" (read: implementation detail) that within a single node all logical characters must have the same width. (You can, for example, store a single one-byte character in a run of two-byte characters as an overlong-encoded two-byte character, but this is just an optimization.)
Short ropes degenerate to a flat array.
(You have to do a workaround for single code points that encode multiple logical characters. You split them into N parts encoded in the private unicode range or something similar, and when displaying them recombine them if they are in the correct order, otherwise normalize them. Although I'm up in the air about this. Should reverse("st") be "st"? Or "ts"? (That's the single unicode character "st", for those that are confused.))
Ideally, you put character encoding directly within nodes.
That way most things "just work". Running a string through the encoder twice doesn't do anything, as it detects the encoding is the same as the target encoding and doesn't do anything. Reversing a string "just works". Indexing a string is sub-linear time, but gives decent results. (Indexing a string and getting invalid unicode as a result is never fun!) Concatenating strings takes sublinear time even. This works really well with immutable data structures, or quasi-immutable data structures. (there's some tricks with rewriting ropes to take maximal advantage of structure-sharing that preserve the illusion of an immutable data structure without actually being immutable.)
And if you really want you can start doing fancy things like allowing lazy generators within strings, or lazily decompressing / reading data from disk.
To store on disk? Yeah, go with UTF-8. (Or my personal favorite pet encoding: compressed UTF-32.)