Obviously the world is a big place and there is room for lots of paradigms and worldviews and we aren't supposed to judge too much.
But come on. If new code isn't working naturally in UTF-8 in 2021 then it's wrong, period.
Obviously the world is a big place and there is room for lots of paradigms and worldviews and we aren't supposed to judge too much.
But come on. If new code isn't working naturally in UTF-8 in 2021 then it's wrong, period.
Paradoxically, trying to do "the right thing" and being an "early adopter" of (the now called) UCS-2 was a "mistake", as both Java and Windows can attest, by getting "stuck" supporting the worst possible Unicode encoding ad-infinitum. UTF-8 is the "obviously correct" choice (from the hindsight afforded by us talking about this in 2021).
I still find it funny that emojis of all things are what actually got the anglosphere to actually write software that isn't completely broken for the other 5.5 billion people out there.
I thought that Chinese and Japanese are the only languages that UCS-2 has trouble fully representing. I believe all the other living languages can actually be represented by UCS-2.
So using UCS-2 would actually work for almost everyone except maybe 1.5 billion people.
I was actually quite bemused to discover that some code review software I was using allowed me to "cursor" halfway through a smiley face emoji and enter a space (typing too fast to pay attention)... causing the infamous "box characters" because I'd accidentally split the smiley down the middle.
I get the need for extreme backward compat in browsers, but... this seems like one of those things that just might be worth fixing. Maybe a "use utf8" directive? :)
Because unicode support isn't binary. Being able to pass along and not mange blobs of unicode is already a lot better than ASCII-only.
It can be. But critically for this conversation, fixing your code to support emoji and other non-BMP characters doesn't necessarily fix those problems.
One of the great things about using Rust is that I don't have to have this argument. There doesn't have to be a debate about whether we should invest in fixing subtly broken code. Rust string-handling which generally works for Europeans will also work for Chinese!
This isn't the fault of the languages which were designed for UCS-2 (then known just as "Unicode"). But the fact that Rust emerged after UTF-8's ascendance means that Rust's users mostly get to avoid the UCS-2/UTF-16 legacy tarpit.
I doubt that is any more true than Java. You can easily write code in Rust that assumes you can split a String anywhere and get two valid strings, that you can compute the length of a String and get information about how long the printed representation will be, that you can find a substring in that string by simply iterating through UTF-8 code points etc. All of these assumptions are about as wrong in UTF-8 Rust as they are in UTF-16 Java.
And I say that as someone who is developing a language which only has UTF-32 support.
The problem is that even with UTF-32, doing things like splitting strings is inherently unsafe, so you are still going to need a Unicode library to do proper splitting by grapheme cluster. In practice, almost all string splitting works on ASCII text, and assumes everything else is data that should not be manipulated. For this, UTF-8 is perfectly acceptable.
" >From ken Tue Sep 8 03:22:07 EDT 1992 "
As discussed last week in the quoted history https://news.ycombinator.com/item?id=26735958 of UTF-8's mail messages: https://doc.cat-v.org/bell_labs/utf-8_history
It appears that Unicode arose between 1990 and 1991 for the initial versions https://en.wikipedia.org/wiki/Universal_Coded_Character_Set#... with the first published version of the specification in 1993:
"ISO/IEC 10646-1:1993 = Unicode 1.1"
As discussed another time on Hacker News https://news.ycombinator.com/item?id=20600195 https://unascribed.com/b/2019-08-02-the-tragedy-of-ucs2.html
It was around 1996 when it became clearly obvious (software got shipped to end users who cared and they complained back) that UCS-2 (16 bit characters) would be insufficient.
+++
Pragmatically this is forever, as long as backwards compatibility must be maintained, the exiting APIs which are built around the crazy 16 bit standard need to exist; but there's little reason they have to be native, rather than wrappers for UTF-8 compatible APIs.
It would even be a good time to standardize on a single user space programming API and have implementations on every operating system. Preferably including basic drawing and font layout functions. So that finally, most programs could be written once, compiled on a platform of choice, and work.
UCS-2 is in fact broken, but UTF-16 is a valid encoding of Unicode, which can be implemented correctly. So is UTF-32, although I can't imagine why anyone would want to use that one.
I can imagine why someone would want to use UTF-16, though: interoperability with Windows, where it's the native encoding. It isn't "wrong, period" to do Unicode in a way which is more convenient for the platform.
There is, of course, a ton of work to really implement Unicode correctly, and UTF-16 and UTF-32 can make it tempting to do the wrong thing, instead of biting the bullet and implementing all of the many ways in which codepoints coalesce into grapheme clusters, and making sure all functions for working with strings can recognize the distinction.
But it certainly can be done in any of the full encodings.
Even on Windows, it's much easier to just use UTF-8 internally and perform conversion to/from UTF-16 when you're about to do/return from a WinAPI call.
So do you think UTF-8 is always the best internal string representation? Or just for English speakers?
For Mandarine what would be optimal?
When you add in interoperability concerns, since so much text these days is UTF-8, for Mandarin at least UTF-8 is a perfectly defensible choice.
(A harder problem is Japanese — Japan really got screwed over with Han unification, so choosing Shift-JIS over any Unicode encoding is often best.)
FWIW I covered the space requirements of various encodings and various languages in this talk for Papers We Love Seattle:
(Source: I wrote a search engine library.)
This statement needs more support. I think “screwed over” is a bit harsh, since I’m not aware the impact on Japanese was anymore than the rest of CJK. Despite the Han unification controversy, Unicode has been heavily adopted in Japan. The space requirements are basically the same as all CJK. Half-width kana is heavier since they are one byte in shift jis but they’re relatively uncommon.
On the web you can work around this using the lang attribute to tell the browser how text should be interpreted.
It's notable that, for example, traditional and simplified Chinese does not have this problem because they are encoded separately.
Another problem is missing characters. Some people have complained of not being able to write their own name. I'm not sure to what extent this has been solved through Unicode updates.
[1] Although on the other hand you could argue that I'm spending more conscious effort reading Japanese text than a native speaker would.
Including language metadata with text was felt to be especially important for Japanese and Korean users. I was told the difference was like having "595 kg" displayed as "5P5 kg". That is, it's possible to decipher the intended meaning but it looks wrong and it takes a moment to work out what was meant. Depending on the language some glyphs can be mirror images, have extra strokes, strokes missing or in different places or at different angles.
Does this mean that “UTF-8 is not a good representation for non-European alphabets?” It may be less efficient but the difference does not seem shocking to me, considering that for most applications, the storage required for text is not a major concern—and when it is, you can use compression.
How so? Delphi for example has wide character-based strings as default, what's wrong with that?
All of those practices are immediately wrong once UTF-32 came in existence and UTF-16 became a variable length encoding. But even if that hadn't happened, what you want to be operating on is not characters, but grapheme clusters, which are equivalent to a vector of chars. Otherwise you won't handle the distinction between ë and ë or emojis correctly.
edit:
For example, we do a lot of string manipulation in Delphi. We might split a string in multiple pieces and glue them together again somehow. But our separators are fixed, say a tab character, or a semicolon. So this stiching and joining is oblivious to whatever emojis and other funky stuff that might be inbetween.
How is this doing it wrong?
I mean yea sure you CAN screw it up by individually manipulating characters. But I don't see how an UTF-8 encoded string in itself prevents you from doing the same kind of mistakes.
Obviously C is better than A or B, because you want people to have a good experience with your software. But weirdly, system A (broken always) is usually better than system B (broken in weird hard to test ways). The reason is that code that’s broken can be easily debugged and fixed, and will not be shipped to customers until it works. Code that is broken in subtle ways will get shipped and cause user frustration, churn, support calls, and so on.
The problem with UCS-2 is it falls into system B. It works most of the time, for all the languages I can type. It breaks with some inputs I can’t type on my keyboard. So the bugs make it through to production.
UTF-8 is more like system A than system B. You get multibyte code sequences as soon as you leave ASCII, so it’s easier to break. (Though it really took emoji for people to be serious about making everything work.)
if you have been handling Unicode and using wide characters, you have not been handling Unicode properly
I agree that UTF-8 is a better encoding overall for the majority of cases. I don't think that means UTF-16, which for example Delphi UnicodeStrings are[1], is not proper.
edit: maybe this is a language confusion thing. For historically tragic reasons, we're stuck with "char" as the basic element of string types in lots of languages. In Delphi a "widechar" is technically a code unit[2], and may or may not represent a code point. This is how I interpreted the OP. Maybe he meant wide characters as code points, in which I would agree.
[1]: http://docwiki.embarcadero.com/RADStudio/Sydney/en/Unicode_i...
[2]: https://en.wikipedia.org/wiki/Character_encoding#Terminology
The ergonomics of the language guide you in that direction when, as you say, a "char" doesn't actually represent a character. Or even an atomic unicode codepoint. And when string.length gives you an essentially meaningless value.
Luckily, code like this will also break when encountering emoji. Thats great, because it means my local users will complain about these bugs and they're easy for me to reproduce. As a result these problems are slowly being fixed.
Having worked in tech support for a piece of very expensive (~$100k per install annual support/license fee in the late 90s) enterprise software that had a GA release shipped to customers with a syntax error in an install script, I would state that more like “code that’s non-subtly broken is less likely be shipped to customers before it works.”
At work we have a codebase that does a lot of string handling. Both in reading and writing all kinds of text files, as well as doing string operations on entered data. Several hundred kLOC of code across the project.
We had one guy who spent less than week wall-time to move the whole project, and the only issue we've had since is when other people send us crappy data... if I got a dollar for each XML file with encoding="utf-8" in the header and Windows-1252 encoded data we've received I'd have a fair fortune.
- It isn’t the number of bytes, unless your string only contains ASCII characters. Works in testing, fails in production.
- It isn’t the number of characters because 16 bits isn’t enough space to store the newer Unicode characters. And even if it could, many code sequences (eg emoji) turn multiple code points into a single glyph.
I know all this, and I still get tripped up on a regular basis because .length is right there and works with simple strings I type. I have muscle memory. But no, in javascript at least the correct approaches require thought and sometimes pulling in libraries from npm to just make simple string operations be correct.
Rust does the right thing here. Strings are UTF-8 internally. They check the encoding is valid when they’re created (so you always know if you have a string, it is valid). You have string.chars().count() and other standard ways to figure out byte length and codepoint length and all the other things you want to know, all right there, built into the standard library.
At least .length tells ypu how much memory the string will occupy - not likely to be very important in a memory managed language, but at least it has one potential use. I don't see a use for the number of code points in a string.
- We don't use arrays of bytes because different languages have different native string encodings. Converting a UTF8 byte offset to a language that uses UCS2 is slow and complicated, and it opens the door to data corruption.
- And we don't use grapheme clusters because what counts as a grapheme cluster keeps changing with each unicode version. (And libraries are big and complicated).
So inserting an emoji into a document is treated the same way we would handle inserting a small list at some offset into a larger list.
> At least .length tells you how much memory the string will occupy
How much memory the string occupies isn't something I've ever wanted to know. I do sometimes want to know how many bytes the string will take up when I store it or send it over the network - but in those cases you pretty much always want UTF-8. And string.length doesn't help you at all with that.
I would have expected that your input method would normally insert grapheme clusters, and that these clusters would be treated as indivisible if they were inserted with a single 'keystroke'. I would also expect that you anyway need to count clusters when presenting something like a 'character count' to the user, as I don't think they would be very happy if a program reported that 'année' was 6 characters long.
> How much memory the string occupies isn't something I've ever wanted to know. I do sometimes want to know how many bytes the string will take up when I store it or send it over the network - but in those cases you pretty much always want UTF-8. And string.length doesn't help you at all with that.
This is what I was referring to, should have called it number of bytes instead of memory. Also note that probably the most used application protocol on the internet, HTTP, doesn't support UTF-8 and defaults to ISO-8859-1 (extended ASCII). For example, HTTP headers are not UTF-8 strings and you can't treat them as such. And even when sending a UTF-8 body the Content-Length header needs to be set to the number of bytes, not the number of unique codepoints, so again you'd use .length.
Where the carot can go is dependent on the editor, not the underlying protocol.
> I don't think they would be very happy if a program reported that 'année' was 6 characters long.
CRDT / OT edits don't make a lot of sense to the user in their raw form no matter what format you use for offsets. "Insert x at position 5043" is equally meaningless if 5043 stores a byte offset, a grapheme cluster index or codepoint index.
> Would it be significantly different if you allowed them to place the caret between the parts of a UTF-16 surrogate pair?
Yes - if you managed to insert something in the middle of a UTF-16 surrogate pair, the string contents would become invalid - and that causes weird language dependant problems. Rust will panic(). In comparison, a broken country flag just renders weirdly - which isn't ideal, but its fine. Mind you, in both cases you'd need one of the editors to do something weird to insert those characters in the document. As you say, the input method will normally treat grapheme clusters as indivisible anyway. But I much prefer invalid states to be impossible to represent in the first place when I can. I don't want a lack of input validation to allow wonky edits to crash my rust server. Invalid grapheme clusters are a much smaller problem in comparison.
And this all skips over how difficult it is to efficiently convert a UCS-2 offset position in javascript into a position in a UTF-8 string in rust. Its much easier to just count in codepoints everywhere.
> And even when sending a UTF-8 body the Content-Length header needs to be set to the number of bytes, not the number of unique codepoints, so again you'd use .length.
No, that'll break as soon as you insert non ASCII characters into your document. string.length will not tell you the number of bytes your string takes up. It will tell you the number of UCS-2 elements in your string, which is the number of codepoints + 1 for each UTF-16 surrogate pair. Which again, I've never wanted to know. You can't use that to calculate the UTF-8 byte length, where each codepoint takes somewhere between 1 and 4 bytes depending on its unicode table value.
Your example 'année' has a string.length of 5, but a UTF-8 byte length of 6. (From new TextEncoder().encode('année').length). If you set Content-Length to 5, bad things happen. Ask me how I know :/
My point was exactly about whether languages should enforce string encoding. I maintain that strings should just be arbitrary byte arrays, and only text processing methods should enforce the appropriate encodings - that would mean that accidentally splitting a UTF-16 code point would be as painful as accidentally splitting a grapheme cluster: anything that wants to render the resulting string will have a problem, but anything in between won't.
> It will tell you the number of UCS-2 elements in your string, which is the number of codepoints + 1 for each UTF-16 surrogate pair.
Oops, here you are completely right. I was under the mistaken assumption that Java String.length() would return the number of 16—byte chars in the string, when in fact it returns the same useless number as the chars().count() method. Sorry about that!
And a point of clarity - Java’s String.length() does not return the same value as rust’s chars().count(). The former returns the useless UCS-2 count. The latter returns the number of Unicode codepoints. Java’s length() will count many single codepoint emoji as having length 2 (same as javascript, C#) while rust will correctly, usefully count one codepoint as one character. (As will swift and go, depending on which methods you call.)
UTF-8 is a nuisance as an in-memory representation because the characters are variable size. You can't get the length of a string without parsing it start to end, and you can't get a character by index without parsing and counting all the previous ones. 16-bit characters (wchar_t, Java char, whatever NSStrings are made of, etc) work fine in 99% of the cases.
UTF-8 is indisputably a good encoding for when you're sending something over the network or putting it into a file or a database.
In other words, they don't work :-).
UTF-16 is also variable-length. Sometimes a character fits in 16 bits, and sometimes it doesn't. From a practical view it's worse than UTF-8, because tests are less likely to detect bugs before shipping.
Even UTF-32 is, in reality, variable-length. Many code points are combining characters, so you need multiple code points to get a single grapheme.
If your language or API requires you to do something, then you'll need to do that. But unless there's an API requirement, in most situations UTF-8 is the best choice for network, storage, and processing. There are exceptions, but they're just that... exceptions.