Let’s Stop Ascribing Meaning to Code Points
manishearth.github.io
manishearth.github.io
By the way, I've sometimes heard that it was the popularity of emoji that finally made American and Western European coders learn about the particulars of unicode. Is this just a jokingly told anecdote, or did emoji really help?
I'm sure plenty of Western programmers were aware of this stuff, though. Just not enough.
https://en.wikipedia.org/wiki/Code_page#ISO.2FIEC_8859-relat...
FWIW (more or less proper) Unicode support seems to be pretty much an implicit, hard requirement today for almost anything facing users.
http://www.24heures.ch/suisse/Les-noms-de-l-Est-mutiles-lors...
I doubt that emoji helped further popularize Unicode in Europe, because in my experience (anecdotal of course) developers were already well aware of the benefits of Unicode (and UTF-8 in particular for the web) before emoji got adopted from Japanese telephones and embraced en-masse. I suspect that in Europe the need for Unicode arose sooner. For Dutch for example; with the introduction of ISO-8859-15 in 1999 (subtly different from ISO-8859-1) you now often had two character encoding standards in mixed use for a single language, and when you had to render documents with names of people from various countries, things just got confusing (in the US diacritics tend to get dropped instead). UTF-8 offered a clean break from this.
It seems plausible that emoji may have helped Americans break the remnants of the ASCII barrier though.
Or perhaps you're just talking about Chinese where I'm also considering Kanji and others. Which isn't from the most spoken language in the world, but still, top 10!
For Korean and Vietnamese the extra Chinese characters might apply to ancient names (geographical usually I would wager), but neither language has a lot of ideographs in modern usage.
It's certainly an interesting topic though.
> Now, Rust is a systems programming language and it just wouldn’t do to have expensive grapheme segmentation operations all over your string defaults. I’m very happy that the expensive O(n) operations are all only possible with explicit acknowledgement of the cost. So I do think that going the Swift route would be counterproductive for Rust.
I'm not sure I understand the logic here. How is Rust exempt but not, say, Python? Just because it's a systems language? I get why you wouldn't make it the default in Rust, but the logic applies to all languages, doesn't it?
Aside: I was talking about this with a friend just yesterday. Is there any language that separates the concept of a string and a user-facing string? Specifically for localization purposes, the way gettext works.
eg. I wish that in Python you could do this:
err = "internal_error"
msg = t"Internal error!"
error_msg would be easy to pick up for translation, rather than have to do some cludgy import as _ and _("Internal error!"). Hm, now that I'm writing this here, I guess it's still not a type you'd want to operate differently on than just plain unicode strings.You deal with strings as bytes (opaque data) (edit2: didn't really mean bytes here. see child comments) more often than strings as text in systems programming, basically.
Also, you'll want to write stuff like fast parsers in Rust. For a parser you're usually doing ascii matching anyway.
Swift hides the actual encoding of the string from the user. That's not something we can do in Rust.
> I get why you wouldn't make it the default in Rust
That's all I'm saying? I'm totally for having EGC operations there, I don't want them in the default operations. "it just wouldn’t do to have expensive grapheme segmentation operations all over your string defaults"
Basically, Rust tries to keep costs explicit, unlike languages like Python. This would be against that philosophy.
Edit: I'd also like to point out that indexing by code point or grapheme is a rare operation in Rust anyway.
Swift supports unicode in source code so its parser is unicode aware. I'm not sure there is a good case for parsing in a non-unicode-aware way.
If you're working with strings as opaque data why are you even using a String type?
> I'm not sure there is a good case for parsing in a non-unicode-aware way
I'm not saying that you should parse in a non-unicode-aware way, I'm saying that most parsing code deals with ASCII delimeters, so grapheme clusters are a distraction (and will mess up parsing, since most formats do not consider EGCs. An HTML closing tag followed by an accent would totally break the parser otherwise). You rarely need to index, though.
Basically, in parsing, you're looking for characters which themselves are single code points. So it doesn't matter.
> If you're working with strings as opaque data
Well, not entirely opaque :)
But the String type is basically "here is a valid piece of UTF8 data", and you can pass that around. That's a distinct enough purpose to have a String type, even if you don't mess with its innards.
IME the most used String APIs are those for interpolation, concatenation, and search, all of which don't need to be EGC aware (though search being EGC aware is sometimes a good thing). Slicing sometimes gets used in conjunction with the others.
I read the Swift Book when Swift first came out and it bothered me a lot that the book didn't say what the memory representation of strings is. In contrast, Rust documentation is very clear about this stuff.
Now you're really confusing me. Are you talking about simply byte strings vs. unicode strings? Most modern languages, including python, have that distinction. What am I missing?
I'm saying that most string handling in Rust is just lugging them around and copying them and concatenating/interpolating them.
On top of that, the few times you do iterate through them, you're looking for specific characters or character sequences (e.g. while parsing), and EGCs aren't so important there. They can be, but usually aren't.
Basically, personally the only times I've seen .chars() being used is in parsery things (where your grammar is defined over code points, not EGCs, so using EGCs is counterproductive and wrong), in the Servo DOM (where the spec tells you to operate on chars), and code where you're just verifying that the string contains or does not contain code points of a certain type (e.g. whitespace).
None of this is ascribing meaning to the concept of a code point, it's just using them as a crutch to reason about things like equality. (Of course, equality in Unicode is a whole other can of worms).
That's true in some sense [1] for unix-ish OSes, but doesn't hold on other operating systems.
[1] Even on unix-ish OSes it's pretty obvious that filenames are supposed to be encoded in something and aren't just uint8_t*. There's just no one telling you what that encoding is, so normally it's guessed (eg. $LANG). For filesystems that actually have a defined encoding (eg. NTFS\Win32, JFS, HFS+, APFS and the whole bunch of "non-unix OS FS") these are re-coded by the filesystem driver in the unix-ish OS. Eg. ntfs-3g transcodes filenames to UTF-8; cifs takes a "iocharset" option; and we of course have convmv(1).
The idea that "file names and generally everything in the OS is a bunch of bytes" is pretty much only shared by unix-ish systems.
Not quite, but Nim has a concept of a tainted string, which comes from outside your (web) application and has to be escaped or explicitly casted into a regular string (the Python web framework Django has something similar). You could use the same mechanism of "distinct types" to create a user facing string vs an internal string if you want. I used this in a little Unicode related experiment to create a number-of-bytes and number-of-chars type, and made my strings take these as index, so I would never confuse them. Pretty neat language feature I thought.
Nim does have something new though, its "distinct" types are more general than just tainted strings, and can be used whenever you have the same raw data type, but different meaning (and want the compiler to prevent accidential mixing). Of course, Haskellers will say that's old and boring...
This is maybe ok for very simple cases, but doesn't work as soon as you need context (in Python this is either done by extra _ parameters or through special comments) or pluralization.
Another algorithm for unicode strings which wants to index by code point is unicode regular expression matching. I heard that you do lose something here if you can't assume all code points encode to same length. Unlike other algorithms you mentioned which only iterate forward, depending on implementation regex needs to backtrack.
I once heard that complexity of implementing regex directly on UTF-8 is one reason why Python will not use UTF-8 internally.
burntsushi's utf8-ranges library used by the Rust Regex lib probably makes this easier to deal with. I wonder how it performs in relation to the Python implementation (which I assume is in native code)
fn rfind(s: &[&str], key: &str) -> Option<usize> {
(0 .. s.len()).rev().find(|&i| s[i] == key)
}
No need to explicitly worry about rune boundaries.
This, like forward find, returns a byte index into a UTF-8 string. Explicit byte indices into UTF-8 strings are troublesome.Yes, I know, I mention this in my blog post :)
Rust's indexing methods on strings actually do operate on byte indices. You can't do `"foo"[2]`, but you can do `"foo[2..3]"`, and it will panic if you are not on a boundary. Nobody uses this (like you say, it's troublesome) :P I've seen folks just break out the byte/char iterator whenever they need something more complicated, because usually they do need to iterate anyway.
Rust internally stores slices as byte indices (well, pointer + byte length), but the compiler enforces the validity of that.
The reason I bring it up is that UTF-16 is a very commonly used internal encoding for strings, and regular expression engines for platforms which use UTF-16 internally are generally designed to work on that encoding without conversion to UTF-32. Examples that spring to mind are most JavaScript regular expression engines, the ICU regular expression engine, and java.util.regex. If variable-width encodings were so difficult to deal with then you would expect these to have run into some problems with it; however, as far as I can tell, they have not. This is evidence against variable-width encodings being difficult for regular expression engines to deal with.
(Sanity check for my argument above: If you read through https://swtch.com/~rsc/regexp/regexp2.html, which covers both backtracking and non-backtracking regex engines and includes code, none of the code examples would need more than trivial modification to deal with variable-width encodings.)
IIRC Java also allows unpaired surrogates? But does consider non-BMP codepoints to be single characters. I think. Idk.
So from JSs point of view it is a fixed-width encoding, one where all byte values are valid. `"\uD83D\uDC68".match(/\ud83d/)` works (U+D83D is a high surrogate, and "\uD83D\uDC68" is U+1F468 MAN), so the regex engine operates on code units, not code points.
I think for Java the regex engine would have to deal with it as a variable-width encoding, though.
Edit: Hacker News removes emoji from comments.
(Note that java.lang.Character is a char, so wrong. I think there is no CodePoint type or similar that wraps int.)
It's unclear to me what you mean by "Java Characters". A Java char is a UTF-16 code unit, and the java.lang.Character boxes a char (16 bits unsigned). However, since Java 5, java.lang.Character has static methods for astral-aware UTF-16 operations where a code point is a Java int (32 bits signed).
> the dot matches the whole pair without splitting it
Yeah, so UTF16 in Java is being used as a variable-width encoding. Nice to know.
This requires that matches match UTF-8 substrings, only. That's easy for explicit strings, such as "abc". For more general forms, "." must match one rune, not one byte. But that's not hard.
Python's problem is that strings are random-access indexable by rune (in the Go sense). Python either needs a representation that's rune-indexable, or it needs to generate an index array for strings that need one. There's an argument for the second approach, because most loops don't need a random-access index. "for i in "abcdef: ... " does not, for example.
Yeah, one of my points in the blog post is that random-access rune (code point) indexing isn't very valuable. Python is stuck with it, but if anyone in the future needs to make a similar choice, I hope they don't use O(1) rune indexing as a reason unless they have a very specific use case where O(1) rune indexing actually matters.
(I, for one, am very happy with the "rune" terminology, "char" is way too overloaded. It took me a minute to understand why they did that in Go when I first picked it up, but when I realized the reason I found it brilliant)
I like "rune" as terminology. A rune is one Unicode code point. A grapheme is a sequence of runes which should not be split.
I'm a bit wary of this because I fear indexing in Python is currently used more than this scheme can handle (i.e. too much for it to be considered a win), if erroneously. I'm not too sure. I certainly have written bad python code using indexing in the past, but that's just me (and many years ago).
Typically this is found when parsing something, eg.
if line[0] == '#':
# skip comment lines
continue
Or prefixes = {...}
for line in lines:
prefix, remainder = line[:4], line[4:]
if prefix not in prefixes:
raise ValueError('Invalid prefix ' + prefix)
parsed = prefixes[prefix](remainder)By default RE2 matches directly against the UTF-8 encoding, see this neat state diagram:
https://swtch.com/~rsc/regexp/regexp3.html#step3
In terms of backtracing engines, perl also stores strings internally as UTF-8 (on most platforms) and its regexp engine runs directly on them. It can also match "eXtended grapheme clusters" with \X.
Why? This is certainly not a requirement for at least some levels of Unicode support.
> I once heard that complexity of implementing regex directly on UTF-8 is one reason why Python will not use UTF-8 internally.
It's really not that bad. Both RE2 and Rust's regex engine operate on UTF-8 directly, and this is one of the reasons why these engines are fast. If your automaton is byte based, then you can do a lot of nice tricks in your DFA implementation. (In fact, it's a critical reason why ripgrep doesn't slow down when handling Unicode features, unlike GNU grep.)
For Rust at least, the trickier parts of generating the UTF-8 automaton are available for reuse in the utf8-ranges crate: https://docs.rs/utf8-ranges/1.0.0/utf8_ranges/ (Which is used in at least two regex implementations I've written.)
I'd say the complexity claim might be coming from some other component of the implementation, or something more related to how Python strings are generally represented, rather than it being a fundamental complexity of building a regex engine on UTF-8 directly.
And, there is still no completely reliable way to guess the cell/column width of a symbol just by looking at it, especially if your goal is to be consistent with any other interpreter of the string (like a text editor). Still heuristics at best for any case outside common tables like CJK.