The string type is broken (2013)
mortoray.com
mortoray.com
let s = String::from("noël");
println!("Printable? {}", s);
println!("Countable? {}", s.chars().count());
println!("Reversable? {}", s.chars().rev().collect::<String>());
println!("First three characters? {}", s.chars().take(3).collect::<String>());
Output: Printable? noël
Countable? 4
Reversable? lëon
First three characters? noë
Repo for the "noël" example and "cats" example: https://github.com/joelparkerhenderson/demo-rust-string-issu...Try:
let s = String::from("noe\u{0308}l");
println!("Reversable? {}", s.chars().rev().collect::<String>());
println!(
"First three characters? {}",
s.chars().take(3).collect::<String>()
);
you get l̈eon which (according to the article) is wrong; likewise "noe" as the first three chars, dropping the diacritic.There's no value in a String type if it doesn't behave for text. One can simply use a "Array<char>" to convey proper semantic meaning.
I updated my post with more info and a link to a code repo.
The code correctly shows me the reverse "lëon" with the umlaut, and the first three characters "noë" with the umlaut.
1. One code point: U+00EB. This is the "precomposed" form.
2. Two code points: U+0065 U+0308, aka e followed by ¨. This is the "decomposed" form, also known as a "combining sequence" since the diaeresis combines with the base character.
If your string type is a sequence of code points, then reversing a decomposed string will tear the combining sequence, and apply the diaeresis to the wrong character. Most string types are affected by this (or worse), Rust included.
The two forms get rendered identically, so you more or less need a hex editor to figure out which form you've got. I forked your repo and switched it to decomposed (the diff looks like a noop), and now it produces the wrong output:
https://github.com/ridiculousfish/demo-rust-string-issues
Another Rust example using explicit Unicode literals: https://play.rust-lang.org/?version=stable&mode=debug&editio...
One could reasonably conclude that precomposed forms are just better and easier. But they're considered legacy: we can't encode every possible combining sequence into a code point, so we might as well go the other way and decompose whenever possible. That's what Normalization Form D is about.
So what would the right definition of a string be that implies everything the author would consider "correct" behaviour?
I've actually run into the same question before and so far my conclusion wasn't even that the implementation of strings is wrong on most platforms, but rather that "string" as a concept just doesn't make sense. Including simple things like the length of a string or concatenation.
If somebody knows a useful definition, I'd be interested ;)
We are almost getting away with this confusion by using brute force - "check out the CPU power of my phone". We're probably struck with it for interweb purposes, but I'd like to see a nice clean alternative, or a cleaned-up version of unicode ... something we could use in small embedded/IoT scale systems.
It can additionally be an indication as to whether regular expression matching works, though that is usually handled by a library, and not the core string types.
My contention was always that if a "string" does not reverse text, then it's no better than an "array<char>". The existence of a "string" type implies it should do more.
It's more meaningless the more a naive algorithm is broken... But shouldn't it at least use those reverse letters Unicode has somewhere? Or is it just a matter of putting the right to left symbol at its start?
is a pretty common idiom in the shell when you want to keep "all but the last N chunks" from some lines of input.
At least for me. I'm sure there's probably a better way, and performance of this is definitely notably poor.
And of course it's not like I'm actually writing the utility. ;)
And for "drop the last X lines or words", you would do better with some specialized command.
(although if you are taking arbitrary user data I suppose there is nothing stopping you doing something stupid like putting an escaped '/' in a filename, which would break this.)
Swift stakes out one of the most interesting points in the design space by completely encapsulating the internal representation and requiring all accessors to be explicit in whether they operate on the string's encoding (e.g. as UTF-8), its Unicode code points, or its grapheme clusters. All positions within the string are represented using abstract iterators that don't reveal their numeric quantities.
Diametrically opposite in the design space is Go (close to C and C++), in which a string is simply an array of bytes, conventionally treated as UTF-8. To access code points requires explicit decoding (or a range loop), and to access grapheme clusters requires a text package.
Ruby is another interesting choice: a string is a pair of a (mutable!) byte array and value indicating how to decode those bytes into code points. This representation works well when you have to deal with strings in obscure encodings as it preserves the original bytes, so, for example, you can read file names from a directory then delete those files.
There are a great many other designs in wide use---Haskell, Java, and Python being the most unnecessarily complicated---but these three designs seemed the most defensible to me.
Raku passes all the tests. He mentions he didn't find any language that upper-cases "baffle" correctly but Raku does.
Here's my REPL test
> my $n = "noe\x[0308]l"
noël
> $n.chars
4
> $n.flip
lëon
> $n.substr(0, 3)
noë
> my $b = "baffle"
baffle
> $b.substr(2, 1).uniname
LATIN SMALL LIGATURE FFL
> $b.uc
BAFFLE
> "noe\x[0308]l" eq "no\x[00EB]l"
True
HN doesn't handle the cat emojis, so they are not included, but they worked fine.We call it NFG for Normalized Form Grapheme.
The JVM and JavaScript backends currently use the built-in string handling features. Which means that they are just as broken as the rest of the code running on them.
---
Of further note, you can use ignorecase in regexes to match "baffle".
"baffle" ~~ /:ignorecase BAFfLE /;This is probably the only right way because 1. it admits multiple implementations (like packing into tagged pointers) and 2. a string isn't simply an array of anything. Strings are made of all of bytes/codepoints/grapheme clusters/words/lines/etc at different levels.
What does make it a bit complicated is that there are about four other different string types to choose between, besides the default String type, which should never be used for anything performance-sensitive.
Though I would bet this mainly has to do with being a language born in the age of the web. It seems like something that would be very hard to back-port (as opposed to designing for it from the beginning), but proper handling of Unicode is practically table-stakes for a new language in today's world. I would guess (I don't know) that Swift, Go, etc also do a good job with this stuff.
Elixir gets it right: https://replit.com/@natanbc/reverse
According to this they added it to the standard library at one point but then broke it back out because the lookup tables were too large
https://play.rust-lang.org/?version=stable&mode=debug&editio...
It looks like it still misses some of the cases (specifically with decomposed vs precomposed).
As a side note, it looks like the playground editor also doesn't correctly render some of the lengths (the cursor is 1 character off on the decomposed case)
Edit: Having read about them, I guess I assumed that if my keyboard inputs an ë it's actually inputting those two underlying characters. Or to put it differently, I assumed there weren't two separate representations for the same character; that the "decomposed" and "precomposed" versions are one in the same, at the byte level. Learned something h̶o̶r̶r̶i̶f̶y̶i̶n̶g̶ new!