For added confusion, Rust has a `char` type which is actually 32-bits. You can create arrays of them, but the resulting string would be in utf-32 and thus incompatible with the normal `str` type.
For added confusion, Rust has a `char` type which is actually 32-bits. You can create arrays of them, but the resulting string would be in utf-32 and thus incompatible with the normal `str` type.
let s = "hello world".to_string();
for ch in s.chars() {
print!({},ch);
}
will iterate through a string character by character. That's the most common use of the "char" type - one at a a time, not arrays of them.Although the proper grapheme form is:
use unicode_segmentation::UnicodeSegmentation; // 1.7.1
let s = "hello world".to_string();
for gr in UnicodeSegmentation::graphemes(s.as_str(), true) {
print!("{}", gr)
}
This will handle accented characters and emoji modifiers.
A line break in the middle of a grapheme will mess up output.By the way, open season for proposing new emoji starts tomorrow.[1]
[1] http://blog.unicode.org/2020/09/emoji-150-submissions-re-ope...
$ cargo install unicode_segmentation (chokes)
$ cargo install unicode-segmentation (seems to work)
and in Cargo.toml
[dependencies]
unicode-segmentation = "1.7.1" (seems to work)
yet in the code, its:
use unicode_segmentation::UnicodeSegmentation;
Why couldn't they be consistent using a dash vs. an underscore ?You can’t use -s in Rust identifiers, so they need to be normalized to _ to be referred to in code.
Huh? That seems clearly untrue, eg `fnfoo` vs `fn foo` or `x && y` vs `x & &y` or `x<<shift_or_type > ::foo` vs `x < <shift_or_type>::foo`? Presumably for some or all of those, one version ends up being a error (eg bitwise and with a pointer from `x & &y` probably doesn't work), but that's not at the level of tokenization.
It can't be interpreted as a single identifier.