Not a Rust expert, but to my understanding, 'string' is an array of characters, not necessarily living in the heap.
The object 'String' (with capital S) might be, but with &str I can have a constant array of characters that is not in the heap and not even in the stack: it's in the code (.text code segment[1]?, or EEPROM or flash, for embedded folks).
If I understood it correctly, &str is a slice with a pointer to that piece of code. Of course being const, it cannot be changed. I can copy it into a String and manipulate it (append, insert, etc.), and that's where the heap is used.
&str, as an hypothetical struct, it may be in the stack, initialized in-scope with (probably) a pointer to somewhere in .text plus some bytes for length or other information.
Are you talking about them being a container with a pointer to the actual array, and also a size and etc?
For added confusion, Rust has a `char` type which is actually 32-bits. You can create arrays of them, but the resulting string would be in utf-32 and thus incompatible with the normal `str` type.
let s = "hello world".to_string();
for ch in s.chars() {
print!({},ch);
}
will iterate through a string character by character. That's the most common use of the "char" type - one at a a time, not arrays of them.Although the proper grapheme form is:
use unicode_segmentation::UnicodeSegmentation; // 1.7.1
let s = "hello world".to_string();
for gr in UnicodeSegmentation::graphemes(s.as_str(), true) {
print!("{}", gr)
}
This will handle accented characters and emoji modifiers.
A line break in the middle of a grapheme will mess up output.By the way, open season for proposing new emoji starts tomorrow.[1]
[1] http://blog.unicode.org/2020/09/emoji-150-submissions-re-ope...
$ cargo install unicode_segmentation (chokes)
$ cargo install unicode-segmentation (seems to work)
and in Cargo.toml
[dependencies]
unicode-segmentation = "1.7.1" (seems to work)
yet in the code, its:
use unicode_segmentation::UnicodeSegmentation;
Why couldn't they be consistent using a dash vs. an underscore ?Huh? That seems clearly untrue, eg `fnfoo` vs `fn foo` or `x && y` vs `x & &y` or `x<<shift_or_type > ::foo` vs `x < <shift_or_type>::foo`? Presumably for some or all of those, one version ends up being a error (eg bitwise and with a pointer from `x & &y` probably doesn't work), but that's not at the level of tokenization.
It can't be interpreted as a single identifier.
You can’t use -s in Rust identifiers, so they need to be normalized to _ to be referred to in code.