Beyond that though, &[T] implies a slice of Ts, that is, multiple Ts in a row. But a &str is a slice of a single string. So &[str] would feel wrong; that is, a &str is a slice of a String or another &str or something else, but isn't like, a list of multiple things. It's String, not Vec<str>.
Basically, Strings are just weird.
Are you talking about them being a container with a pointer to the actual array, and also a size and etc?
For added confusion, Rust has a `char` type which is actually 32-bits. You can create arrays of them, but the resulting string would be in utf-32 and thus incompatible with the normal `str` type.
let s = "hello world".to_string();
for ch in s.chars() {
print!({},ch);
}
will iterate through a string character by character. That's the most common use of the "char" type - one at a a time, not arrays of them.Although the proper grapheme form is:
use unicode_segmentation::UnicodeSegmentation; // 1.7.1
let s = "hello world".to_string();
for gr in UnicodeSegmentation::graphemes(s.as_str(), true) {
print!("{}", gr)
}
This will handle accented characters and emoji modifiers.
A line break in the middle of a grapheme will mess up output.By the way, open season for proposing new emoji starts tomorrow.[1]
[1] http://blog.unicode.org/2020/09/emoji-150-submissions-re-ope...
$ cargo install unicode_segmentation (chokes)
$ cargo install unicode-segmentation (seems to work)
and in Cargo.toml
[dependencies]
unicode-segmentation = "1.7.1" (seems to work)
yet in the code, its:
use unicode_segmentation::UnicodeSegmentation;
Why couldn't they be consistent using a dash vs. an underscore ?You can’t use -s in Rust identifiers, so they need to be normalized to _ to be referred to in code.
Huh? That seems clearly untrue, eg `fnfoo` vs `fn foo` or `x && y` vs `x & &y` or `x<<shift_or_type > ::foo` vs `x < <shift_or_type>::foo`? Presumably for some or all of those, one version ends up being a error (eg bitwise and with a pointer from `x & &y` probably doesn't work), but that's not at the level of tokenization.
It can't be interpreted as a single identifier.
Not a Rust expert, but to my understanding, 'string' is an array of characters, not necessarily living in the heap.
The object 'String' (with capital S) might be, but with &str I can have a constant array of characters that is not in the heap and not even in the stack: it's in the code (.text code segment[1]?, or EEPROM or flash, for embedded folks).
If I understood it correctly, &str is a slice with a pointer to that piece of code. Of course being const, it cannot be changed. I can copy it into a String and manipulate it (append, insert, etc.), and that's where the heap is used.
&str, as an hypothetical struct, it may be in the stack, initialized in-scope with (probably) a pointer to somewhere in .text plus some bytes for length or other information.
Just as `[u8]` is to `Vec<u8>`, `[char]` is hypothetically to `Vec<char>`.. and `Vec<char>` is basically a `String`, no?
_edit_: Though looking at the docs, `char` is 4 always bytes, so i guess that's where the breakdown would be? `char` would need to be unsized i guess, but then it would be an awkward `[unsized_char]`, which is like two unsized types.... hence `str` maybe?
`str`'s data layout happens to be `[u8]`, but it's type provides additional guarantees about the structure of the data within it's internal `[u8]` (for example, forbidding sequences of u8 that don't encode valid utf-8).
I think `char` _would_ work, if it was similarly unsized like a single piece of `str`. The Problem is .. as i see it, that `[unsized_char]` seems odd.
`str` is not a slice; this is itself already a wrong statement. A slice is a dynamically sized type, a region of memory that contains any nonnegative number of elements of another type.
A `str` is dynamically sized, but is not guaranteed to contain a succession of any particular type. It's simply a dynamically sized sequence of bytes guaranteed to be Utf8.
`strs` aren't slices; all they have in common with them is that they are both dynamically sized types.
`Vec<char>` is also not the same as `String`; a string is not a vector of `char`, which is already a type that has a size of 4 bytes.
This all results from that Utf8 is a variable width encoding, and since slices are homogeneous, all elements have the same size.
I've taken to calling "String" "owned string" but that's a little unwieldly and not quite suited for complete Rust beginners who don't know the concept of ownership yet. Same with "string reference" for "str".
Why is 'str' a "primitive type"? What about 'str' means it has to be primitive, instead of being a light-weight wrapper around a '&[u8]' (that obviously enforces UTF8 requirements as approriate).
let hello: &[u8] = "Héllo".as_u8_slice();
// hello[1] != "é", I believe it would be the bottom half-bytes of "H".
// hello[2] != "é", I believe it would be the top half-bytes of "é".
To use &[u8] would be very non-ergonomic.It makes sense to have a special type because it can supports all of this stuff through separate methods on the type. (Or through nearby APIs). It’s confusing, but I think it’s the right choice.
Although, I think the most confusing thing about rust strings isn’t that &str isn’t &[u8]. It’s that &str isn’t just &String or something like that.
Really, there's no single, good option that covers all use cases. Which is why Rust ended up with multiple string types, along with ways to convert between them.
Yeah I know. This is really confusing when learning rust. Coming from another language where there's a single rust type its really tricky to figure out when to use what!
For reference, I now use String for large strings, string builders, and when the string is owned - like in a struct. And &str when taking in a string as an argument to a function. Or when iterating or something like that where you don't want to move the string. And then for small owned strings (like labels or IDs) I reach for the InlinableString crate. That handles small strings inline and puts large strings on the heap, like what Apple does for Swift and Objective-C.
I understand why its this way; but its a pity the best answer rust has is so complex and confusing.
In addition, even if you gave up on O(1) indexing by grapheme cluster index, anything faster than O(n) indexing would require some sort of tree data structure to translate grapheme cluster indices to byte indices. Such a structure would be expensive to maintain and not needed the vast majority of the time; in a language that prioritizes performance, it simply couldn’t be part of the default string type no matter what the syntax.
There was a PR implementing it as a wrapper around &[u8], but it didn't really provide any actual advantages, so it was decided to not do that.
The majority of the functions on `&str` seem to make sense for all `&[T + !Sized]` where `type str = [unsized_char]`.
That said, I'm also not sure in the general case.
e.g. Think of it as an enumeration like :
unsized_enum unsized_char {
Char1(u8),
Char2(u8, u8),
Char3(u8, u8, u8),
}
let c: unsized_char = ...; // is a complied error: “cannot assign a variable stack size.” or something.
let parsed: &[unsized_char] = unsafe { /* parse some bytes */}; // this complies because the implementation of some hypothetical &[T + !Sized] appropriately abstracts the allocated memory region.
Unlike a normal enum it doesn’t always take 3 bytes of memory, it’s variable memory such that [unsized_char] might only take 1 byte if there is only 1 unsized_char::Char1 inside.Of course this means that you can’t index into the slice because the boundary is not clear.
What I’m describing is not different than the internals of str. I’m just trying to armchair design a generic implementation such that str and slice are more unified.
With the actual `&str`, this treats 2 and 4 as byte indices. But if you’re going to treat str as a slice of a specific type, the indices really should be in units of that type. `&s[2..4]` should return the second and third grapheme clusters, or codepoints, or whatever you want to define `unsized_char` to be. The same applies to `len()` and all the other methods that return or accept indices.
But if it’s in units of `unsized_char`, how do you actually implement indexing? You could do an O(n) scan from the beginning of the string, but that’s clearly unacceptable. Or you could set up some kind of acceleration data structure behind the scenes, but that would be expensive to maintain, and would conflict with Rust’s goals of explicitness and cheap FFI.
If on the other hand you keep indexing with byte offsets, you end up with almost no operations that work, and do the same thing, for both &[unsized_char] and normal slices. You wouldn’t be able to write any useful generic code that operates on &[T] for T: ?Sized, at least not without branching on which kind of slice it is. And even from a human user’s perspective, the syntactic similarity between two things that work differently could confuse more than it helps. Again, a matter of explicitness.
As a side note, it is possible to define your own "unsized" slice type which wraps `[u8]`. This can be useful for binary serialization formats which can be subdivided / sliced into smaller data units.
I don’t think any other languages do that. Instead, most of them are implementing as much as they can while viewing the storage as a blob of UTF8/UTF16 bytes/words, and throw exceptions from the methods which interpret the data as codepoints.
Strings are used a lot in all kinds of APIs. For instance, strings are used for file and directory names. The OS kernels don’t require these strings to be valid UTF-8 (Linux) or UTF-16 (Windows).
To address the use case, Rust standard library needs yet another string type, OsString. This contributes to complexity and learning curve.
So, they offloading complexity to programmers. Being such a programmer, I don’t like their attitude.
> other languages will fail at runtime
In practice, other languages usually printing squares, or sometimes character escape codes after backslash, for encoding errors in their strings. That’s not always the best thing to do, but I think that’s what people want in majority of use cases.
The only choice is whether it's explicit and managed by the language, or hidden, and you need knowledge and experience to handle it yourself without language's help. If you want "squares" for broken encoding, Rust has `to_string_lossy()` for you. It's explicit, so you won't get that error by accident.
Avoiding "mojibake" in other languages is usually a major pain. For example, PHP is completely hands-off when it comes to string encodings. To actually encode characters properly you need to know which ini settings to tweak, remember to use mb_ functions when appropriate, and don't lose track of which string has what encoding. There's internal encoding, filesystem encoding, output encoding, etc. They may be incompatible, but PHP doesn't care and won't help you.
I would want it to be implicit.
Ideally, for rare 20% of cases when I care about UTF encoding errors, I’d want a compiler switch or something similar to re-introduce these checks, but I can live without that.
> For example, PHP
When you compare Rust with PHP it’s no surprise Rust is better, many people think PHP is notoriously bad language: https://eev.ee/blog/2012/04/09/php-a-fractal-of-bad-design/
I like C# strings the best, but I also have lots of experiencer with C++, and some experience with Java, Objective-C, Python, and a few others. None of them have Rust’s amount of various string types exposed to programmers, many higher-level languages have exactly 1 string type.
Interestingly, some dynamic languages like swift use similar stuff internally, but they don’t expose the complexity to programmers, they manage to provide a higher-level abstraction over the memory layout. Compared to Rust, improves usability a lot.
Then simply use the `lossy` function. Or create an alias to it if you don't like the name.
A &str is really a different kind of thing from other slices. In any other slice, each element in the slice takes up a constant number of bytes, but this is not the case for a &str.
Strings in Rust are (normally) represented as UTF-8. Both `String` and `str` represent data that is guaranteed to be valid UTF-8.
This means that if Rusts UTF-8 strings were represented as normal slices, they would have to be slices of UTF-8 code-units.
Rust wants to provide a safe and correct String data type, and therefore, indexing a string on a byte (code-unit) level would be incorrect behavior.
Having a custom type `String` and `str` instead of just a `Vec<u8>` enables you to have more correct behavior implemented on top of the data type that doesn't implement normal slice indexing and such.
---
As a note, even though you probably don't want to normally, you can quite easily access the backing data of your string using `String::as_bytes`
Calling &str a "string slice" is really more about the contrast with String, and how the relationship there mirrors the relationship between &[T] and Vec<T>. It's more of an analogy than a concrete description of the interface.
&OsStr and &Path are the same way.
But, str is a subset of &[u8], the type's contract is that it is unsafe to have non valid UTF-8 data, hence https://doc.rust-lang.org/std/str/fn.from_utf8.html can error, offering the unsafe variant https://doc.rust-lang.org/std/str/fn.from_utf8_unchecked.htm...
This is all very different than &[char] which would be an array of 4 byte characters (or, a UCS4 string)
&[T] is an “array-slice”, even if it's called just “slice”.
See this example ([..] is the syntax to create a slice of something): https://play.rust-lang.org/?version=stable&mode=debug&editio...
Because unicode, &[T] make easy to write wrong code (that asume here T = Char).
It CAN'T be Char, because char is larger than u8:
https://doc.rust-lang.org/std/primitive.char.html
and it mean unicode point.
In other words: Rust is using types to PREVENT the wrong behavior.
https://doc.rust-lang.org/stable/std/primitive.str.html#meth...