D's advantage over Rust is its familiar syntax and semantics. C, C++, Java, C#, Python, etc. programmers feel at home once they learn some differences. I heard others say "D is what C++ should have been" and "Compiled Python".
I did use Rust for a brief period in a project where the experienced Rust programmer among us was throwing '&' characters here and there to make the code compile, seemingly randomly in many cases. Personally, I remember fighting with impedance issues with the many different string types of Rust. All of this spells a steeper learning curve to me.
I think D is familiar to programmers of many other languages.
const(char)[]
meaning you can treat it like any other array.https://dlang.org/phobos/std_utf.html#.byUTF
"Throws: UTFException if invalid UTF sequence and useReplacementDchar is set to UseReplacementDchar.yes"
My guess is that this is a mistake and should instead say UseReplacementDchar.no since it makes sense to throw an exception if you can't use U+FFFD here, rather than do both.
Anyway, in my view this is bad the same way the Billion Dollar Mistake is bad, and Rust made the right choice here. Arrays of stuff are great, but they aren't strings. Having to sprinkle "or maybe not" cases all over these libraries because of course these might not really be strings, results in exception fatigue from your developers, which in turn results in lower quality software and more effort for the conscientious developers who stick it out.
D's strings are less stupid than C's (and thus some of the C++ strings) but they're still just arrays which are maybe but maybe not actually text.
On the other hand, D's dstrings are more like text because they are not only UTF-32 but also random-accessible code points. (D does not address multiple representations of graphemes at language level. For example, at language level, ğ is different from "g and combining breve" but there are std.uni and std.utf modules that help.)
Bytes. It's an array of bytes. D's char type isn't actually restricted to UTF-8 code units, char x = '\xFF'; works just fine even though that's not UTF-8.
Having string be a magic builtin type does not eliminate the problem of dealing with invalid UTF sequences.
Invalid UTF sequences are inherent to the Unicode design, and programmers are left on their own to deal with it. The options are:
1. ignore them
2. use the replacement char
3. throw an exception (or other error indication)
D enables the programmer to pick which they need, on a case by case basis.
If you have type safety, you can make the choice just once.
Rust's String::from_{utf8,utf16}_lossy turn valid UTF-8/16 sequences into strings, and "fix" invalid ones with U+FFFD
Meanwhile String::from_{utf8,utf16} attempt the same but with an Err instead of replacement on failure if that's what the programmer wants.
Imagine if all D's numeric functions took the same attitude as its string functions, insisting on being passed arrays of bytes so that each function can parse those bytes, decide if this is actually a 16-bit unsigned integer (for example) and if so do what's expected otherwise perhaps return an error. We'd spot right away that this was not a practical design.
D's choices here are conventional, but I've come to expect a lot more and so I'm disappointed when I can't have it.
#23405 was resolved as fixed a week ago. It isn't fixed. I guess at least I didn't waste my time filing the bug.
It's because Rust gives very low-level control over strings. If what you want is a string the same way as they work in other languages, including D, then use String. If you want very fine control over memory allocation or want the string to be fixed length, then you use a type that guarantees those properties.
String will work with anything, at the cost of having little control over its allocation or representation.