Edit: this was supposed to be a response to the other response to Steve’s comment. On my phone now, can’t easily edit.
However, rust is very well-engineered, and I’d be interested in knowing what factors made its decision a win.
The discussion seemed to me to support the conclusion that Rust will not introduce SSO right now, but the conclusion the maintainers actually drew was that Rust will not introduce SSO ever. Of course, that could always be revisited!
1. Conventions and prevailing idioms. In C++ copying strings is common, both unintentionally due to the implicit copy constructor and intentionally as a means of defensive programming. In Rust, copying strings is relatively uncommon: Rust has no copy constructors (copying must be done explicitly with the `.clone()` method), and the borrow checker obviates the need for defensive copies and therefore passing string slices is overwhelmingly preferred (and recall that string slices in Rust (`&str`) are 16 bytes, in comparison to 24 bytes for C++'s `std::string`).
2. Alternative optimizations. For example, in Rust, initializing a `Vec` (and by extension `String`) does not perform a heap allocation for zero-length values. IOW, you can call `String::new()` (or even `String::from("")`) in a hot loop without worrying about incurring allocator traffic. Any hypothesis that short strings are more common than long strings will likely also hypothesize that the most common short string length is zero (and this appears to be borne out in practice; this optimization is really important for `Vec`, so important that, for Servo, it often makes `Vec` faster in practice than their own custom `SmallVec` type which strictly exists only on the stack), so this optimization goes a fair ways towards satisfying the "I actually do have many short strings and I'm not just copying around the same string a lot" use case.
3. Weighing trade-offs. The pros of SSO are better locality and less memory usage, with the cons of larger code size and additional branches on various operations.
Putting it all together, this gives the Rust devs reasonable incentive to take the conservative path of not implementing SSO for the default string type. That's not to say that we'd be incapable of finding Rust code that would benefit from SSO, or that C++ chose their default incorrectly (different contexts allow different conclusions), or that the Rust devs will always have this stance (if performance benefits of SSO were to be clearly demonstrated on Rust code in the wild in such a definitive way as to justify changing the default behavior, then I don't think they couldn't find a way to make it happen).
G++ used to have CoW strings, but I believe they dropped them because for common cases the atomic reference counting was more expensive than potential copying.
Inlining strings save space and allocator traffic at the expense of a more complex string implementation and (depending on the platform/implementation) strings that do not start word aligned.
There are other flags you might want to put on an underlying text sequence:
- that the text is exclusively ASCII in UTF-8 strings. This means the character count and indexing can be optimized to constant-time operations
- latin-1 in UTF-16 strings. This means you can do an alternate 8-bit encoding for space savings, as well as constant-time count/indexing.
- mark that the string represent a single grapheme cluster (aka a printable character). This may allow you to combine unicode text processing to a single data structure.
- mark that the representation is a slice. In this case, you'd typically store an offset rather than a capacity. This again helps you combine higher level types on top of a single data structure and text processing libraries.
These features (and especially combinations of them) impact the hot path for text processing. If your language is already doing custom allocation (such as using a copy constructor), the allocator savings are nil, so it becomes a space vs code trade-off.
That said, as mentioned the other features do make this hurt less for many cases.