In defense of all those usages (except the LSP one, which is indefensible), the original pitches for Unicode[1] literally said that it was intended to be fixed-width, an international ASCII of sorts; that was to be achieved by restricting it to “commercially-relevant” text (and Han unification). Then it turned out there are plenty of very rare Han characters people really, really want to see encoded (for place and personal names, etc.). Of course, in hindsight an encoding with nontrivial string equivalences (e.g. combining diacritics) was never going to be as simple to handle as ASCII.
Getting the hi and lo UTF-16 code points can use the BSD licensed
uint16_t le16toh(uint16_t little_endian_16bits)
and uint16_t htole16(uint16_t host_16bits)They're C99 standard functions and should be converting between "wide strings" and "multibyte strings", which should be native UTF-16 and UTF-8 if your current locale is an UTF-8 locale.
Apparently this works on Windows since Windows 10 version 1803 (April 2018).
There are also "restartable" variants `wcsrtombs` and `mbsrtowcs` where the conversion state is explicitly stored, instead of (presumably) a thread-local variable.
C11 added "secure" variants (with an `_s` suffix) of all these which check the destination buffer size and have different return values.
https://learn.microsoft.com/en-us/windows/apps/design/global...
https://stackoverflow.com/questions/72143553
> it looks like switching back to UTF-16 would unblock your experience with Resource Editor
Edit: reportedly changed in VS2017 15.9
Either way, I suggest to the readers who might feel upset over this statement to explore something outside of C and C++, liking which, when it comes to strings, is nothing short of Stockholm syndrome.
I'm working on a UTF-8 string library for C# and across the last 6-8 months explored string design in Rust, Swift, Go, C, C++ and a little in other languages. C and C++ were, by far, most horrifying in the amount of footguns as well as the average effort required to perform trivial operations (including transcoding discussed here).
Strings are not easy. But it does not mean their complexity has to be unjustified or unreasonable, which it is in C++ and C (for reasons somewhat different although overlapping). The problem comes from the fact that C and C++ do not enjoy the benefit of the hindsight that Rust had designing its string around being UTF-8 exclusive with special types to express either opaque, ANSI or UTF-16 encodings to deal with situations where UTF-8 won't do.
But I assure you, there will be strong negative correlation here between complaining about string complexity and using Rust, or C#/Java or even Go. Keep in mind that Go's strings are still a poor design that lets you arbitrarily tear code points and foregoes richness and safety of Rust strings. Same, to an extent, applies to C# and Java strings, though they are also safe mostly through a quirk of UTF-16 where you can only ever tear non-BMP code points, which happen infrequently at the edges of substrings or string slices as the offsets are produced by scanning or from known good constants.
If, at your own peril, you still wish to stay with C++, then you may want to look at QString from Qt which is how a decent string type UX should look like.
I wrote more about this here: https://blog.burntsushi.net/bstr/#motivation-based-on-concep...
I mention gecko as an example repository that contains data that isn't valid UTF-8. But it isn't unique. The cpython repository does too. When you make your string type have the invariant that it must be valid UTF-8, you're giving up something when it comes to writing tools that process the contents of arbitrary files.
You sometimes need a way to operate on entirely arbitrary sequences of bytes. This is mostly easy, it's been a long time since non-octet bytes were relevant in most situations, so the vast majority of the time you can just assume they're all octets.
You sometimes need a way to operate on arbitrary text. This inherently requires knowing how that text is encoded, but as long as you know that it's mostly easy.
You sometimes need a way to operate on text-like things that aren't necessarily text, like the output of old CLI programs that used the BEL character to alert the user to events. Or POSIX filenames. Or text where you don't know the encoding. This is where the bugs lie, where we make unchecked assumptions about the data that turn out to be invalid.
You'll notice that I didn't say "Go's string design is good and we should all use it." I made an argument that's Go's string design is not poor and provided an argument for why that is. In particular, I described trade offs and a particular pragmatic point on which abdicating the UTF-8 requirement makes for a more seamless experience when dealing with arbitrary file content.
> but as long as you know that it's mostly easy. [..snip..] Or text where you don't know the encoding.
You don't know. That was my whole point! I gave real-world concrete examples of popular things (Mozilla and CPython repositories) that contain text files that aren't entirely valid UTF-8. They are only mostly valid UTF-8. If I instead treated them as malformed and refused to process them in my command line utilities or libraries, I would get instant bug reports.
> Go strings aren't necessarily text.
I would generally consider this to be an incorrect statement. The more precise statement is that Go strings may contain invalid UTF-8. But the operations defined on strings treat strings as text. For example, if you iterate over the codepoints in a Go string, you'll get U+FFFD for bytes that are invalid UTF-8. By your own reasoning, U+FFFD must be considered text because it can also appear in a Rust &str/String. Despite the fact that a Go string and a []byte can represent arbitrary sequences of bytes, a Go string is not the same thing as a []byte. Aside from mutability and growability, the operations on them (both those provided as a library and those provided by the language definition itself) are what distinguish them. They are what make a `string` text, even when it contains invalid UTF-8.
There are deep trade offs here, but the UTF-8-is-required does have downsides that UTF-8-by-convention does not have. And of course, vice versa.
[1]: https://docs.rs/bstr
The choice to have this in a language with the safe/ unsafe distinction works very nicely because in so many languages you'd have this promised UTF-8 type and then in practice everybody and their dog uses the unsafe assumed conversion because it's easier, but in Rust you're pulled up short because that conversion needs an unsafe block, your local style may require (and good practice certainly does) that you address this with a safety comment explaining why it's OK and... it just isn't, in most cases. So you write the safe conversion instead unless you really need not to. This is a really nice nudge, you can do the Wrong Thing™, but it's just easier not to.
MFC had a similar problem for years. It had a CString class which was ANSI or Unicode, depending on how your compiled your app, but moderately often you needed the other one, so it should have had a CStringA and CStringW too, with nice conversions between them.
In most of these file format cases what you've got is &[u8] or &[u16] and maybe it's the NonZero variant instead, so I think it's fine to be explicit that's what is going on and maybe in the process remind you to check - is this data UTF16LE? UTF16 with a BOM? UCS2 with a nod and a wink? Just arbitrary 16-bit integers and good luck?
But like I said, I favoured "picking nits" long before I learned Rust, so mileage may vary.
That said, even now in 2024, it's not clear how much of a bet Windows is making on UTF-8 versus UTF-16.
Since references can't have destructors (they don't own the data like an OsString does), it means that the standard library can't give you a newly-allocated string without leaking it. Since obviously it isn't going to do that, the &OsStr must instead just act as a view into the underlying &str. And the conversion can't enforce any extra restrictions on the input string without breaking backward compatibility.
The overall effect is that whatever format OsStr uses, it has to be a superset of UTF-8.
[0] https://doc.rust-lang.org/1.77.2/std/ffi/struct.OsStr.html#i...
Then when 32 bit came to be we had the other variations on top.
Got to love leaky abstractions.
Reference to “Celestial Emporium of Benevolent Knowledge”?
As a meta-joke I was also considering:
APL (j/k)
It's a problem with Unix filenames where the encoding is just a convention. A Python program that doesn't take great care can crash on a parameter that takes a filename even if that filename is just passed to a function like 'open' so no sanitisation or conversion is necessary.
This is a very real problem in hivex, our Windows registry library, where the Python bindings don't really work well. The Windows registry is a hodge podge of random encodings, essentially whatever the program that wrote the registry key thought it was using at the time. When parsed through Python as a string this means you'll get unicode decoding errors all over the place.
Also more in this article: https://changelog.complete.org/archives/9938-the-python-unic...
UNIX filenames are NOT necessarily printable text, they're byte strings. Don't treat them as printable text. They're sequences of bytes not containing 0x00 or 0x2F, but with no encoding.
Same for Windows registry keys. Don't mistake byte strings for text.
Text is a byte string with an encoding which describes which byte values are valid and which characters each byte (or sequence of bytes) corresponds to. If the encoding information is discarded, it stops being text and becomes just a byte string.
But they nearly always also happen to be printable text. Treating them as printable text is soooo convenient 99.9% of the time.
Having them be almost printable text, but not quite, is an API design wart that is begging to be used incorrectly. It's the inverse of an affordance.
The POSIX filename rules are a "here be dragons" sign. If you can avoid dealing with them, it's best to do so. If you can't avoid it, you'll need to parse your inputs into valid text, and fall back to a byte string handling path (or just fail with error) if they're not text. You can't safely just assume they're text & treat them as such, it doesn't end well.
(IIRC, at least one of the BSDs has actually moved to forcing filenames to be UTF-8 and refusing path names that aren't UTF-8. Would only that Linux moved down that path as well so that we could be done with this farce.)
I'd agree it'd be nice if we could restrict filenames to valid UTF-8. But that's not the API that existing filesystems provide, nor what (most of) the existing OSes enforce.
Threading both of these needles at once basically requires viewing paths as potentially-invalid encoded text.
Don't get me wrong it's still painful and annoying and bug prone, but the point is, it's encodings, it's always going to suck no matter what.
How else would you decode a string without knowing its encoding? You can either guess (and risk invalid result/decode errors) or store this information somewhere. This is universal and true in every language. In most cases today people choose to guess utf-8.
>>> 'abc' + 'def'[1]
abce
>>> b'abc' + 'def'[1]
TypeError: can't concat int to bytes
It's been responsible for a lot of bugs in my code. They copied the "bytes is an array of integers" thing from Java. Big mistake. Python is not Java.Them treating filenames as strings is another bug factory on Linux. Almost no one unit tests filenames with invalid encodings, so the result is a whole pile of python3 programs will fall when given perfectly valid input, whereas those same programs in python2 were fine.
It's a very odd outcome given python harks from Linux.
Rust, C#, Java and Go are fairly straightforward in this regard.
`Encoding.{Name}.GetString/GetBytes`
Just not my cup of tea though.