edit: I note that HN filters out utf8 emojis..
edit: I note that HN filters out utf8 emojis..
These days one of the systems level programming languages (Rust) is natively UTF-8.
This would have been a huge source of drama in say, 1993 (when UTF-8 is basically brand new) and maybe even in 2003 (by which point it's clear Unicode is a success, but still conceivable that UTF-16 "wins" in some sense) but in 2023 people just shrug - obviously it's UTF-8, why not?
C in principle allows to abstract over the system encoding.
> Until recently, Windows has emphasized "Unicode" -W variants over -A APIs. However, recent releases have used the ANSI code page and -A APIs as a means to introduce UTF-8 support to apps. If the ANSI code page is configured for UTF-8, -A APIs typically operate in UTF-8. This model has the benefit of supporting existing code built with -A APIs without any code changes.
This requires Windows 10 version 1903 or newer. As the version number suggests, this feature is only four years old, and using them prevents your program from working on older versions.
[0]: https://learn.microsoft.com/en-us/windows/apps/design/global...
Note that Rust has OsString to handle both of these cases, but it’s somewhat hard to use.
There's no reason all the rest of your software needs to pay for Microsoft's mistake.
> and converting to/from UTF-8 on every Windows API call is not only a pessimization both in runtime and memory usage, but also introduces complications in having to handle possible conversion failures (involving unpaired surrogate characters, in particular).
If the text isn't actually Unicode, then all you're demonstrating by not handling that in another language is that you didn't care about correctness. There's no magic here, unpaired surrogates aren't somehow actually valid Unicode so long as you stick with C.
How is it a mistake when UTF was not even a thing when the NT kernel was developed?
Traditionally, strings in C are encoding-agnostic. By default, the encoding of the current locale is assumed, and applications are expected to convert to other encodings when specifically needed. Forcing everything into “there can only be one encoding we can work with” isn’t a great solution. One can wish that everything was UTF-8 from the beginning, but that’s not the world we live in.
Except that the geniuses who maintain the extant Unix systems still couldn't bring themselves to at least start transitioning pathnames to utf8, so now we're forever stuck with the terrible OsString cruft and wasting brain and CPU cycles converting from and to OsString in some unsatisfactory manner.
At least ZFS (which also incidentally is pretty much the only non-terrible filesystem) has an utf8only flag.
And there is no serious backwards compatibility problem. You just introduce a mount option to enforce utf-8, flip it to default after a say a decade, and people who then still have file systems with pathnames in latin-1 (or other craziness) can then flip it off and get a few extra decades to migrate their stuff. In the meantime the rest of the world just writes software under the assumption that it will be deployed to non-crazy installs, and software becomes magically more reliable and quicker to write.
And it's not like even 1% of software now would handle non-utf-8 filenames robustly anyway (I mean even if you are aware of the problem and diligent about it, there is absolutely no good way to deal with non-utf8 filenames if you need to output them as text for consumption by humans or other programs, which is close to 100% of cases, because at the very least you will need to do so in error messages if there is an IO problem).
And even if you settle for utf-8 you'll still have to deal with differences with filesystems automatically canonicalizing names or being case insensitive.
Exactly what you would want to happen (this is the actual error message you will get, BTW)?
> (cd /mnt/funkyfilesystem && touch $'\370' && tar cf my-bad-file.tar $'\370')
> tar xf /mnt/funkyfilesystem/my-bad-file.tar
tar: \370: Cannot open: Invalid or incomplete multibyte or wide character
tar: Exiting with failure status due to previous errors
Or are you trying to tell me that your life isn't complete without tar or a git checkout silently dropping some nonsense byte sequences into your filesystem? In both cases you can map the names to something saner explicitly and continue with whatever you were doing, or checkout to e.g. a tmpfs where you flipped off utf8-enforcement. BTW, unless you work with fairly unusual git repos or tar archives on a regular basis, this is very unlikely to ever happen to you.> And even if you settle for utf-8 you'll still have to deal with differences with filesystems automatically canonicalizing names or being case insensitive.
(Case-preserving you mean? I think truly case-insensitive died with 8.3 DOS filenames). These are also annoying, but a much more minor issue, e.g. you don't need to create a new weird string type just because of them.
But hey, if enough tools and applications stop supporting non-utf-8 names, it is possible that in a a few decades you might get what you want as de-facto non-utf8 files will just disappear.
Let's say you get an EILSEQ on accessing the dirent with readdir and now have to debug how you got some corrupted filename you clearly didn't want in the first place, remount and fix it. How is that worse than having the corrupted filename and not knowing about it before something more insidious happens?
Uh, not exactly, not quite. UTF-8 was invented by Unix people who were fed up with all the alternative proposals at the time. Those people did not then go and "fix Unix" in any way regarding this. Those people were none others than Ken Thompson and Rob Pike, who were among the creators of Unix -- you could only have gotten closer to "those people" being the creators of Unix by having had Denis Ritchie at that diner table on that fateful day.
As far as POSIX and Unix `open(2)` and friends go, pathnames are opaque binary with only two special byte values: NUL (0x00, because it's C strings, which are NUL-terminated) and '/' (because that's the path component separator). Any codeset and form that is compatible with that will "work" -- for some value of "work" where if the codeset isn't Unicode in UTF-8 then you'll be sad.
Nothing keeps users / sites from declaring that thou shall use UTF-8 on the filesystem, or, even better, that thou shall use UTF-8 locales only, as the latter is the only simple way to [mostly] get the former.
> At least ZFS (which also incidentally is pretty much the only non-terrible filesystem) has an utf8only flag.
Note that ZFS can't tell if the strings from user-land are in UTF-8. ZFS can only tell if they're not valid UTF-8.
Another good one is being completely honest about your height (mostly just filters out American expats but still).
Anyway, the only reason I used emoticons was because I use this keyboard:
http://www.exideas.com/ME/index.php
... but it apparently had some side-effects.