EDIT: oh lord Hacker News strips emoji. You get the idea even though HN ruined the illustrations. Not my fault HN is broken.
Just because the worst case can't improve doesn't mean that making the average case better is worthless.
(I don't think the talk where the stuff about Arabic was mentioned was Plain Text by Dylan Beattie, but I haven't re-watched it to confirm. So maybe it is. Can't remember the name of any other talk about the subject right now.)
IIRC the lower levels of Windows will happily work with filenames that are not valid Unicode strings, for example if you use the kernel API rather than Win32.
But what about Win32? If you create a file before normalization and then open it using the normalized form, will it open the same file or return file not found?
What about other systems? For example AWS' S3 allows UTF-8 keys, with no mention of normalization[1].
On the phone so can't try myself right now.
Anyway for general text I agree, but for identifiers, filenames and such I prefer to treat them as opaquely as possible.
[1]: https://docs.aws.amazon.com/AmazonS3/latest/userguide/object...
Just made two files named ä.txt which happily sat next to each other, one being NFC and other NFD.
So yeah, don't mess with the normalization of filenames.
I don't use Windows, so I can't check. Linux literally allows any arbitrary byte except for 0x00 and 0x2F ('/' in ASCII/UTF-8). It's a problem for programming languages that want to only use valid Unicode strings, like Python. Rust has a separate type "OsString" to handle that, with either lossy conversion to "String" or a conversion method that can fail. Python uses the custom use Unicode range to represent invalid byte sequences in filenames. It's all a mess. JavaScript doesn't give a damn about the validity of their UTF-16 strings.
(Note that Rust's OsString is different from it's CString type. Well, I guess under Unix they're the same, but under Windows OsString is UTF-16 (or "WTF-16", because it isn't actually valid UTF-16 in all cases).)
I tried using U+13161 EGYPTIAN HIEROGLYPH G029[1], which resulted in a string of length 2 as expected.
Using both chars (code units) and just the first char (code unit) worked equally fine. In Windows Explorer the first one shows the stork as expected, while the second shows that "invalid character" rectangle.
So yeah, treating filenames as nearly-opaque byte sequences is probably the best approach.
Äh, nu går vi och gör något annat.
Win32/NTFS treats them as two different filenames, as it doesn't normalize them before storing/comparing.
The Linux kernel doesn't validate filenames in any way, so a filename in Linux can contain any byte except 0x2F ('/', which is interpreted as directory separator) and 0x00 (which signals the end of the byte string).
ETA: of course some file systems have other limitations, for example '\' is not valid in FAT32.