APFS is not safe to use with names which have Unicode normalisation issues
eclecticlight.co
eclecticlight.co
It's the logical end product of the odyssey from Apple's original philosophy of a resource and data fork model for files to the UNIX stream-of-bytes model for files. The UNIX model traditionally kept metadata about files separate (anyone else remember naughtily hand-editing directory files with sed(1), back in the day?) but gradually picked up a bunch of new stuff that has to be stored somewhere. Meanwhile, the Mac platform adopted UNIX binaries with the move to NeXTStep underpinnings (we're going back to the late 90s here, and the adoption of OSX over traditional MacOS), obviating the need for the original resource bundle, which got rid of those annoying errors about missing bundle bits but left us with a legacy of .DS_Store turdlets in every directory to hold file metadata that was formerly stored in the resource fork.
As a destination it's laudable — design consistency is almost always laudable — but it's the end goal of a very messy process and it looks like APFS isn't quite there yet; NSFileManager and NSURL need some way of distinguishing files with different unicode representations, and the Finder in particular needs to be robust. I'm guessing this is going to be fixed when High Sierra finally ships, but isn't supported in Sierra at present, hence the OP's alarm at the situation.
Also, in regards to the headline, there are tens or maybe over a hundred million non-english speakers using iOS already running APFS...
You did, 11 times so far this thread. I counted.
"Field type" in an RDBMS controls what can be written (e.g. "valid UTF8 strings"), what will be read back (e.g. the use of the Unicode replacement character), and what special values like NULL will cast to.
An RDBMS field's collation controls how values in the field will compare for equality, and what will happen when you sort on that field.
Unique constraints in RDBMSes function on top of collation—so if the collation says two values are equivalent, the unique constraint will prevent you from inserting the new one.
Great properties, if you can get them. It's too bad that the filesystem API is as low-level as it is, actually; if POSIX had some concept of "readdir(2) pre-sorted by a given field" like e.g. Windows does, then you could require that userland programs rely on FS-level collation rather than allowing them to collate the resulting values however they feel like (usually meaning "naively.")
And the n-i string comparison function produces a boolean if I remember correctly, not a trinary. So it doesn't define a collation. But it could define a collation, that's true.
The problem with moving sorting into the kernel is that you now need to have the collations there (English? French? something else? "Unicode" is not enough), and user-land needs to tell the kernel what collation to use for any given process or thread or system call. That's ETOOMUCHWORK for everyone, so it doesn't happen.
Moreover, like it should be EBADIDEA.
The former resource fork is in ._$filename. You won't see it, unless you copy the file to smb share or zip it.
However, 99%+ of Mac users do use FAT-formatted USB sticks, ZIP files, or other fs/mechanism/whatever that does not support resource forks where the compatibility littering kicks in, so they, or the people they share their files with, will see ._$filename files too.
On the other hand, Finder litters with .DS_Store files everywhere, even on HFS+.
I really doubt that 99%+ of Mac users use flash drives / zip files, though. I suspect that well more than 1% of users never use anything like that. A lot of people only have a single computer and share through e.g. just mailing files to people or using Dropbox, or not even that.
I'm not sure that there are more than 1% of mac users do not exchange files with other people, possibly on other platforms. What do they use the computer for, then? iPad would be more suitable then.
My position on this is mixed. I've had to write code to deal with normalization changes in archives and it's quite tedious. The HFS+ approach of normalization is in many ways the best choice considered in isolation but in practice it's really expensive to support since most filesystems, tools, APIs, etc. predate Unicode and everyone else chose the bag of bytes approach.
The comparison with other filesystems does not hold since applications for other OSs have always been developed with no normalization at FS level, and hence it was done by the applications, or through the use of high-level OS APIs. Mac applications, on the other hand, expect it to be the responsibility of the filesystem. This is explained pretty clearly in the article.
On a side note, I wish Apple would have taken this opportunity to switch normalization from NFD to NFC, which basically everything else uses. The distinction causes complexity and often issues in software which share data between Apple platforms and other platforms, such as version control systems for instance.
EDIT: according to pilif's comment, they did, which is awesome!
> Mac applications, on the other hand, expect it to be the responsibility of the filesystem. This is explained pretty clearly in the article.
Again, the article made a huge sweeping claim without supporting it. That's simply not true in either way – many apps on every OS don't handle this at all, some handle it consistently everywhere, and what a “Mac application” means varies widely from “clean, modern Cocoa” to “uses a lot of C, etc. libraries”, “cross platform C++”, “C# port”, “Electron shell”, etc. You can't make any statement which is true for every single one of those categories, much less for every code path which eventually results in a filesystem call. I've run into cases where something mostly worked until you hit their integrated ZIP, Git/SVN, etc. support and found a new way that a filename was constructed.
My point wasn't that everything is fine but simply that this is complicated and no decision results in avoiding problems. Not normalizing allows for confusing visually-identical files; normalizing results in errors or data loss which will be blamed on the OS.
Looks to me like the article made claims and backed them up with examples and screenshots.
Rather than saying the article is wrong, can you demonstrate /why/ it is wrong? Using its examples and concerns (Finder, console, and scripts)?
I think you might to re-read my entire comment: note that I'm not arguing that the technical details are wrong, only that they're insufficient to support the huge “APFS is unusable” conclusion.
As previously noted, Windows and Linux work the same way and they are used by more people in individual non-English locales than the total number of Mac users. Would you say “NTFS is unusable by non-English users” is a useful statement?
There's plenty of room to say that a particular tool needs improvement, or that people making systems which copy or archive files should check for pathological cases, but it doesn't help anything to overstate the case so broadly.
It's not a problem on Windows or Linux filesystems because Windows and Linux don't provide a half-assed normalization scheme that lets me fairly easily create files that can't be accessed. If the Cocoa libraries did no normalization, then the resulting behavior might be obnoxious from a human-interface perspective, but I don't think the article would describe it as "little short of catastrophic".
I'm sitting here on my US English keyboard typing scancodes that look just like they did in 1990, so I'm not the best authority on how big of a problem it really is, but I'd guess it's going to result in a lot of bugs. Anyone who's ever tried to use a Mac with a case-sensitive HFS+ partition should be able to tell you that programmers can't even "normalize" their filenames consistently strictly within their native language.
This is only true if you're talking about the kernel APIs. Unfortunately, filenames come from a variety of sources and it's easy to find tools which inconsistently normalize them – e.g. simply copying and pasting a name from a Word doc, web page, etc. which has different normalization than whatever originally created the file – or which produce either duplicate error messages or confusing error messages because the normalization form used in a file doesn't match the normalization form written on disk.
I've encountered variations of this problem on all three systems. No approach is going to handle 100% of the filenames in the wild and all of them will require extra care in the user-interface which may or may not have been done – e.g. the Windows Explorer still provides no way to tell why Café.txt and Café.txt are not the same file – and fixing the cases where programs are internally inconsistent. APFS switching will expose some programs which were unsafe before but since it's consistent with the other common filesystems it'll remove the need for every archive, version control, etc. system to either special-case or break.
The reason it's not that expensive is: a) this requires no memory allocation, and b) most characters in most strings require no normalization!
Notionally you just look at pairs of next codepoints, and if the second one isn't combining and the first one is canonical for the chosen NF (a very fast check for ASCII!), then there's no need to normalize the first, otherwise you gather the combining codepoints and normalize, producing one normalized character and restarting the process where you left off. Most of the time the first codepoint requires no normalization, so the fast path is fast -- not as fast as a normal strcmp() or memcmp(), but still pretty fast.
Why? It's done once per open, not once per I/O.
In ZFS it's once per-open()/stat()/and so on. But still, not at all on readdir(), and anyways, it's highly optimized. For an all ASCII filename the slow path is never taken, and for a mostly ASCII filename the slow path is only taken for non-ASCII codepoints that are followed by combining codepoints (that check is itself a slower-than-the-fast path, but still faster than the slowest path).
Basically, it's noise.
Specifically, ZFS has a normalization-preserving, normalization-insensitive behavior -- a lot like case-preserving but case-insensitive behavior, but for normalization forms rather than case.
The way this works is that there's a) a string comparison function that can provide normalization- and/or case-insensitive comparisons, b) a character-at-a-time normalization function of sorts used for directory hashing (since ZFS hashes directories).
This is much better than normalizing on create, which is destructive and obnoxious when the form you normalize to is not the more common form and you live in a sea of applications that don't do normalization.
That's not a question of being true or not but referring to different things. The distinction I was trying to make is that HFS+ will force every filename into NFD. With ZFS, the filename as received from the APIs should be the same byte sequence which was used to create it and that avoids an entire category of bugs where e.g. a program writes a file and fails to find it in a later readdir() call. As an example, Subversion and Git both had numerous bug reports over the years where it was impossible to simply checkout a repo containing a file which used a different normalization form.
Say you create a file with a name that has different NFC and NFD forms, and you create it with the NFC form. Then you go try to create it with the NFD form, well, if doing an exclusive create (O_EXCL) then you'll get EEXIST, else you'll open the existing file.
I threw up a little.
Of course, it didn't take long for that to change, but MS had already put a lot of resources into switching to UCS-2.
Well-meaning developers learn how to support Unicode with wchar_t only to discover that wchar_t is the worst way to support Unicode.
https://www.moria.us/articles/wchar-is-a-historical-accident...
It was a very important article for 2003, but it should be honorably retired.
Its plan of action isn't as good as the modern "UTF-8 Everywhere", and even its motivating examples are becoming less relevant: when was the last time you went to the 'Encoding' menu in your web browser and guessed which codepage the page author meant to use? When was the last time you worried about whether your e-mail was 7-bit clean?
[1] https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...
When I argued strenuously for n-p/n-i behavior in ZFS (see elsewhere in this thread) I was told that gee, Unicode doesn't describe n-i string comparison the way I was proposing, so we can't do it :( However, fast n-i does follow from the spec, so eventually it got done. I can't take credit for the code, though I did contribute an optimization.
This is not very good. You want normalization-preserving/insensitive behavior instead.
It's no more a "bag of code units" than Linux filesystems store a "bag of code units". Windows will barf back whatever wchar_t array you give it, just like Linux will barf back whatever char array you give it.
So I can take the byte string [0xFF, 0xFF, 0xFF] and use that as a filename on Linux, or I can take the sequence of 16-bit values [0xD800, 0xD800, 0xD800] and use that as a filename on NTFS. They're not made of code units because they're not Unicode strings.
Sequence of bytes. "spiroagnew.txt" and "xtt.growapenis" are not the same filename.
You need to continue to write this on APFS. Nothing really changed in that regard. It's now just also possible to store denormalized filenames.
Short-term this will cause a lot of inconsistencies with applications using low-level APIs as the files currently existing will be NFD normalised, but any user-given path will very likely be more or less equivalent to what NFC would do.
So in the short term, this will be a mess (unless the APFS conversion on install-time also does a once-time conversion of the normal form), but in the long term, I believe this is the way to go.
And for that matter, the same should happen with regards to case-insensitivity which is even worse as ensuring case-insensitivity might actually be dependent on locale, which means that depending on the user's locale two file names might or might not be identical under case-insensitivity rules.
Unfortunately, it look like there's just too much legacy code around that plain doesn't work with a case-sensitive file-system on the mac (I'm looking at you, Adobe).
BTW: On my 10.13 Beta 1 setup with an APFS converted boot drive, unless I manually create a NFD encoded file name on the command line (which you can't do accidentally), everything is NFC both in the UI and on the command line. This also applies to files I haven't touched since the conversion to APFS.
I've seen something in a couple comments now to this effect, does nobody here do anything with ZFS at all? It's a pretty great filesystem on not just illumos but FreeBSD, Linux, and macOS, and it has full native normalization support (and for all 4 forms too). Normalizing is optional, but it's definitely there and with Macs I have been using NFD for compatibility purposes for 5 years now with no discernible performance problems in day-to-day use. I know ZFS isn't remotely a majority but I don't think it's that obscure either.
In particular ZFS has normalization-preserving/normalization-insensitive behavior, which is far superior to HFS+'s opinionated normalization-on-create (to a form that is different from the common input modes' output!).
The choice of APFS is that a filename is a sequence of bytes. Nothing more, nothing less (feel free to correct me if I'm wrong here).
If you want to see the kind of issues that path normalisation brings, check out this: https://github.com/thibaudgg/rb-fsevent/blob/master/ext/fsev...
I'd like to believe that most developer would prefer the current behaviour over that proposed by the article (normalisation).
If there are issues in Apple's Finder or other high level APIs, I'd like to believe they will be fixed before the final release. These are fixable bugs, IMHO.
I've always disliked the practice of messing around with file paths (storing them, concatenating them, etc). I preferred the way that the Classic MacOS typically dealt with filesystem references and aliases instead of file paths. This also meant users can rename or move files and references still work.
You got a similar siltation with NTFS where the file system supports long paths while the userland WinAPI does use a much smaller length limit (unless you jump through "\\?\ hoops, which does not work for relative paths), which renders certain files "unopenable" by certain applications.
[0] Regarding opening the wrong files, anybody remember the Android zip vulnerabilities? https://googlesystem.blogspot.com/2013/07/the-8219321-androi... It's not that hard to imagine that some macOS software does security/sanity checks on files using the low-level non-normalized API but then opens the (wrong and unchecked) file later with normalizing high level API, or vice versa. Having this API discrepancy built into your OS certainly makes these kinds of things more likely.
Sadly it's not.
> The choice of APFS is that a filename is a sequence of bytes. Nothing more, nothing less (feel free to correct me if I'm wrong here).
That is incorrec. APFS treats filenames as utf-8 strings and depending on the API you are using normalization is still taking place but on different levels. For instance all Cocoa APIs will perform normalization but the underlying syscalls and posix APIs will not.
APFS should have been normalization-preserving/normalization-insensitive.
"The case-insensitive variant of APFS is normalization-preserving, but not normalization-sensitive."
It should just always be form-preserving and form-INsensitive.
There is absolutely no reason to want the combination they give. It makes me suspect that they didn't bother separating case- and form-insensitivity, although I can't imagine why one wouldn't. It seems like.. a mistake. Someone misunderstood.
Previously normalization (while terrible) was at least consistent. Now different APIs have widely different behaviour and it's even worse on the application level. What used to be a annoyance has now become absolute chaos with no guidance from apple.
My understanding of the article is effectively: AFPS on macOS in the current form is unusable for non-English users and I strongly agree with that. While it might be partially usable, in particular if you upgrade from an earlier mac was performed, it has a ton of really terrible edge cases most users will quickly run into.
I don't think that is true. https://developer.apple.com/library/content/documentation/Fi...:
"Filenames in APFS are encoded in UTF-8 and aren’t normalized."
So, names must be UTF-8; not every sequence of bytes is valid UTF-8. What they removed from the file system are the Unicode normalization tables.
The FS needs to be normalization-preserving/insensitive. http://cryptonector.com/2010/04/on-unicode-normalization-or-...
The problem is that most input methods produce something close to NFC while HFS+ decomposes to something close to NFD. Which means that if you cut-n-paste non-ASCII Unicode names from a finder into any app that doesn't normalize, then you'll have problems.
The solution we came up with for ZFS was normalization-preserving but normalization-insensitive name comparison and directory hashing. This produces the best interoperability via NFS, SMB, WebDAV, and so on, and local POSIX access.
EDIT: i.e., ZFS doesn't know or care about encodings. Its n-p/n-i behavior means that for non-UTF-8 that happens to appear as valid UTF-8 there is some potential aliasing behavior going on, but this is exceedingly unlikely, and we decided that it was worth taking that risk.
If you're ever going to do some sort of "displaying" of data, you cannot store it as bytes. You need to know what characters things are supposed to be presented as.
You could imagine not settling on a specific encoding, but you must know the encoding. Unless your plan is to show a list of numbers to users.
The filesystem is much better off handling filenames as a number of arbitrary bytes. Let people who want to put weird bytes in their filenames see ugly filenames along the lines of "\x00 Can you see this?"
Can they not read the filenames anymore because their shell/window manager is different?
Or worse, you speak multiple languages, or learn new ones, and need to... switch codesets? How do you then access your old files?!
We must switch to Unicode. Full stop. If there are imperfections in Unicode script support, then we must fix those, but otherwise we must adopt Unicode.
In fact, I think the normalization HFS+ does is more problematic. For example, fish shell can't complete file names when you use un-normalized characters in the input. [0][1]
Edit:
0: "Unicode normalization issues with HFS+" https://github.com/fish-shell/fish-shell/issues/474
1: "Completion does not work for special characters" https://github.com/fish-shell/fish-shell/issues/1794
Anyway, it looks like the issues raised in this (April) blog post will not actually apply to iOS 11 or macOS High Sierra.
I've been saying for years, to anyone who will listen (and many who won't!) that normalization-preserving/normalization-insensitive is the only sensible behavior for a filesystem.
(I've said this many times in the IETF in the context of NFSv4 and WebDAV and such. Every time I've noticed the subject of stringprep for filesystem protocols come up.)
http://cryptonector.com/2010/04/on-unicode-normalization-or-...
What I mean is that unicode normalization is really hard, and it should be it's own module that can be used regardless of the fs.
app->fopen->unicode normalization->APFS/HFS/FAT...
app -> CFFile -> unicode normalization -> fopen -> APFS
which screws you because anyone can just call fopen on their own without using the core foundation libraries, leading to inconsistent states in the filesystem. You can be higher level than the filesystem, but only a little bit. You can't be higher level than some API that developers will regularly use (unless you do like ZFS and normalize at lookup rather than create).> APFS now supports an on-disk format change to allow for normalization-insensitive Case Sensitive volumes. This means that file names in either Unicode NFC or NFD will point to the same files.
Which means that both versions support normalisation-insensitivity.
(Edit: There is also a one-line mention of this in the iOS 11 document [1] but it doesn't say if it is the default.)
[0]: https://developer.apple.com/library/content/releasenotes/Mac...
[1]: https://developer.apple.com/library/content/releasenotes/Mac...
APFS is mentioned at the very end.
I would say the only downside to this approach, is a user wouldn't be able to distinguish two files with the same name apart, but it's hard to imagine how they'd get to creating such a situation in the first place without the developer rule above being violated.
Absolutely, yes. File names are text by their very definition; that we've been treating them as "bags of bytes" is a historical tragedy. At the very least, file names need to be displayed, as text, to the user, so they should be stored as text, that is in some well-defined encoding, and yes, it should be the job of the filesystem driver / kernel to enforce that it's not writing garbage out to disk.
Reinventing that wheel in every system that in any way interacts with the filesystem is bad engineering, and doomed to fail.
Further, I don't see why the typical user should need to know or understand the differences between 'e\N{COMBINING ACUTE ACCENT}' and '\N{LATIN LOWERCASE E WITH ACUTE ACCENT}'. Likewise, I don't see why each and every piece of code should be forced to handle that. Developers will get this wrong. In fact, the article seems to say even Apple can't get it right, in that Finder will not correctly show the directory contents in some instances, and fails to open files in some instances, telling users the file "doesn't exist".
The point is that it doesn't need to be. I would entertain that not everyone might not want to use Unicode: in that case, the FS should still have a well-defined encoding, such that I can still arrive at a string to display to the user. The point is not that "Unicode is best" but that storing file names as "bags of bytes" is incredibly user unfriendly, and there needs to be a straightforward, no bullshit method to display and transmit the names of files.
But I would also argue that Unicode is the best we've got presently, and it would be pragmatic for a filesystem to simply adopt it outright. It's overwhelmingly the dominant character set in use today, especially if you ignore deprecated junk that Unicode is a strict superset of.
Filesystems should do the same. Either pick one way and stick to it (expose it as UTF8, with a sensible normalization and collation) or offer several options that can be specified when creating a volume and do on the fly conversions for clients that need it.
Anyways, it's possible to allow UTF-8 and non-UTF-8 on the same filesystem, and still provide form-preserving/insensitive behavior... ZFS does it.
In this case: yes. Specifically the FS should implement normalization-preserving/insensitive behavior. http://cryptonector.com/2010/04/on-unicode-normalization-or-...
This is actually an oversimplification. Mac NFD does not decompose characters in a few specific ranges[0]:
> U+2000 through U+2FFF, U+F900 through U+FAFF, and U+2F800 through U+2FAFF are not decomposed
But I do expect applications built for HFS+' pseudo-NFD to struggle with non-normalizing APFS.
[0]: https://developer.apple.com/library/content/qa/qa1173/_index...
Other systems consider filename a string of octets and interpretation what these octets mean is left for userspace to decide.
For Linus Torvalds colorful opinions on HFS+, see here: https://plus.google.com/+JunioCHamano/posts/1Bpaj3e3Rru (in the comments).
NFD over NFC is simply trading data over CPU. NFC requires 3 passes over a string, NFD only 2, whilst NFD is usually a few bytes longer. Linus talking about NFD corrupting data is of course nonsense. He should stay technical and only rant about things he has an idea about. The python3 NFKC normalization format is nonsense, but not really important. NFD is fine, because faster.
Shifting interpretation of byte encodings from utf-8 to user-space is typical Linux non-sense, but he inherited the existing mess. Using utf-8 is miles better than bytes.
Case-insensitivity in HFS+ is of course legacy nonsense. This should have been to target to get rid of, not normalization.
The 2nd big security issue would be to forbid mixed scripts in a name (filename). I blogged about it here, and OP uses the same examples I used: http://perl11.org/blog/unicode-identifiers.html Invisible combining marks, right-to-left overwrites or similar / spoofs security problems are only fixable with normalization and more TR31 (mixed scripts, confusables), though confusables can only be handled in user-space, not the FS.
ZFS does normalization on LOOKUP. This is much better, as it preserves whatever form you used, but is normalization-form insensitive, which is precisely what users need (and would want, if only they could be expected to understand what the heck is going on!).
Of course this is all going to break now, of course it will take years until Oracle fixes this.
rsync -rltv --iconv=utf-8-mac,utf-8 from_a to_b
and similar sshfs: https://github.com/osxfuse/sshfs/issues/14#issuecomment-1859...I need to use that iconv options everytime when dealing with umlauts and mac<->linux communication. So do I still need that on APFS or not?
> BTW: On my 10.13 Beta 1 setup with an APFS converted boot drive, unless I manually create a NFD encoded file name on the command line (which you can't do accidentally), everything is NFC both in the UI and on the command line. This also applies to files I haven't touched since the conversion to APFS.
Effectively you should most likely still do that, yes. Apple's apps more or less expect normalization.
The fact that Cocoa apps create different filenames than the terminal is not great. It's even worse that some UI (like Finder) seem to cause even more confusion is not helping.
ZFS allows you to store non-UTF-8 if you like, but for all valid UTF-8 it implements normalization-preserving/insensitive behavior, which is the best possible compromise (IMO).
That's not really the issue here. Even if you can avoid normalization in some languages the problem will come back in others. For instance what are you going to do about invisible characters, control characters or more? Traditionally software attempted to just ignore the problems and let it blow up in other places. For some fun issues just try whitespace and bell characters in filenames and navigate around in your shell.
Normalization is just one of many issues with filenames.
Incidentally, ASCII was actually a multi-byte codeset... since one could combine most lower-case characters with BS (backspace) and overstrike with apostrophe, backtick, tilde, comma (for cedilles), and double-quotes (for umlauts), or with the same char for bold, or with underscore for underline. nroff(1) still uses this for bold and underscore, no?
I guess it's another relic of the idea that computer output goes to a printer rather than a display?
Note that BS only produces overstrike when followed by certain characters (which we might term "combining" for the fun of it), while most will just change the character at that location. A tty spinner is just |BS/BS-BS\BS|BS... with some delay between each -- no overstrike there.
https://en.wikipedia.org/wiki/Latin-1_Supplement_(Unicode_bl...
Some Cyrillic and Greek letters even look like Latin letters. And in some fonts '1' (one) and 'l' (ell) look the same even in Latin.
Dealing with these is much harder than with normalization forms. And confusables are a source of serious security headaches.
> Any filename (directory entry as identifier) must forbid mixed scripts.
So I'm not allowed to name my files “A=πr²” or “10kΩ Резисторы”?“10kΩ Резисторы” is Cyrillic, Latin, and Greek (U+2126 canonically maps to U+03A9). A proposal that disallows that but allows “10kΩ Resistors” is a political non-starter.
But why oh why did they have to specify FOUR normalization forms?
Also, for Hangul, though decomposition is the better form to use normally (since it's phonetic, not syllabic), conversion to pre-composed syllabic form is very useful sometimes.
Even without pre-composition there would have been multiple equivalent forms of writing any character that requires more than one combining codepoint.
The only way to have avoided normalization altogether would have been to have no combining codepoints, and only a complete set of pre-compositions. This would have been less flexible, and very obnoxious for Hangul and possibly others.
I believe the need for normalization was simply unavoidable. It's not the fault of Unicode but the fault of humans' script designs. And it's OK.
Ignoring compatibility normalizations there are really only two and neither of those two you can remove. So that's a bit of an odd question. If you want to remove the compatibility forms that's fair but nobody is force to use them. You can consider them "partially normalized" or "not normalized at all" for all intents and purposes
But that's not really what is happing here. What is happening is that the renderer decides to merge some code points into one glyph. That was true long before unicode as well (for instance ligatures) or even on old terminals with overstriking. Unicode just made this more explicit.
So, we need composed characters, but surely combining characters aren't needed, then? Well, no, because it's unreasonable to include a codepoint for every single combination, especially once you have multiple diacritics on a single character, for example. Also, compatibility factors in here too, because legacy encodings likewise have combining characters.
Also, it can be useful to decompose in some cases. For example, it's useful for removing diacritical marks, which can be useful for fuzzy searches. Hangul is really phonetic, not syllabic, so there it's the reverse: decomposed forms are most useful, except when you need to think in syllables.
https://nodejs.org/en/docs/guides/working-with-different-fil...
The gist is that normalization should only ever be used for comparison (if needed, e.g. "do these two files have filenames that would look the same to user"), and never for changing data (filenames are user data and should be stored verbatim without normalization). HFS+ should never have used normalization in the first place. You can think of normalization as essentially a lossy hash function (you cannot get back to HFS+ NFD form once you normalize it to NFC - HFS+ NFD and NFD proper are not the same thing - many people don't realize this). Using normalization for anything other than temporary comparison leads to data loss.
Almost every script but Latin-1 has TR31, TR36 and TR39 issues. http://www.unicode.org/reports/tr39/
(Incidentally, NFC is closed to new compositions, though new compositions can be added to Unicode.)
> How does Apple File System handle filenames?
> APFS has case-sensitive and case-insensitive variants. The case-insensitive variant of APFS is normalization-preserving, but not normalization-sensitive. The case-sensitive variant of APFS is both normalization-preserving and normalization-sensitive. Filenames in APFS are encoded in UTF-8 and aren’t normalized. […]
> The first developer preview of APFS, made available in macOS Sierra in June 2016, offered only the case-sensitive variant.
(I understand not treating confusable characters as the same. That's a different story.)
For me it seems more like a problem with Unicode. I can see why it is the way it is from a certain perspective. Very connivent.
But it has broken the underlying abstraction layer.
Keeping your issue reports neutral and avoiding hyperbole is an underrated skill.
Apparently that was a wrong assumption.