A brief history of APFS (2022)
eclecticlight.co
eclecticlight.co
As a result, this means that… if you create a Korean-named file yourself in Finder: then the file name is normalized in Unicode NFD (decomposed) form, but if you create a Korean-named file in a zsh shell, the file name is normalized in Unicode NFC (composed) form.
You don’t realize the difference until you try an `ls | xxd`, but the bytes are different, even though they look the same.
It’s an invisible mess that people don’t realize until they do. And a lot of programs gets subtly crazy when fopen(A/B) succeeds but listdir(A) doesn’t list B.
Be faithful and preserve what was entered rather than trying to "second-guess" the user.
UNICODE details like normalization can be handled up in the application layer.
A filename is an identity and if identities are ambiguous, you run into trouble, examples in this thread abound.
Putting it another way, do you work with relational databases? Primary keys? Those are identities. I don't even want to think about a world in which I can't roundtrip data through a database and have things NOT come out the way I put them in. "Human labels" or not ;-)
Of course it's fine to search for things case-insensitively or otherwise normalized. And so we have "find -iname" and "locate -i" etc. etc. That's totally reasonable.
But that worked in an OS that had no command line - UNIX users are allergic to binary configuration formats and demand to be able to type path names.
The fact that inhuman labels can sneak in stops me
> said no Linux user ever
You just weren't listening
> you run into trouble, examples in this thread abound
Bag of bytes approach is not trouble free
> Primary keys?
Do those accept any bytes in a bag?
> can't roundtrip data through a database and have things NOT come out the way I put them in
So you just want normalisation-preserving filesystem?
But I see that you care so little about human users that you think they should learn to avoid Á outside of letters to grandma...
Stops you from what? I'm not sure what you mean there. No one mentioned inhuman. Someone, not me, instantiated a "labels for humans" concept and I roll with that, for the sake of conversation.
>> Primary keys? > Do those accept any bytes in a bag?
Well, a "bag" (unordered) would not be the right term, but we understand what the original author meant (a bytestring), so: Yes, in fact they do. For instance, PostgreSQL:
create table bla (joyfulkey bytea PRIMARY KEY)
We don't want identifiers mangled. Should you be so unfortunate to have to use bytestrings as primary keys, then they'll come out the way you put them in. Because they're unambiguous identifiers. Not prose.
> So you just want normalisation-preserving filesystem?
I want an identity-preserving filesystem. And I have it!
> But I see that you care so little about human users that you think they should learn to avoid Á outside of letters to grandma...
I care a lot about human users. I give them nice file pickers so that they can easily choose to open files named ÁÁÁ.docx. Or search for those, using a system where both query and index is normalized (as is the common approach). And the other kind of users, also human, who find they care about the identity of files since they're writing that search engine and have to pass them in an argument vector to get the PDF viewer to open the file that the user found, will be most happy if they don't need to worry about the file system second-guessing them on identities.
PS There's a lot of you-this you-that ad-hominem things in your comments. That's not really the best way to hold a discussion.
Historically, you configured your OS to match (what you thought was) the encoding used on your disk, but that starts to break when you have network disks or when users exchange floppies.
In practice, I think about every Unix/Linux file system at least assumes ASCII encoding for the printable ASCII range. How else does, for example, your shell find “ls” on your disk when you type those characters?
So, not even in Unix/Linux is a file name “a bag of bytes”
I think file systems should enforce an encoding, and provide a way for users of the file system to find out what it is.
Apple’s HFS, for example, stored explicit ‘code page’ ID so that the Finder knew, for example, to use the Apple Cyrillic encoding to interpret the bytes describing file names on disk for that disk.
The modern way is to just say file names are valid Unicode strings, with UTF-8 being the popular choice. Unfortunately that introduces the normalization problem.
As you say, it can just show the sequence of bytes (plus some visual hints that they are escaped). It is not like nonsensical filenames with seem composed of random numbers are uncommon (e.g. look into your .git dir)
> How else does, for example, your shell find “ls” on your disk when you type those characters?
This is purely a presentation layer issue. The filesystem assumes nothing.
What you're suggesting would mean that a Korean user saving a file on their desktop and naming it in Korean would in turn see the filename turn into garbage, which, I hope we can all agree, is not a reasonable suggestion.
Creating 2 different systems carries risks: it has a learning curve, a support cost, and breaks everything that came before.
Perhaps, in a utopian green field, we redo a systems language, a portable systems API, and a common operating system with a compatibility layer to the old way on top.
Not changing impacts many more users (all the future billions) and costs even more: the learning/support/etc. is all there, and it's harder to learn about a buggy system since you have more papercuts
Also, it doesn't break everything, we're not living in a dystopia
2) Change what to what? There are multiple filesystems systems (and OSs) with different and sometimes incompatible filename encodings. Who picks the winner?
3) There is irony on dropping backward compat and requiring UTF-8, an econding whose claim to fame is backward compatibility with ASCII based systems (and not coincidentally was designed by UNIX Elder Ones)
1) so? That will be true for anything
2) sequence of bytes to human encoding. Who picked the loser??? This would similarly be some unidentifiable group of people. You can start by picking the winner for yourself, even at a conceptual level
3) no irony, sequence of bytes is not ASCII
The filesystem shouldn't arbitrate localization issues. It's the wrong place to do it in terms of performance, area-of-responsibility, and portability.
NUL-terminated UTF-8 in memory and size+data on disk is the best encoding.
Normalization is a multifaceted problem that needs to be handled at or before where the filename argument hits the standard library. Sometimes, normalization is NOT desirable and it is preferrable to normalize for comparison only to other strings and to faithfully store whatever was provided verbatim.
Side-effects and doing 737 MAX MCAS underneath users is the road to ruin.
Command line argument is treated as (unicode) string, and file name are (sometimes) treated as bag of byte .....
Czech language has Á, Č, Ď, Ě, É, Í, Ň, Ó, Ř, Š, Ť, Ů, Ú, Ý
Á can be represented in Unicode by its spesific codepoint, or can be made up of two - Letter A and the squiggly bit.
So when a user tries to fOpen a file called ÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁÁ
Should by application make 2^32 requests to the file systen to open a file for every possible combination of bytes this steing could be represented as?
Still not great.
I'm aware of the normalization problem, but that should be handled in the API layers above the filesystem. In which layer exactly is up for discussion though (but IMHO there should always be a fallback to access files via their "raw bytes path identifiers", if for nothing else than debugging).
"But that's subjecting the user to limitations of the system!" you might object, and while you'd be right we should also ponder whether a filesystem is a word processor or a system for locating data using filenames as identifiers, and consequently, what sound handling of identifiers is.
And let's be realistic. Most users will choose the file to open from a file picker dialog. fopen() is not their interface, their interface is a file picker.
The intersection of people for whom fopen() is their interface (no file pickers) but for whom learning lessons on the value of using unambiguous identifiers is somehow out of reach, would be exceedingly small.
The unambiguously encoded filenames wouldn't be in the natural language of the user, so you're asking the user to learn another language. Users shouldn't have to learn another language to use the computer.
For prose of course it wouldn't be acceptable if you can't use your full language, but this is a filesystem, and we're talking identifiers, not prose. We accept that for a wide range of everyday things — I don't think many people are getting worked up about web forms not accepting their phone numbers when entered in Latin numerals.
Using bytestring labels sans normalization is just the rational choice, and actually quite accommodating. People _can_ use their natural language. If you can encode your prose in bytes, then there you go, that'll be your filename ;-) and if you can't remember what diacritic-combining strategy you used when you created the file, and can't manage to fopen() it anymore by typing out the filename, use a file picker or search. And then, perhaps, go "that was a hassle, I'll avoid those fu͝n͜n͏y͝ sq̧u͜igg̶l̢įȩş for my next filename, computer are stupid" - not the end of the world, not the worst outcome.
Then there is Vietnamese which has a lot of accents and other diacritics. Not sure how much you can do without them.
Perhaps a language like Chinese doesn't have so many ambiguities, as in there is only one way to encode a given character in Unicode?
What is the threat model that makes having both versions exploitable? Historically, many vulnerabilities (e.g. several git vulns) have come from filesystems doing normalization rather than the lack of it.
Buy generally anything that leads to unexpected behaviour can be exploited in some way, if not technical then to mislead the user.
The question is, why is 'failing that' happening here? And if there's 2 files encoded differently with the same name, how would the user differentiate them?
What I needed to eventually do is to go on Windows, install WSL there, then run some script that recursively renamed them from the first to the other (on the FAT volume), and only then I could copy them from FAT to APFS, on the Mac.
(I could probably skip the WSL and do it with some Windows-only tooling, but that would take me even longer.)
so this little exercise involved mac, windows, linux (and plan 9, if you count the wsl filesystem shenanigans)
Well, it turns out that Unicode can represent a lot of characters by way of different possible encodings (for example é could be a single codepoint, or the codepoint of e followed by the codepoint of an acute accent).
Basically I had a Git repository and didn’t know about `git config core.precomposeUnicode`. And when the repository is synchronized to a Linux system, on the Linux side it can sometimes have 2 files with the same-looking file name but different normalizations. (Because I think ext4 doesn’t normalize Unicode?) That took me about an afternoon to fix.
https://en.wikipedia.org/wiki/Code_page
Yeah, native Linux filesystems don't really know anything about unicode at all at their heart. They really work more on raw bytes, simply reserving 0x00 (NUL) and 0x2F ('/'). Anything else goes in a filename as far as the kernel is concerned (including evil stuff like incomplete multibyte sequences, or other invalid UTF-8 byte sequences). User space is welcome to and encouraged to treat filenames as UTF-8 on modern systems, but that's not really enforced anywhere strongly.
But a tweet yesterday said they didn’t do a mock upgrade. Every single iOS update for a year did that so they could catch and fix a ton of bugs before the big switch.
Pretty genius.
I was reading about macos' history with filesystems a while back and apparently they were considering switching to zfs at some point but instead decided to design apfs as their own in house next gen filesystem. In my experience with it it seems to lack almost all of the features that would have made zfs interesting. One of the most obvious usecases would be for something like time machine. I have a parallels virtual machine that is hundreds of gigabytes in size and as far as I can tell that each time you modify even a single byte in that vm it will need to backup the entire file all over again. not only is this infeasible in the amount of space it would require it's also infeasible in that every time I do a backup i would need to spend the time transfering the whole thing all over again. this is also one of the biggest problems that next gen filesystems were designed to solve. is it really the case that apple's next generation filesystem doesn't support snapshots or block level replication when one of their most obvious usecases for it would be time machine backups?
https://arstechnica.com/gadgets/2020/01/linus-torvalds-zfs-s...
(We didn't talk about iPhone, so either it had not been announced yet, or development was still so closed down, that it never came up as a topic of consideration.)
My recollection was with their ZFS work, while it was clear to everybody that Time Machine should be a big benefactor of ZFS, they were still unhappy with the high RAM requirements (and also not thrilled about the high CPU requirements). Since laptops were then the bulk of their sales, and these machines shipped with much more constrained specs, there was a lot of uncertainty about when/if they were going to make ZFS the default file system.
(Looking it up, the base 2006 Macbook had 512 MB of RAM, and shared RAM with the Intel GMA 950 GPU.)
I recall there always was a secondary backdrop of concerns with the ZFS license and also the fate of Sun, but the engineers I spoke with weren't involved in those parts of the decision making process.
From one of the co-creators of ZFS:
> Apple can currently just take the ZFS CDDL code and incorporate it
> (like they did with DTrace), but it may be that they wanted a "private
> license" from Sun (with appropriate technical support and
> indemnification), and the two entities couldn't come to mutually
> agreeable terms.
I cannot disclose details, but that is the essence of it.
* https://web.archive.org/web/20121221111757/http://mail.opens...* https://arstechnica.com/gadgets/2009/10/apple-abandons-zfs-o...
What the lawyers say is different.
What the business folks who control the license say was very different. FUD and an unforeseen lawsuit hanging over your head at some random point in the future when Sun or Oracle needs a revenue lift at an end of a quarter is a recipe for ulcers.
Bonwick was the Sun Storage CTO and after the acquisition a vice president at Oracle.
https://www.theregister.com/2010/09/09/oracle_netapp_zfs_dis...
This is a likely an issue with Time Machine’s design rather than any limitation in APFS. Time Machine works with full files and uses hard links to link different versions. Unlike HFS+, APFS does support CoW (Copy on Write) and snapshots. You can verify this yourself by duplicating a file on your local drive and seeing that it doesn’t actually create one full copy of he file and occupy double the space.
(I know this isn't the response you were looking for. I've also experienced corrupt TM unrelated to HD encryption, not sure why)
The reports have continued for years, right into this one: https://forums.macrumors.com/threads/how-is-this-an-acceptab...
I recommend surveying the extant reports on the issue. I've read several analyses that assert that corruption is indeed inevitable, and my experience was consistent with that. But hey, roll the dice if you want. I'm just providing information.
They did change it recently. It uses snapshots now (I think from Big Sur on?)
APFS does have snapshots but I don't think it has replication.
Time Machine is so bad and has seen so little work, I feel like Apple really wants to drop it and move everyone to iCloud subscriptions instead. And yet they still don't have a cloud backup product, which is even stranger.
What I would love for APFS to have thought is block-level deduplication. Seems like an obvious fit for an SSD-optimized COW system anyway, surprised they haven't implemented it.
There's a way to explicitly compress files if you want. You can get the "afsctool" and run it on the directory with your sources for example and gain some space back. Unfortunately, this gets applied per-file only. There's no way to mark a directory to always compress everything in it. I don't think you can do it per-volume either.
It would be great if those options got exposed to the user. I'm not going to hold any breath though, it may be one of those "Apple knows better" issues.
Hard links are such a rare exception in any system that either ignoring them or compressing regardless would be fine.
Or check what others have done. You can do "btrfs property set /path compress zstd", so someone's solved it before.
Or add an config option
Why is that hard?
Under Linux, you can simply assign /dev/sdX to a particular drive, just create an udev rule. Under macOS, that's impossible to my knowledge.
Presumably they have sufficient control at that level for their current lineup, but I’m not sure what it means for aftermarket disks and external disks.
Just curious how folks apply this day to day outside of how Apple leverages it.
Oh well.
https://news.ycombinator.com/item?id=17852019
Apple's not big on f/oss. They maintain their hardware moat (historically) by differentiating with extremely proprietary software.
> Apple can currently just take the ZFS CDDL code and incorporate it
> (like they did with DTrace), but it may be that they wanted a "private
> license" from Sun (with appropriate technical support and
> indemnification), and the two entities couldn't come to mutually
> agreeable terms.
I cannot disclose details, but that is the essence of it.
* https://web.archive.org/web/20121221111757/http://mail.opens...* https://arstechnica.com/gadgets/2009/10/apple-abandons-zfs-o...
I don’t. They worked on ZFS, even released read only support, and some people think they came close to making such an announcement (e.g. http://dtrace.org/blogs/ahl/2016/06/15/apple_and_zfs/) but I don’t think they ever made any statement that ZFS would be the replacement for HFS+.
That's interesting to think about. Would have propped up the xServe idea better. A timeline where an Apple server was a somewhat credible competitor to Linux servers might have changed some things.