A ZFS developer’s analysis of Apple’s new APFS file system
arstechnica.com
arstechnica.com
The question I'm sure Apple engineers have were - How often do we see BitRot occurring off media, and is the media that we're deploying sufficiently resistant to bit-rot?
And, with APFS's flexible structure, this is a feature that can be added at a later time. Probably made sense to deliver in 2017 something that was rock solid that they could build on, than to either (A) push out the delivery date, or (B) not fully bake all features of the file system.
https://www.usenix.org/legacy/event/fast08/tech/full_papers/...
That's exactly what you want – a clean failure which prevents other software from silently propagating corrupt data. Even more importantly, with corruption it's usually much easier to recover from another copy if you notice the problem quickly.
Think about what happens with e.g. photos – you copy files from a phone, USB card, etc. and something gets corrupted on local storage. With integrity checks that means that the next time you load your photo app, go to upload photos somewhere, run a backup, etc. the filesystem you immediately get an unavoidable error telling you exactly what happens. If it's recent, you can go back to the source phone/camera and copy the file back.
With all current consumer filesystems, what you instead get are things like a file which looks okay at first – maybe because it's just displaying an embedded JPEG thumbnail, maybe because software like Photoshop is enormously fault-tolerant – but doesn't work when you use it in a different app or, because you didn't get an error, the underlying hardware file affecting your storage becomes worse over time until something fails so badly that it becomes obvious. By the time you notice the problem, the original file is long gone because you reused the device storage and you have to check every backup, online service, etc. to find which copies are corrupt.
The user would have to be informed, since Apples hardware has no storage redundancy. That will be perceived as an admission of failure on Apples part; your $2k device just murdered your data.
There is not really a UX flow apart from telling the user to recover the file from backup.
Technically it is the dominant strategy to have data checksums. - Even if you assume the Hardware is very awesomely perfect.
I agree with all your points, but just want to offer a possible UX solution:
Since the vast majority of data on people's hard drives are video and images, where minor data corruption results in (in most cases) just visual artifacts, we could have a pop-up dialogue that says: "A higher quality version of this file is found on your backup. Do you want to restore it?" when the corrupted file is a video or image and there's a confirmed backup of it.
If there's no confirmed backup then just silently ignore the corruption since the user won't notice anyways.
There absolutely is, since that could be performed automatically. Inform the user about the corruption, rename the corrupted file, then restore the most recent backup in its place.
In Apple's case, they have complete control over iCloud and could offer this easily for any file which is stored there. They could also add some sort of metadata API so services like CrashPlan, Backblaze, etc. could register the presence of other copies in a generic manner. Third party services could also integrate background scrubs into their existing application.
In each case, the first time that dialog appeared you'd likely have a customer for life from anyone who's gone through the hassle of losing a personal memory, important document, etc. or make a panicked search for other/older copies.
Also, I wouldn't be surprised if Apple adds a cloud-based backup for macOS once APFS is the default filesystem, since change sets would be extremely efficient.
I really like the hack of a block pointers being a data structure and containing the birth time (tx id I guess) and how that avoids needing managing bitmaps.
The advantage of doing send/receive vs rsync was also a nice explanation.
It is not often that you see a technology and think "Oh, this is great stuff", this is like that with ZFS. Haven't played much with it. But get that idea just from learning about it so far.
Both Jeff and Bill are great at communicating and explaining the technology. Like how they tag team, with minor funny bits here and there.
Regarding the main issue here, checksums -- yeah I don't see how Apple engineers could have watched this and said "Meh, don't need data checksums". Maybe they do have a secret vault with magic new holographic storage, immune to cosmic rays and other vagaries of physics, who knows.
It looks like he has a PDF up on his website:
I understand he's the APFS lead dev, he gave half the "Introducing Apple File System" session alongside Eric Tamura (the local FS manager)
"APFS, the Apple File System, was itself started in 2014 with Giampaolo as its lead engineer...(he built the file system in BeOS, unfairly relegated to obscurity when Apple opted to purchase NeXTSTEP instead)"
http://arstechnica.com/apple/2009/10/apple-abandons-zfs-on-m...
Edit: And this wasn't slow and of considerable overhead like what we have right now with indexers for OS-wide search.
The really cool part is that it was also portable – there were multiple email which seamlessly interoperated because they were both just querying BFS, which made it faster for developers to experiment with alternative UIs since you didn't have to reinvent or tune a lot of core functionality.
It's up to the people who cares to organize and assemble teams etc etc .. but only rarely it goes somewhere I believe. I still hope for a sane cpu, sane gpu, sane dsp, all very very open so that a solid software can be written without too much reversing and friction (I saw that tagged pointers architectures were back in research on risc-V so who knows).
It's time for an integrative phase after the last 20 years of growth.
I've never understood how people seem to ignore the fact that you can run more than one CPU intensive application at the same time or even have breathing room for light applications while most cores are busy heavy processing.
I can easily saturate all the cores you throw at me with the usual tasks I'm busy with at a computer, so give me more cores, more memory. And I run highly parallel tasks like image or video processing or compiling code.
If Apple can actually pull off this turnover so quickly, does that suggest the complaints about Apple's declining software engineering quality were overblown?
Edit: Ted T'so in this talk(1) (at the 8 min mark) discusses the taskforce that birthed Ext4 and Btrfs and its estimate (based on Sun's experience with ZFS, Digital's with AdvFS) of 5-7 years at a minimum for a new file system to be enterprise ready--an estimate which definitely proved optimistic with regard to Btrfs. Will APFS be different?
Snapshots are read-only, unlike Btrfs (but like Linux' LVM, which has been around for quite some time), and we won't even delve into online rebalancing and device addition.
APFS is a very welcome addition given that it brings Apple file systems into the XXIst century, but the reason for its speedy implementation is that they purposefully constrained its goals — a good call, in my opinion.
If you have the same workloads as Apple's customers and large amounts of SSD storage, you'll find Btrfs and ZFS both fit your needs on Linux.
Besides being old, there's nothing much to the "robustness of HFS+". It's not like users are losing data left and right, as its being painted. In fact it's an FS running just fine on about a billion devices...
https://blog.barthe.ph/2014/06/10/hfs-plus-bit-rot/
> A concrete example of Bit Rot
> ... HFS+ lost a total of 28 files over the course of 6 years.
BTW he didn't lose any data, since he had backups. If he had a checksummed filesystem, but not backups, he would still have lost data. Checksums, like RAID, aren't backups!
To quote:
>I understand the corruptions were caused by hardware issues. My complain is that the lack of checksums in HFS+ makes it a silent error when a corrupted file is accessed. This not an issue specific to HFS+. Most filesystems do not include checksums either. Sadly…
The complaints are usually about minor parts of the OSX/iOS stack -- parts Apple might not even particularly prioritize.
A faulty filesystem on the other hand is something entirely else altogether and something they can't ship unless it's good.
> an estimate which definitely proved optimistic with regard to Btrfs.
SUSE has had Btrfs support for enterprises since SLE11SP2[1,2] (2012). And it's been the default for the root partition since SLE12[3] (2014). So it wasn't overly optimistic at all, it was actually a very accurate estimate. The same support apples for openSUSE, but I think they had it for longer.
[1] https://www.linux.com/news/snapper-suses-ultimate-btrfs-snap... [2] https://www.suse.com/communities/blog/introduction-system-ro... [3] https://www.suse.com/documentation/sles-12/stor_admin/index....
If I look at btrfs patches on lkml, on one hand side I can see some fixes for data loss, but on the other they're usually close to "if you change the superblock while log is zeroed and new superblock is already committed and there's a power loss exactly at the point, you'll get corruption" - which are just really obscure edge cases people are unlikely to ever hit.
So what can I look at to get a realistic picture of what's going on? (what would SUSE point me at)
(for the negative results, I know of the recent filesystem fuzzing presentation where btrfs comes out worst, but honestly I don't consider it interesting for real world usage - car analogy, I'm interested how the car behaves on a typical road, not which fizzy drinks added to the gas tank will break it)
As for unstable and corrupts data, this is just not true. Many more users using it on stable hardware don't have problems. I've used Btrfs single, raid0, raid10 and raid1 for years, and haven't had corruptions at all ever that I didn't myself induce.
I have stumbled upon, just days ago, parity corruptions in raid5 however. The raid56 stuff is much much newer and hasn't been as well tested, and has been regarded as definitely not ready for prime time. So that's a bug, and it'll get fixed.
Bunches of problems happen on mdadm and LVM raid also due to drive SCT ERC being unset or unsupported, resulting in bad sector recovery times that are longer than the kernel will tolerate. That results in SATA link resets, rather than the fs being informed what sector is bad so it can recover from a copy and fix up the bad sector.
So there are bugs all over the place, it's just the way it goes and things are getting quite a bit better.
It is totally true that Apple can produce their own hardware that doesn't do things like lie to the fs about FUA or its equivalent of req_flush being complete when it's not, or otherwise violating the order of writes the fs is expecting in order to be crash tolerant. But we're kinda seeing Apple go back to the old Mac days where you bought only Apple memory and drives, 3rd party stuff just didn't happen then. The memory is now soldered on the board and it looks like the next generation of NVMe and storage technologies may be the same.
Windows and Linux will by necessity have file systems that are more general purpose than Apple's.
Hence it's not very surprising that the supported btrfs in SuSE is a sub-set of all the available features: the ones that are most mature.
This is an illustrative patch from a couple years ago, not sure if still current:
http://kernel.opensuse.org/cgit/kernel-source/plain/patches....
The type of developer you would have working on a filesystem are likely going to be from a different world than those who work on UI applications. Speaking about their apps, the quality problems Apple faces are usually design-based rather than functionality. When working on a filesystem, you're not going to be forced into developing a crappy application by some designer who ruins the entire application.
http://www.evanjones.ca/tcp-and-ethernet-checksums-fail.html
When The CRC and TCP Checksum Disagree: http://conferences.sigcomm.org/sigcomm/2000/conf/paper/sigco...
Alternately, just look at "netstat -s" for any machine on the Internet talking to a bunch of others. Here's the score for the main web host of the Internet Archive Wayback Machine:
3088864840 segments received
2401058 bad segments received.APFS isn't designed as a server file system. It's meant for laptops, desktop and most importantly (to Apple) mobile devices. Note that most of the devices are battery powered. That means "redundant" error checking by the FS is a meaningful waste.
That's not to say they might not add error checking capability in the future, but it makes total sense to prioritize other things when this file system is mostly going to be used on battery powered clients basically never on servers.
I mean, it'll tell you that your only copy of a file got corrupted, but it'll still be corrupted...
Never underestimate the value of a reason to play Quake 2.
At the very least, a checksum failure might tell them (or the tech they're consulting) that they have a data problem, rather than, say, an application compatibility problem.
Some filesystems can be configured to keep two or more copies of certain filesystem/directory/etc. contents. Two copies is enough information to do something useful.
How big an issue this is in practice I don't know.
Do it in CoreStorage, not filesystem X.
The cause is the change in editorial standards for the site which has been decreasing for years, and so every year is at a new all time low. Not to mention the really invasive ads and even sponsored content.
As for advertising, how do you expect them to hire writers and editors or run a high-traffic website without advertising? They've offered paid subscriptions for years and subscribers don't see ads at all but not enough people have taken them up on it.
The lack of subscriptions are probably a symptom of the editorial quality going down. You can say there's no evidence of it, but when lots of people complain about it that is evidence of it. I was a long time Ars reader and the b.s. finally got too thick that I have stopped going to the site regularly in the last year.
It used to be a site for intelligent, balanced articles about tech. Now it's got shills like DrPizza who basically just reprint whatever Microsoft's PR department emails him.
I stood around with Giampaolo, Tamura, and other Apple folks. I had a significant leg up on journalists; I know the subject well and the Apple folks know and respect me based on technology I've worked on. (And they didn't kick me out when I revealed that I didn't actually have a conference pass). I knew (some of) the questions to ask, and they were remarkably open with their answers.
? ZFS is copy-on-write.
> With APFS, if you copy a file within the same file system (or possibly the same container; more on this later), no data is actually duplicated. Instead, a constant amount of metadata is updated and the on-disk data is shared. Changes to either copy cause new space to be allocated (this is called "copy on write," or COW). btrfs also supports this and calls the feature "reflinks"—link by reference.
> I haven't seen this offered in other file systems (btrfs excepted), and it clearly makes for a good demo, but it got me wondering about the use case. Copying files between devices (e.g. to a USB stick for sharing) still takes time proportional to the amount of data copied of course. Why would I want to copy a file locally? The common case I could think of is the layman's version control: "thesis," "thesis-backup," "thesis-old," "thesis-saving because I'm making edits while drunk."
CoW is one of those features that is superficially questionable until you start noticing the bits and pieces of workflow it really makes faster and easier. Given the keynote was towards an audience of developers, I'm really surprised that there wasn't a demo showing how much faster deploys, etc are with such tech.
Could someone clear up how this is can be determined on the filesystem rather than scheduler level (I suspect it cannot be, or the article is making bogus claims)?
HFS+ implements scheduling of background QoS threads, especially on HDDs - you can see it working with 'spindump'.
I'm not sure I quite understand this approach. It's sounds like a pure NIH kind of method (which is admittedly common for Apple). I.e. of course one can always try to reinvent the wheel, but why is it bad first to analyze what already exists and to evaluate good / bad sides of that? Or his approach is simply always to make everything from scratch and not to look at anything else?
It's a legal defense strategy.
ZFS is covered by multiple patents.
If someone who have never read about any of ZFS's designs and patents independently reproduces one of ZFS's patented features, then the courts could rule that that was non-infringement.
That's why the author said of Giampaolo: "...but didn't delve too deeply for fear, he said, of tainting himself". Reading too much about a patented product effectively taints yourself from being able to freely create your own designs.
That being said IANAL and have no idea whether residual knowledge has been explored w.r.t. the GPL or CDDL.
Apple makes a mistake in 10.11.6 and ships a kernel without DTrace. Everything runs but a few nerds notice and file bug reports. 10.11.7 quietly ships and all is right with the world.
In contrast, shipping a new filesystem is enormously invasive: billions of hours are spent globally rewriting EVERY storage pool in existence. Some percentage of those will fail due to hardware problems or corruption which happened years ago (or even somewhere else) but was previously unnoticed, flooding support and the news with dire predictions. Every crash or performance issue noticed for the next year will probably be vocally blamed on the new filesystem, even if there are clear signs pointing elsewhere.
You don't want to deal with that any more than you have to.
The question isn't whether Apple will force this on users soon or without notice but rather the observation that they're going to be very careful not to do this more often than actually necessary. That's why despite having implemented and shipped ZFS in the 10.5 beta series they removed it prior to release. Sure, it might have gone okay but if that changed later there would be no easy way to go back without forcing users to migrate existing data and if that hit a licensing/patent case, the other side's lawyers enjoy the extra leverage which that would give them.
As an aside, “There are many filesystems in the kernel” is technically true but misleading: HFS+/HFSX is used constantly on every device, FAT/ExFAT are used regularly by many users, and everything else is a rounding error. A few Mac users use NFS or SMB, even fewer use UFS, etc. and no iOS device uses any of those.
And I'm not sure you need snapshots and many of those industrial features on a watch anyway. Sometimes simplicity is actually a plus. So who said they need "one shoe fits all" to begin with? It's never a good approach.
You don't necessarily need snapshots on a watch, but I can see it coming in useful on a phone when you're editing pictures and video.
Snapshots on a watch could be useful for system recovery. Say something went wrong during an update you could potentially use the snapshot to revert to a previous version.
On the other hand, computers are gradually moving towards miniaturization, so I suppose all this really is quite transitory, and soon enough all such considerations wills simply be irrelevant.
The main point is that upgrading "in-place" does not require any of the metadata to be in the same place between old and new filesystem. With a mildly flexible destination filesystem, the old filesystem can stay there until you're completely sure the conversion is a success. You end up with two read-only filesystems sharing a partition, and you choose which one to go forward with. Cancelling at any point is trivial, even if the conversion process crashes.
Much of the operating system is proprietary, but I didn't think the kernel or HFS+ filesystem were included in that.
Also, btrfs is GPLv2. So that agrees with my point -- if they didn't mind using free software they could use that code.
I'm less sure about iSCSI targets, but iSCSI is a block-storage protocol. If you can format an iSCSI volume with HFS+ today, you'll probably be able to format it with APFS tomorrow.
(and I would kind of avoid formatting external drives in "funny" FSs as I might want to read them in some other OSs. Unfortunately this usually means FAT32)
But I do have external media as HFS to work with Time Machine
Another option is ext3. The only caveat there is the Windows ext drivers don't fully support ext3 (unless I've missed an announcement). However they do fully support ext2 and ext3 is backwards compatible so you can effective get ext3 support in Windows.
Sadly though, all the good stuff requires 3rd party libraries. It's a real pity everyone can't agree on a standard to replace FAT32. :(
Did Microsoft ever open-source the NTFS specs and drivers?
I know there was a way to write to ntfs from Linux, but it required to install ntfs driver file from Windows.
I hope there is a native ntfs support that supports writing.
That isn't true. You need to use ntfs-3g, which is a free software implementation of NTFS (that allows both reading and writing). It's been stable for 10 years. Using NTFS doesn't require anything from windows and doesn't require proprietary software.
I stated how it was in the past[1]. Anyway according to [1], it looks like ntfs-3g still uses a proprietary version of ntfs.sys
[1] http://superuser.com/questions/139452/kernel-ntfs-driver-vs-...
Also, Trisquel (an FSF-approved GNU/Linux distribtion, meaning that it doesn't have any proprietary software within 100km of the distro) has packages for ntfs-3g[2]. So it's _definitely_ entirely free software.
So again, you're wrong on this point. In addition, I strongly believe that you were never correct on this point. Maybe you confused ntfs-3g with the proprietary version that company sells?
[1]: http://www.tuxera.com/community/open-source-ntfs-3g/ [2]: http://packages.trisquel.info/search?keywords=ntfs&searchon=...
Since my main usage for external media was big files (you know, the ones with extension mov, mpg, avi, etc) using FAT32 was not a big issue (unless it was bigger than 4GB of course)
I don't know the answers to your other questions. Sorry.
Here is a Linux script to format a drive as UDF so it is usable with both Linux and Windows: https://github.com/JElchison/format-udf/blob/master/format-u...
UDF works well for thumb drives that are large enough that you might want to copy a file to the thumb drive that exceeds the max file size of a FAT filesystem.
It's actually a shame. A better file system for removable devices was a huge opportunity for an open standard.
Ideally they could have released a filesystem that was backwards (read at least) compatible. Large files could show up as multipart files when viewed as FAT32.
Every OS in the universe supports UDF now, and its the only real ubiquitous non-proprietary filesystem. I use it on all external storage that I cannot guarantee will be touching Linux machines exclusively.
Apple contrasts [space sharing] with the static allocation of disk space to support multiple HFS+ instances, which seems both specious and an uncommon use case.
Really? I depend on this use case every year to safely test drive pre-release OS versions.
(So, Yosemite on one, and a pre-release of El Capitan on the other.)
Space sharing just promises to allow the same thing but without wasting space.
This seems too obvious and important a use case to call "specious and an uncommon".
What about linking to the actual source rather than a 3rd party: http://dtrace.org/blogs/ahl/2016/06/19/apfs-part1/
edit: and to be clear I'm delighted and flattered to see my work featured in Ars!
Link wherever you'd like, of course, but the more traffic this pulls in, the more ammo I have to be able to get Adam contributing to Ars as a regular freelancer!
(edit - hi, adam!)
(edit^2 - corrections corrected. Apologies for the errors. I am just a simple caveman. Your mathematics confuse and annoy me!)
For example, when I was a kid growing up in NY, one of the local radio stations I listened to was WPLJ, 95.5 MHz. That's 95,500,000 cycles per second, not 100,139,008.
Go back nearly 100 years, the Chicago area got a radio station called WLS[1], one of the original clear channel stations. It broadcasts at 870 KHz. That's 870,000 cycles per second, not 890,880.
Much as computer people would like "kilo", "mega", "giga" etc to mean 1024^(whatever), there's a lot of precedent for doing things the old fashioned way!
As Wikipedia explains: tera-, from Greek word "terastios"="huge, enormous", a prefix in the SI system of units denoting 10^12, or 1 000 000 000 000
SI is a well accepted standard. Just because it's more logical for chip designers to implement memory chips using powers of 1024 isn't a good enough reason to ignore SI.
That or rewriting it for the plural.
A snapshots lets you freeze the state of a file system
Doesn't really read well. A snapshots as a construction in English even in context really grated as I read it.- using DTrace for ZFS operations insight
- state of OpenZFS, comparing illumos, FreeBSD, NetBSD and Linux
- myths and/or often cited problems (ECC, COW unfit for DB load, VM images etc.)
- ZFS version and feature flags wrt portability between illumos and FreeBSD versions
- pool flexibility work that will make it easier to remove devices
- comparison to btrfs, hammer{1,2} and flash filesystems
- the topic of zfs pointer rewrite
- garbage collection and safely removing traces of a file/directory in a COW fs
- rebalancing story compared to HAMMER and future work in this space
- built-in ZFS encryption (independent of system crypto volume support)
I should say that I'd only support an article like that if Ars allows parts of the written text to be incorporated into OpenZFS wiki/documenation.
We did run a big piece by Jim Salter a couple of years ago on next-gen file systems that focused on ZFS and btrfs (http://arstechnica.com/information-technology/2014/01/bitrot...), but yeah, I'd love to have more filesystem-level stuff showing up. The response is generally very, very strong—turns out people really like reading about file systems when the authors know what they're talking about!
edit -
> I should say that I'd only support an article like that if Ars allows parts of the written text to be incorporated into OpenZFS wiki/documenation.
That's more complicated, unfortunately. I am not a lawyer etc etc and I am only speaking generally here, but Ars and CN own the copyright on the pieces we run (though syndications like Adam's piece today are different), and wholesale reuse of the text without remuneration isn't something that the CN rights management people like. Fair use is obviously fine, so quoting portions of pieces as sources in documentation is not a problem, but re-using most or all of something isn't (necessarily or usually) fair use.
(again, not a lawyer, my words aren't gospel, don't take my word for it, etc etc)
Also, Dr. McKusick documented some of the internals in his living FreeBSD kernel book, but as you said, this might be beyond Ars's scope.
Though, I've seen some deep technical content on Ars, so why not give it a try.