Apple’s big test of data integrity
eclecticlight.co
eclecticlight.co
This feels weird. So they have a very low number of possible systems to support. The OS is immutable. Why would they go with a reinstall at that point instead of "we know which part failed, so we'll re-download that one block from Apple service"?
(I guess that one ties in with: so you've got a low number of systems and an immutable OS - why does it take 15min to apply an update)
Well, it means they have to keep track of every build, and keep every version of every OS live, and that's not something they really do. Especially in cases of security fixes, they'll take down an old version from the update servers.
As for 15min, they're actually testing fast patching right now, although it's currently limited to security updates (rapid response).
That's not a lot of work. I would be extremely surprised if they were not already doing it. If someone reports an issue going back to a specific version, they need to be able to install that specific one.
I think you're overestimating the effort / underestimating Apple resources here. Ubuntu keeps all released packages, going back to 2004 online for example. (And isos, even for the effectively dead powerpc) Microsoft's symbol server is likely also a way bigger project than just preserving images of all versions.
> and that's not something they really do
Are you actually speaking for Apple engineering here? The version not being advertised on the update servers does not mean it's not available.
Also, there's the question of supply chain audit. I doubt their security would accept not being able to say "this is/isn't what we shipped at the time".
This is written a bit confusingly, if you don't already knows how it works (even though a later reply to a comment tries to clarify it). Because it sounds like the disk is fully and instantly validated at every boot, and "the startup process halts" if there is a mismatch.
Instead, depending a little on the actual implementation, most hashes will very likely only be validated when the corresponding block is actually accessed (up the entire path to the root, where the higher level hashes may be validated now or may have been validated already; this again is an implementation detail).
The result is also, that some blocks may never be validated during normal operation, simply if some files are never accessed. Unless there is some explicit whole disk validation sometime, e.g. during installation, though any corruption happening after that last "full disk check" will still lay dormant until access or the next full check.
But since this is a tree of hashes, it still provides all security and integrity benefits to anything that does get accessed. What never gets accessed does not matter per definition.
>In case of mismatch, the system assumes the data has been tampered with and won’t return it to the requesting software.
Whether this then leads to a kernel panic later on probably depends on if it's the kernel that needs this file, or a userspace process
In most situations, you likely just want the regular I/O error and then whatever happens downstream of that.
Also, is there a post about this zfs on apple silicon error thing?
This isn't dunking on Apple putting together a reliable product in their own walled garden using their own full array of vertically integrated options at all, but at the same time the ZFS team can be entirely correct given a very different problem space.
While it has to happen occasionally, it does seem to be a pretty rare event in my experience. Drive actually failing are way more likely to happen it seems.
I don't buy that.
- a process opaque to us.
- seems to be not failing at scale enough to make it to press.
- therefore the chance of [random data corruption] is remote
versus ex. [in 2016](https://physicsworld.com/a/cosmic-challenge-protecting-super...), couple UK PhDs laboriously chart 55,000 bit flip errors on a 100 node cluster.
This allows you to only hash those parts of the tree you actually want to read.
In this case they're not selective, so the tree approach doesn't save the reads. (Just makes it easier to identify which part failed)
That's uncached, so it's not that quick at boot. 9GB would take ~19s then. Block would go faster, but not an order of magnitude faster.
It's a Merkle tree, apparently, so parallelizing the checksum should be trivial.