Filesystems can experience at least three different sorts of errors
utcc.utoronto.ca
utcc.utoronto.ca
https://www.man7.org/linux/man-pages/man8/btrfs-device.8.htm...
Man page includes definitions of the 5 kinds of errors tracked.
The article mentions structural errors. Sounds like this is the detection of an inconsistency. These aren't counted on btrfs, but are logged. Anytime the read or write time tree checker finds a problem, it results in the filesystem going read-only to prevent (further) confusion from ending up on disk. These are exceptionally rare, I've never seen one on any of my filesystems; but have seen it catch things like bitflips due to bad RAM, i.e. the checksum was computed correctly on already corrupted (meta)data.
https://patchwork.kernel.org/project/linux-btrfs/patch/024e4...
* - not a perfect analogy, but I'm hard-pressed to think of good off the shelf tools for what they're looking for. I guess reiserfsck's infamously side-effect laden --rebuild-tree would probably be closest...
(I am acquainted with zdb -r and import -T; neither helps you if there's not enough metadata consistent to get enough of a pool structure in memory to 'import', but one could still conceivably salvage some data in that case.)
The point of those is that the only time a pool is truly unusable is when you can't even import it.
The issue is, I think, people are expressing a desire for a tool that can still salvage data even when you've gone through all the (128/N) -T options and found them all unhelpful.
For more on this, see https://utcc.utoronto.ca/~cks/space/blog/solaris/ZFSScrubLim...
(The tl;dr is that a fsck on an ordinary filesystem has to walk the directory tree to find everything. However, ZFS maintains a separate list of active inodes and a scrub can just walk over them and check the checksums of all of their data blocks. It doesn't have to, for example, read a directory's contents to find further files to scrub.)
The checksum/mirroring mechanisms cannot fix any structural error when the filesystem is doing something and finds and inconsistency.
ZFS chose to not have a fsck out of pure arrogance, not because scrub is a proper substitute. ZFS developers believed that corruption bugs produced by code can be fixed by providing Bug Free Code (tm). That, and the fact that errors due to media corruption will be fixed with checksums and mirroring, made them believe that they could make fsck a thing of the past. Other modern file systems mimicking ZFS are developing a fsck, despite having scrub-like functionality.
...as I said, and your reply proves again:
> It's surprising the large amount of people I have found who are incapable of conceiving the notion of what this post calls "structural error" in modern file systems such as ZFS
People _really_ wants to believe ZFS has some kind of magic.
What's your point here? That ZFS can't correct logic bugs in it's own implementation that would lead to structural errors? (What system could?)
Another example that I have here right now, is a directory that says the following on "ls"
# l pg_stat_tmp/
ls: cannot access 'pg_stat_tmp/global.stat': No such file or directory
[...]
-????????? ? ? ? ? ? db_0.stat
The btrfsck says stuff like "parent transid verify failed" and "ERROR: child eb corrupted". scrub finishes without errors.> That ZFS can't correct logic bugs in it's own implementation that would lead to structural errors?
Again, I can't speak for ZFS, but the problem with btrfs is that for example in ext4 you have fsck that will fix such errors (sometimes losing the affected files). But in btrfs, the fsck is mostly "beta and do not use and it can't fix that, just move the data elsewhere, create the FS from scratch and move data back and hope it won't happen again".
This type of error can be detected with sufficiently robust integrity checking e.g. some type of durable Merkle tree, but ensuring that integrity checking can reliably detect phantom writes has a relatively high performance cost so many storage systems just assume it will not happen, given the low prevalence and high cost. FWIW, I think this is the correct tradeoff for many storage systems, where it is not a highly probable source of data loss in practice and some types of replications architectures make it relatively straightforward to detect after it has occurred even if not immediately.
IIRC they do distinguish between command/interface errors and medium errors, which is somewhat analogous to the filesys I/O and integrity errors discussed.