ZFS's own recent file system corruption issue is in roughly the same category of edge case, but accessible to reasonable if niche workloads.
ZFS's own recent file system corruption issue is in roughly the same category of edge case, but accessible to reasonable if niche workloads.
Combine ext4's dumb but robust approach to journaling and robust metadata layout (for example inodes are statically allocated) with fsck.ext4 (which got refined for years) and you can recover from any situation.
To give you an example, fsck.ext4 will happily carve a working filesystem out of random data, as long as there is a valid superblock. Seriously - try it yourself:
# create a working filesystem image
dd if=/dev/zero of=test1.img bs=1M count=256
/sbin/mkfs.ext4 test1.img
# write file with random data
dd if=/dev/urandom of=test2.img bs=1M count=256
# copy superblock into random file
dd if=test1.img of=test2.img bs=1024 count=4 seek=1 skip=1 conv=notrunc
/sbin/fsck.ext4 -fy test2.img
sudo mount test2.img /mnt/somewhereLike: "Wow, but why should I care?" I'm not sure being able to fsck a random disk image shows resiliency. Doing this could do all sorts of nasty things to data you actually care about.
So -- you think redundant metadata is a bad thing? Try wiping your metadata and then trying to fsck that random disk image.
Again, has this been the source of data corruption you're aware of? This seems like a "Maybe, it could be this way, but I don't know" kind of take.
What application does this have for recovering user data? Would it not be more interesting to see how much of the user data can be recovered? Or are you implying that the random data can symbolize the user data here and much of it would be recovered?
— someone who lost data due to btrfs bug
Stable process is cherry pincking thousands to tens of thousands of patches from the current master kernel branch into years old kernel branches, spraying tens of thousands of emails at original patch authors, hoping (With some limited testing on top) that the resulting frankenkernels will still work and all these patches will have satisfied dependencies applied, too.
I don't trust this process very much, and just run the latest stable branch for the latest kernel release, only. Staying with older stable release branch for too long seems too risky, unless you're some bigcorp that can afford the testing required, or you're running some highly mainstream setup that is probably covered by tests done by the stable team, and testing teams they cooperate with.
You know it's not that hard to debunk this:
> You'll then be left with a kernel-6.X.? directory, containing both an unpatched 'vanilla-6.X.?' dir, and a linux-6.X.?-noarch hardlinked dir which has the Fedora patches applied. [1]
WOW fedora patches! Sounds to me like they're not shipping vanilla.
I have problem with backaptching 10s of thousands of changes to years old kernels by people who don't really understand the changes or consequences, like 5.4, 5.15, 6.1 or whatever. Not patching up 6.6 or 6.7 kernel with a few out of tree patches, where they mostly understand what they're doing and can test the limited set of changes they're applying.
Anybody who cares that much is very likely already compiling their own kernels (I speak for myself here). It doesn't make sense for distros to do the extra work to support it.
Just expanding on this a bit: my Debian laptop has a Kconfig with modules disabled that only includes the exact set of drivers it needs. It takes ten minutes to rebuild on the laptop when I pull from git. Even if Debian did all the work to let me automatically install the latest vanilla kernel... I'd still build it myself.
I remember that there was a fairly severe one which was caused by patching OpenSSL I think? But I remember the change they made being fairly weird and no one understood why but it was easy to see that it would introduce a vulnerability.
In this case Debian's current process is good - it's kernels track kernel.org stable releases. This debian bug is responsibly flagging "for visibility" that a serious bug has been discussed and fixed upstream.
[1] https://lore.kernel.org/stable/20231205122122.dfhhoaswsfscuh...
"properly sync file size update after O_SYNC direct IO": https://cdn.kernel.org/pub/linux/kernel/v6.x/ChangeLog-6.1.6...
"update ki_pos a little later in iomap_dio_complete": https://cdn.kernel.org/pub/linux/kernel/v6.x/ChangeLog-6.1.6...
This post explains the relationship between the two commits: https://lore.kernel.org/stable/20231205122122.dfhhoaswsfscuh...