Linux: Ext4 data corruption in 6.1.64-1
bugs.debian.org
bugs.debian.org
In the upstream Linux kernel there were two fixes posted months from each other, one for direct io [0] and the other one for ext4 [1]. The ext4 one was marked for backport to stable (CC: stable@vger.kernel.org), the other was not. The problem is that these commits depend on each other for things to work properly. If you have both, you're fine. If you have only the backported one, you have a problem.
Which versions are affected? We know for sure that 6.1.64 is affected, 6.1.55 is not (because it doesn't have the commit). As of right now, 6.1.64 is still marked as "stable" in Debian [2] but if you actually try to install it from the official mirrors (deb.debian.org), you will get error 403. The fix is included in version 6.1.66 which will soon be available.
The issue seems to be only highlighted in the context of Debian but it is not specific to it. The issue is/was in the official upstream release.
[0] https://github.com/torvalds/linux/commit/936e114a245b6e38e0d...
[1] https://github.com/torvalds/linux/commit/91562895f8030cb9a04...
[0] https://github.com/torvalds/linux/commit/7e7efdda6adb385fbdf...
[1] https://github.com/torvalds/linux/commit/076fc8775dafe995e94...
[2] https://cdn.kernel.org/pub/linux/kernel/v6.x/ChangeLog-6.1.6...
The issue is that some people don't have the time, money, and/or willpower to invest in backups. Some people only have one volume, or one copy of each volume at least, and need to recover their data using only that one volume if something gets corrupted. That's what "corruption you can recover from" means. Using a backup is not recovering from corruption; it's replacing the corruption with the pre-corrupted data. Recovering from corruption involves actually transforming the corrupted data back into non-corrupted data.
So you recover your data, but not from the corruption. You recover it from a backup. Similar, but not the same thing.
This is a silent data corruption bug (data is written at the wrong offset, IIUC) with no known method to detect the error condition except to compare the resulting data with a known-good copy. The next backup run will happily accept the corrupted file contents into the backup vault, and unless someone happens to stumble upon the corrupted data, the last-known good copy will eventually be rotated out.
Even linux tools like shred have given up saying they can actually delete data from disks due to how SSD's work these days.
Which emphasizes the importance of enabling full disk encryption immediately whenever you start using a new device--BitLocker if you're on Windows, FileVault on macOS, LUKS on Linux, etc. Trying to decrypt data is much harder than reconstructing deleted data on a stolen drive.
How to do that reliably is another question.
https://support.apple.com/guide/disk-utility/erase-and-refor...
> Note: With a solid-state drive (SSD), secure erase options are not available in Disk Utility. For more security, consider turning on FileVault encryption when you start using your SSD drive.
So if you set up a Mac without FileVault you can never erase everything.
At least with my Lenovo I can do the secure erase.
Of course providing something like 10% of hidden extra capacity would extend the life of the drive significantly, but are manufacturers really doing that and are not mentioning it in their marketing materials? I never heard that they do that.
>but are manufacturers really doing that and are not mentioning it in their marketing materials
SSD manufacturers have been caught out repeatedly with their in-built deletion API claims. Recovery of a significant % of files is possible nearly always without tampering having occurred.
Every SSD has considerably more physical blocks than reported blocks. They have to, SSD bad blocks are common and number of writes is quite limited compared to hard disks.
https://download.semiconductor.samsung.com/resources/others/...
Not sure how you missed it, as it's not a new thing. Straight from the horses mouth, if that helps: :)
https://www.kingston.com/en/blog/pc-performance/overprovisio...
Note that it's not just a Kingston thing, it's common across SSD manufacturers. Or at least with the major ones.
People who want secure deletion generally want their storage device to remain functional afterwards. Same reason they don't simply destroy the disk to get rid of its contents: those things are expensive. Overwriting the entire SSD every time you want a file gone just isn't a solution.
Please don't trash physically good drives with a hammer. It's not good for the environment (or the drives!) when you have such a simple technology at your disposal!
Also have you done any work in digital forensics?
Eagerly awaiting your response.
Assuming the firmware doesn't lie about it does:
* https://www.zdnet.com/article/flaws-in-self-encrypting-ssds-...
P.S. Squirrel!
“How is it not serious if she died from it?”
“She made an unsuccessful recovery”
- how badly can you phrase yes?
[Cuts to other scene]
- No
Such a good show.
From the link in https://news.ycombinator.com/item?id=38591444 :
> file position is not updated after direct IO write and thus [we] direct IO writes are ending in wrong locations effectively corrupting data.
This may be a major problem for hypervisors using ext4-backed disk images. It wouldn't surprise me if qemu opens image files with O_DIRECT by default, but it would surprise me if professional shops were using ext4 as the backing store rather than a provisioning abstraction layer like dm_thin. Still, anyone using kvm on top of ext4 must not be having a good day.
"sudo systemctl stop unattended-upgrades.service" ...wasn't able to prevent unattended-upgrades from going on ahead and just upgrading (to this problematic kernel) anyway.
Unintuitively, the "right" way to disable unattended-upgrades is:
"sudo dpkg-reconfigure unattended-upgrades" ...and choose "No" when asked.
FTFY. Helps keeping them apart to not use the same word for both ;)
systemctl disable --now unattended-upgrades.timer;
systemctl disable --now unattended-upgrades.service
No need to uninstall or mask.That said, I certainly wouldn’t blame anyone for blindly trusting a kernel in Stable. This hurts.
I'm not sure if that's a standard value, but it made me chuckle a little.
Severity: `
“If my file system developer hasn’t been through the Full Metal Jacket experience, I don’t trust the code they write!”
“Linux is going down the toilet! What we really need is leadership from a guy who’s VERY angry all the time!”
Sure, giving good candid feedback is a gift, but degradation isn’t a necessary part of honesty.
It seems like a commit from the main branch was backported to stable branch, but it actually depended on another one which was not backported.
Linux kernel maintainers heavily backport commits as it seems the distinction between bug-fix and new features is hard to do. Also, bug-fix can depend on new features..
I'm always a bit scared when I look at the amount of commits in a single stable kernel point release, plus the fact that these point releases come out every week or so. The velocity of change in what should be a stable kernel is very high.
I'm wondering who exactly wants to run a stable kernel that receives all these updates so fast. If you want all the latest and greatest, just use -latest?
I find the kernel is now the least stable bit of debian stable
I've had LTS kernels break the networking on my workstation (with a 10 year old onboard NIC) and they deliberately broke ZFS in the past
in LTS there shouldn't be ANYTHING other than security fixes
ZFS's own recent file system corruption issue is in roughly the same category of edge case, but accessible to reasonable if niche workloads.
Stable process is cherry pincking thousands to tens of thousands of patches from the current master kernel branch into years old kernel branches, spraying tens of thousands of emails at original patch authors, hoping (With some limited testing on top) that the resulting frankenkernels will still work and all these patches will have satisfied dependencies applied, too.
I don't trust this process very much, and just run the latest stable branch for the latest kernel release, only. Staying with older stable release branch for too long seems too risky, unless you're some bigcorp that can afford the testing required, or you're running some highly mainstream setup that is probably covered by tests done by the stable team, and testing teams they cooperate with.
You know it's not that hard to debunk this:
> You'll then be left with a kernel-6.X.? directory, containing both an unpatched 'vanilla-6.X.?' dir, and a linux-6.X.?-noarch hardlinked dir which has the Fedora patches applied. [1]
WOW fedora patches! Sounds to me like they're not shipping vanilla.
I have problem with backaptching 10s of thousands of changes to years old kernels by people who don't really understand the changes or consequences, like 5.4, 5.15, 6.1 or whatever. Not patching up 6.6 or 6.7 kernel with a few out of tree patches, where they mostly understand what they're doing and can test the limited set of changes they're applying.
Anybody who cares that much is very likely already compiling their own kernels (I speak for myself here). It doesn't make sense for distros to do the extra work to support it.
Just expanding on this a bit: my Debian laptop has a Kconfig with modules disabled that only includes the exact set of drivers it needs. It takes ten minutes to rebuild on the laptop when I pull from git. Even if Debian did all the work to let me automatically install the latest vanilla kernel... I'd still build it myself.
I remember that there was a fairly severe one which was caused by patching OpenSSL I think? But I remember the change they made being fairly weird and no one understood why but it was easy to see that it would introduce a vulnerability.
In this case Debian's current process is good - it's kernels track kernel.org stable releases. This debian bug is responsibly flagging "for visibility" that a serious bug has been discussed and fixed upstream.
[1] https://lore.kernel.org/stable/20231205122122.dfhhoaswsfscuh...
"properly sync file size update after O_SYNC direct IO": https://cdn.kernel.org/pub/linux/kernel/v6.x/ChangeLog-6.1.6...
"update ki_pos a little later in iomap_dio_complete": https://cdn.kernel.org/pub/linux/kernel/v6.x/ChangeLog-6.1.6...
This post explains the relationship between the two commits: https://lore.kernel.org/stable/20231205122122.dfhhoaswsfscuh...
— someone who lost data due to btrfs bug
Combine ext4's dumb but robust approach to journaling and robust metadata layout (for example inodes are statically allocated) with fsck.ext4 (which got refined for years) and you can recover from any situation.
To give you an example, fsck.ext4 will happily carve a working filesystem out of random data, as long as there is a valid superblock. Seriously - try it yourself:
# create a working filesystem image
dd if=/dev/zero of=test1.img bs=1M count=256
/sbin/mkfs.ext4 test1.img
# write file with random data
dd if=/dev/urandom of=test2.img bs=1M count=256
# copy superblock into random file
dd if=test1.img of=test2.img bs=1024 count=4 seek=1 skip=1 conv=notrunc
/sbin/fsck.ext4 -fy test2.img
sudo mount test2.img /mnt/somewhereLike: "Wow, but why should I care?" I'm not sure being able to fsck a random disk image shows resiliency. Doing this could do all sorts of nasty things to data you actually care about.
So -- you think redundant metadata is a bad thing? Try wiping your metadata and then trying to fsck that random disk image.
Again, has this been the source of data corruption you're aware of? This seems like a "Maybe, it could be this way, but I don't know" kind of take.
What application does this have for recovering user data? Would it not be more interesting to see how much of the user data can be recovered? Or are you implying that the random data can symbolize the user data here and much of it would be recovered?
> This should block just the buggy kernel. Which might help with unattended upgrades problem or just being forgetful. It might even uninstall the buggy kernel, though I didn't test that. And it shouldn't impact upgrading to 6.1.66 when it's available.
> create a file: /etc/apt/preferences.d/buggy-kernel
> with the contents:
# avoid kernel with ext4 bug
# 1057843
Package: linux-image-\*
Pin: version 6.1.64-1
Pin-Priority: -1
> (the comment isn't required but is helpful for remembering why this file
is around in 6 months.)- I was expecting the package to either be revoked, or to have a new `6.1.64-2` version being published with the previously good known state.
Maybe it wasn't done because it was not possible (mirrors being write-only, and maybe other complications in publishing a 6.1.64-2).
- I was expecting some guidelines to be published for affected machines. I've seen questions being asked [1], but no answers to them yet (so users are not sure if it would be safe to rollback to the previous kernel version (6.1.55-1) or not).
Maybe it's because the problem is not well understood yet, or maybe there are not enough people available to answer such questions on a weekend?
If someone can provide some context about "how these things work", it would be interesting to learn more about this.
[^1]: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1057843#33
If this all seems archaic, keep in mind that all this has been designed 20+ years ago, when bandwidth was quite a bit less abundant.
Related: "Debian 12.3 image release delayed" https://lists.debian.org/debian-announce/2023/msg00005.html
This doesn't help debians reputation.
I should just move on to FreeBSD.
Context: This is seen in the following environments: > > > * dragonboard-845c > > > * juno-64k_page_size > > > * qemu-arm64 > > > * qemu-armv7 > > > * qemu-i386 > > > * qemu-x86_64 > > > * x86_64-clang
There were some old UNIX variants where O_DIRECT actually bypassed the filesystem cache, but Linux's cache is coherent, so reading a file immediately after an O_DIRECT write completes is guaranteed to give the new value. That is more sane than the cache-incoherent approach (how do you know your write is going to a page that is clean in OS cache in the other UNIXes?), but also eliminates most of the code-path-length benefit of O_DIRECT.
Also, if I remember right, O_DIRECT doesn't bypass the I/O scheduler on Linux. That, and bypassing the cache are the two main benefits of the old API, and you get neither.
As a bonus, filesystems like tmpfs passive-aggressively return error if you try to use O_DIRECT.
Linux achieves that coherency by making O_DIRECT invalidate the page in the page cache. I think it paints a more accurate picture to say the caching is disabled with O_DIRECT, not that it is coherent (although it certainly is).
https://github.com/zabbly/linux
Quote "As those kernels aren't signed by a trusted distribution key, you may need to turn off Secure Boot on your system in order to boot this kernel."
Anybody would be so kind to tell me why I have been downvoted?