A Linux kernel without struct buffer_head
lwn.net
lwn.net
Really makes me consider subscribing.
More on the topic I'm wondering how the remaining users of buffer_head could possibly migrate? Especially since we're talking filesystems I'm guessing we don't want any change in behaviour as it could result in a loss of data
As a systems programmer I believe it will have useful content for me, and I want to support the initiative
This is awesome, but I don't see the `buffer_head` getting replaced anytime soon. It's so baked into existing filesystems, I can't see it going away for at least half a decade.
Also, love lwn.net—great find!
Nobody in GNU/Linux tests old binaries; only whatever they recompiled for their current distro, and screw the rest.
The kernel is pretty good about maintaining syscall compatibility.
> Microsoft had gone as far as emulating undocumented/undefined API behavior for misbehaving/buggy applications
Yeah, I believe Chen writes about code for SimCity. This is where Microsoft loses balance. They keep backwards compatibility to a fault. Literally a fault. At some point the kernel will no longer have `struct buffer_head`, at least not in the mainline, though it may be a while. That backwards compatibility will go away. Microsoft has been and continues to be bad about end-of-life for their software. They keep things around that should be gone, and remove or stop supporting things that should stay.
> Sad but true: Once you document a file format, it becomes a de facto API.
So if I build a system that only uses XFS/BTRFS, do I get anything by explicitly disabling it, or are those filesystems already getting the performance benefit just by not using it themselves? Is BUFFER_HEAD just a goal to show kernel devs what's left, or does it have practical value to end users?
config EXFAT_FS
tristate "exFAT filesystem support"
+ select BUFFER_HEAD
select NLS
select LEGACY_DIRECT_IO
help
which means that if CONFIG_BUFFER_HEAD is disabled then all those filesystems get disabled too (or if you pick a filesystem to enable that still needs struct buffer_head then the CONFIG_BUFFER_HEAD option is enabled).So basically it's only to show the kernel devs what is left. You already got the performance benefit for XFS etc as it just isn't used by those filesystems.
I’m referring to where the article says ext4 “still makes heavy use of buffer heads,” implies this has been the case since ext2 and says “the buffer cache was deeply wired into both the block subsystem and the filesystem implementations.”
Naively I would think the job of the filesystem is to handle reads and writes to disk and that a well designed one would leave caching to other parts of the OS.
The other part of the kernel that deals with data in memory is the cache.
Rather than have two separate representations of "data in memory" and have to translate between them all the time, it makes sense to have one representation of "data in memory" that all the affected systems can use to talk to each other.
Even though the kernel now has multiple representations of "data in memory" due to advances in other parts of the kernel, it's still desirable to have just one representation shared between the cache and the filesystem, which is why people are talking about removing the old ones, and porting old code to use the new ones.
- Since a file is something filesystems do, it seems sensible to quite directly plumb the read() syscall into the filesystem, and the filesystem can then handle reading blocks from the actual disk, or caching previously read blocks. If an OS wants to support multiple filesystem types, it might even provide some common data structures and code for doing this caching of disk blocks. In Linux this being the "struct buffer_head".
- The other option being that managing memory is something kernels do, and it makes sense to centralize this management in a common code that can then make sensible decisions what to do with this memory. Such as giving it to applications that ask for memory, using it for caching file contents (page cache), and the crucial part, balancing these different usecases and making good choices what to evict. If you make this design choice, the read() syscall will not immediately be handed down to the filesystem, but rather it first asks the memory management subsystem whether the data happens to be in the page cache, and only if it isn't, it calls down to the filesystem to read that data from the physical disk, copy it to the page cache (potentially evicting something else to make room), and then copy it back to the application asking for it.
Now it seems that historically quite a few OS'es, Linux included, started off with the first approach. However later on it turned out that the second way of designing the OS is better, and thus Linux was slowly over time modified to mostly use the page cache. Though the struct buffer_head is still in use here and there.
So these filesystems are old. Even if it was a good idea to structure caching and filesystem block access separately (which as the sibling comment says, it isn't), the Linux developers were just getting started back then and it was much more important to get something working than to think about good abstractions.
The journaling layer (jbd) in ext3/ext4 was built on top of buffer_heads as buffer_heads were the way writes got tracked. Rewriting ext4 and jbd2 to use a new data structure to track writes to the journal and disk will be a lot of work. ext4 has a number of issues that make it a less desirable filesystem these days, so it's not clear the work will be worth it any time soon when filesystems like bcachefs, btrfs and xfs do so much better.