Tangentially, I would love to have a FUSE filesystem with a.) minimal build dependencies, b.) some sort of CLI interface, and c.) the ability to, say, forget to flush un-fsynced data to "disk", allowing us to simulate a power failure. There have been a bunch of research projects on this front, and they find bugs spectacularly. I bet this approach would also find errors in distributed systems, but I've yet to find one that really has the right shape for use with Jepsen.
> I/O errors are simulated in both the TCL and TH3 test harnesses by inserting a new Virtual File System object that is specially rigged to simulate an I/O error after a set number of I/O operations. As with OOM error testing, the I/O error simulators can be set to fail just once, or to fail continuously after the first failure. Tests are run in a loop, slowly increasing the point of failure until the test case runs to completion without error. The loop is run twice, once with the I/O error simulator set to simulate only a single failure and a second time with it set to fail all I/O operations after the first failure.
Of course distributed filesystems are even harder but we can see that even writing to a local file is surprisingly complex and in many (most?) apps probably not really guaranteeing proper data consistency.
Imho the whole Posix filesystem API is flawed to begin with and it would be great if a modern replacement would emerge.
Not as elaborate as Jepsen, but there has been some work:
- Kirk McKusick's papers and work on BSD fs log semantics
https://www.researchgate.net/scientific-contributions/Marsha...
I believe the term to look for is "soft updates."
- the BSD/NeXT file test program "fstest.c", used on local and NFS (Samba), which found many bugs in popular fs using simple operations. The ZFS team also has a version.
You can Bing versions of that by using quotes "fstest.c".
- the Luster/Gluster maintainers/consulting team used to just untar emacs on their distributed fs buildouts and see how many nodes left the cluster. (They lived off DARPA funding basically, and were paid to configure and install distributed fs for US govt/military supercomputer installations, and fix the underlying bugs as found.)
- Ironically, the Ceph team did not own any commercial storage devices, so just tested on regular linux machines.
- Reiserfs 3 was the first GA log fs on linux (default on SUSE), so I was one of the earliest US users in production.
SUSE's rep called me a liar at trade show, saying "nobody uses our distro in the US. :) It worked well on email server loads, and could delete 1 million files in a directory in under 1 second. I followed the development of v4, but the "wandering logs" and "dancing trees", etc. kind of wigged me out.
https://en.wikipedia.org/wiki/Dancing_tree
Source: DBA and storage engineer.
EIO: Error Handling is Occasionally Correct https://www.usenix.org/legacy/events/fast08/tech/gunawi.html
Evaluating File System Reliability on Solid State Drives https://www.usenix.org/conference/atc19/presentation/jaffer