That should have been in the initial implementation.
That should have been in the initial implementation.
I'm increasingly realising that most of the code out there isn't tested anywhere near as well as we think it is. Most programmers stop when the feature works, not when the feature is bulletproof.
I've been writing a database storage engine recently. It should be well behaved even in the event of sudden power loss. I'm doing "the obvious thing" to test it - and making a fake, in-memory filesystem which is configured to randomly fail sometimes when write/fsync commands are issued (leaving a spec-compliant mess).
I can't be the first person to try this, but it feels like I'm walking over virgin ground. I can't find a clear definition of what guarantees modern block devices provide, either directly or via linux syscalls. And I can't find any rust crates for doing this programatically. And googling it, it looks like many large "professional" databases misunderstood all this and used fsync wrong until recently. Its like there's a tiny corner of software that works reliably. And then everything else - which breaks utterly when the clock is set to 2038 because there aren't any tests and nobody tried it.
I half remember a quote from Carmack after he ran some new analysis tools on the old quake source code. He said that realising how many bugs there are in modern software, he's amazed that computers boot at all.
I'm unaware of any research newer than Dan Luu's post on filesystem error handling.
https://www.sqlite.org/testing.html
3.2. I/O Error Testing
I/O error testing seeks to verify that SQLite responds sanely to failed I/O operations. I/O errors might result from a full disk drive, malfunctioning disk hardware, network outages when using a network file system, system configuration or permission changes that occur in the middle of an SQL operation, or other hardware or operating system malfunctions. Whatever the cause, it is important that SQLite be able to respond correctly to these errors and I/O error testing seeks to verify that it does.
I/O error testing is similar in concept to OOM testing; I/O errors are simulated and checks are made to verify that SQLite responds correctly to the simulated errors. I/O errors are simulated in both the TCL and TH3 test harnesses by inserting a new Virtual File System object that is specially rigged to simulate an I/O error after a set number of I/O operations. As with OOM error testing, the I/O error simulators can be set to fail just once, or to fail continuously after the first failure. Tests are run in a loop, slowly increasing the point of failure until the test case runs to completion without error. The loop is run twice, once with the I/O error simulator set to simulate only a single failure and a second time with it set to fail all I/O operations after the first failure.
In I/O error tests, after the I/O error simulation failure mechanism is disabled, the database is examined using PRAGMA integrity_check to make sure that the I/O error has not introduced database corruption.Every time I play video games and see that "Don't turn off your console when you see this icon" I die a little inside. We've known how to write data atomically for decades. I find it pretty depressing that most video games just give up and ask the user to make sure they don't turn their console off at inopportune moments.
And I don't even blame the game developers'. Modern operating systems don't bother giving userland any simple & decent APIs for writing files atomically. Urgh.
Look yourself in a mirror. Do you even comprehend yourself? You're an incredibly big bunch of cells, of which only very few of them have any chance of continuing on. If you're male, it's not even real continuation, it's just part of a molecule.
We're hardwired to seek out and go with the most superficial of models, I guess because that's the most efficient way to go about in life.
Sure; but thats the exact reason drug discovery is so difficult. If we understood the human body in its entirety like we understand computers, we could probably cure cancer & aging.
Our capacity to write correct software depends entirely on being able to build mental models of how the machine works. The deep stack of buggy crap that we just take for granted these days makes software development harder. The less understandable and the less deterministic our computers, the worse products we build. And the less effective craftsman we become.
ext2 was first released in 1993 and has had impressive work on being extended far enough to keep up with growing storage needs, but I suspect the original designers figured it'd have been tossed aside well before 2038 would pose a problem. Unfortunately that assumption has proven to be wrong (unless maybe we do junk ext[24] within 15 years' time, but it would be extraordinary unlikely that every last machine running it will do so).
If you told the developers in '08 that ext4 would not be replaced in 30 years time, they would laugh at you (or cry).
ext4 has all the right hooks - if you use the "large" inodes, which appear to be done in a backward compatible fashion to ext2.
except edge cases in the user mode e2fsprogs/mkfs.ext4, etc. where it has to handle both small and large inodes it gets kinda complicated. I made a working patch, but it's just too icky. I think e2fsprogs needs to just deprecate the old small inodes and it would be clean.
Look man, it's Theo Ts'o maintaining the thing. Don't worry it won't be fixed correctly.
Also, I will point out that Y2038 was well known to DMR/BWK/etc. when they built the thing in 1970's. It's as useless as saying today "they should have known about the year 9223372036854776878 problem". In year 9223372036854776873 I don't know anything else other than everybody is going to panic.
A grim/realistic view, but Theo is 55yo. As we get closer to 2038, he may not care much about it anymore, or may even die before. I really hope it's someone else we talk about as maintaining it soon.
And I'm sure there are several orders of magnitude more of them than _that_ xkcd cartoon imply...
[1] Except when that "one" is me. You're all welcome to solve any problems in code I leave behind, I no longer will care. Whether that's a "bus factor" or a "won the lottery factor".)
But projects that have a bus-factor ~1 are often some of the tightest, best ones.
So ¯\_(ツ)_/¯, don't worry so much. Hopefully the owner leaves some keys and contingency plans, but otherwise, carry on as normal. The GPL and other open source licenses are good enough backup insurance if they didn't.
Knuth is 85 and still active. Kernighan is 81, likewise.
Torvalds is 53 and he's just starting to mellow out and grow up (a little, not too much).
Anyone can get hit by a bus at any age. So the age at which you stop being productive or kick it is highly individualize. I don't know Theo Ts'o, but I have no evidence of his eminent demise.