A way to do atomic writes
lwn.net
lwn.net
- Borrow some ideas from generational garbage collection. Young generation in SSD (or mirrored in RAM) with copying GC to get rid of old versions of fast changing data blocks.
- Utilize some deduplication techniques with content based signature.
However, LSMT is for relatively smaller data set, i.e. ordered key-value. It has worse write amplification than a simple append log. The level-0 memtable flushed to the write-ahead-log counts as one write. Writing to the level-1 sorted files counts as 2nd write. Merging the sorted files counts as 3rd write. There're 2~3 writes per change.
Also it doesn't offer help to address the frequent update block problem. All versions of a data change are written to disk. A merge is needed to get rid of the old versions.
But it has a number of good implementation ideas that can be borrowed.
It's mind boggling that this is still an issue, given the number of decades we've had applications (databases, and other things) that need to efficiently guarantee bits are on stable storage.
This is definitely not impossible. It's not even that difficult (and yes, I've written a number of transactional storage systems over the past 30 years).
I worked with a battery backup disk controller once, that had redundant failover to a duplicate unit. It worked great for many months with flawless failover. Except when one of the internal batteries died, the whole system deadlocked and the 24/7 nonstop cluster was down for days.
The more layers, the more unpredictable the failure modes.
That's the thing, you are communicating with your storage controller and these are just promises from your controller, not guarantees. Once you try to read the data back sometimes in the future there is no guarantee that you will succeed. A lot can happen between your controller reporting successful write and you retrieving the data: software bugs and false promises, hardware problems, operational problems, disasters, etc. There is a limit of how high of a probability of data retention a typical single server in a typical server room can achieve.
Saying "we cannot achieve anything because there might be firmware bugs" is technically true, but completely counter-productive. Adding "A little bit better or worse is not that big of a deal" is just bad engineering -- can you imagine doctor saying, "you might get hit by a car at any time, so I decided it is not worth it to heal you"?
You can still lose everything, you can't control all failure modes. But you can plan for and protect against common disasters.
The original discussion was about the lack of OS support for proper flushing and making data stable. I think the lack of decent support for this is a decades-long travesty; claims that the lack of this functionality doesn't matter because "there might be firmware bugs in the storage system, so why bother?" are specious and unhelpful.
The problem is that using that support kills performance, so apps often don't.
But yes, it's absolutely possible for all of this to be vertically integrated, and agreed that it's ridiculous that it hasn't happened. Most systems shouldn't care, but the ones that do care really really care.
My rule of thumb is that if I want serious reliability I make sure my data winds up on three different pieces of hardware and includes CRC codes. So, basically ZFS or some cloud/cluster equivalent like S3.
On the other hand, for day to day work on your dev machine, the drives are so crazy reliable that you just don't worry about it. Which, of course, can bite you if you translate your day to day experience to production at scale. Everything breaks at scale.
[1] https://www.postgresql.org/docs/devel/wal-reliability.html
This. Data is real when it has been acked by a quorum with independent power. Three is because two doesn't give you enough slack to maintain it.
Can you give more details on this? (What is the solution, and what does it cost? What kind of computer you use it on? How large is the memory?) Very curious.
The battery-backed DRAM really is an implementation detail. Things will execute correctly if you don't have it, but it's usually a huge performance win.
In this case we're using Nimble SANs (Nimble is now owned by HP and suffering customer service rot, oh well). The immortal DRAM is fairly small (8GB per controller head?). The storage arrays are petabyte scale, all flash, with many, many 16 and 32 Gbit fiber channel connections. Cost is a few million per SAN instance, of which you need several for real durability (and a replication scheme for remote storage, which I'm not going to discuss here).
For anyone interested, especially the -N variant:
* https://en.wikipedia.org/wiki/NVDIMM
Back when ZFS was still new-ish, and SSDs were still expensive-ish (~2008), people were experimenting with using ZILs (ZFS Intent Log) on these types of devices:
* https://techreport.com/review/16255/acard-ans-9010-serial-at...
SSDs have come down in price since than, so people don't bother with RAM disks as much now.
I will add that the different flavors of fsync, and what they do (and don't do), is always a source of entertainment in database engineering chat rooms.
Also, a while back there was another paper that had a much better API as it allowed atomicity across files: https://github.com/ut-osa/txfs (better, but obviously not perfect: https://twitter.com/CAFxX/status/1017204557003173888)
Isn't this Multiversion concurrency control[1]?
1. https://en.wikipedia.org/wiki/Multiversion_concurrency_contr...
This problem will be solved if, as looks likely, the openat2() syscall is merged for the path resolution restriction patchset. As a new syscall, it can do what open() should have always done and return EINVAL on flags it doesn't understand.
close() flushes the application buffer into the vfs buffer, but it does not guarantee the vfs buffer is flushed to the disk buffer, nor does it guarantee the disk buffer flushes to disk. A successful fsync() call guarantees these things.\*
\* except if the disk drive controller lies about flushing its buffer to disk, but there's nothing you can do about that besides buying a better disk.
No, close() only closes the file descriptor. fclose() potentially also flushes a userspace buffer.
I suppose that's not fast enough for a DB.
Achieving that guarantee is not impossible, but it needs to be explicit. It can't be inferred from other API calls.
That's how you end up with stuff like OS X's "no really, fsync" param [1], or Motorola shipping nobarrier on their phones. [2]
[1] https://github.com/google/leveldb/issues/203#issuecomment-55...
[2] http://taras.glek.net/post/Followup-on-Pixel-vs-Moto:fsync/
My memory is vulnerable to row hammer due to vendors cutting corners while pushing for increased DRAM density.
And my - supposedly - non-volatile storage is broken due to vendors gaming benchmarks with fsync.
Is there any component in a modern consumer computer that isn't fundamentally broken in one way or another?
Look at how much slower CPUs are without those speculative execution tricks. I can buy gigabytes of RAM with a 20 I found in an old jacket. You can explore entire worlds consisting of gigabytes of high res textures and mesh data in near real-time, while downloading 4 new albums off the internet.
Broken?
But do I care? All my computers, except the really ancient phone I use, are snappy.
Of course there are cases were CPU speed matters, but I don't know if they are any less obscure than the cases where timing attacks are a risk.
Yes, it is breaking the promise of what it is supposed to do. If fsync() was defined as "will ensure your data is on disk, unless that's kinda slow, then who knows", then the behaviour would not be broken, just potentially useless for many applications. But if you promise to ensure something is stored on disk, and then don't, that's the definition of being broken.
On windows FILE_FLAG_WRITE_THROUGH ("Write operations will not go through any intermediate cache, they will go directly to disk").
It's all there.
I agree with other poster, just cos you read back the file and it compared byte-for-byte, unless large it's likely to have come from the OS's RAM file cache.
I always wondered about writable disk-backed pages -- back when I last looked it up munmap() manpage made no suggestions about whether it would flush the data or whether you need an explicit msync() to make it happen. Since pages could get written back without an explicit msync(), it's always tricky to speculate how the behavior should be based on observation.
In case you end up creating such a tool, please share it with us.