A File System All Its Own – specialized for SSDs
queue.acm.org
queue.acm.org
Maybe he didn't have old price data for SLC?
It won't be too hard to have a good filesystem that works over raw NAND flash but it will not work on older OSes, it will not work in the enterprise storage market and so there will be less buyers and thus it will cost more so no one will buy it and it will not be made.
Even the enterprise storage folks just want the damn flash devices to just work without the storage folks doing anything with them. It's taken to extremes sometimes and the flash vendors just do whatever they are told since there is a lot of market in whatever the software-defined engineers want. Except the engineers mostly want to deal with high level algorithms and to brag how fast their algorithm is without really thinking about the hardware. Hardware is hard. Besides they can do something with the hardware that is already on the market rather than envision something better.
TL;DR unless someone will hold the stick at both ends (software and hardware) no one will make a reduced layer solution.
When it comes to research, no one cares that much about what home users and enterprise-users-small-enough-not-to-use-custom-software-stacks want now. Case in point: I don't think many IT mangers were that eager to switch to using ZFS in production when it was announced back in 2004 (and ZFS wad been under development for years at that time).
I've considered doing a PhD project pushing and stretching the boundaries of SSD firmware/operating system/filesystems because I think there's a lot of improvement that can be done in this area. The cost of OpenSSD that the sibling comments mention wasn't even that much of a problem. I seriously don't think someone not associated with a research department somewhere would have the time and/or know-how to do original research and implement a working prototype. Hell, I might get a devkit, but I doubt I'll do anything interesting and original at the same time. Which brings us to the actual problem:
Documentation and NDAs. For lots of ICs, microcontrollers, processors you can freely get hardware documentation, programming manuals, etc. For flash controllers and high-density NANDs? Almost nothing at all. Maybe some stuff can be reverse engineered and you get some documentation for OpenSSD. But the NAND manufacturer won't tell you stuff that's really, really important about things like failure patterns, which would allow you to optimize error correction and wear leveling for example.
I don't think that you really need the inner information about NAND to do the original and innovative research. It really depends on the area you want to work on, the SSD firmware level might require that but the SSD makers are already on that route (some better than others). The other level is to not pay too much attention to the differences between NAND chips and just implement something at a higher level to push the hardware-agnostic smarts to the OS.
The OpenSSD also lacks documentation, last time I looked at it there was no info on how to do NCQ on the SATA interface, without which there is no talking about a speedy SSD.
Apple would be well-positioned here if they still cared about their Macs. HFS is due for a replacement anyway after 30 years. (it could be done on iOS devices too, but flash I/O performance doesn't seem to be a the major bottleneck for those uses)
HFS+ has some nice features. I don't know anyone not running OS X using it.
Anobit SSDs were the best I've seen so far in terms of consistent performance.
https://en.wikipedia.org/wiki/List_of_file_systems#File_syst...
I personally believe log-based file systems are a perfect match due to never saving the same file repeatedly to the same location (so provides built in wear-leveling) and one can optimize writes by always clearing the head of the log for the next write.
Take ZFS. I designed the flash integration for ZFS; it's used as a caching tier. ZFS is definitely not optimized for use with flash as its primary backing store. The same is true for some of the other filesystems in the list; offhand: CASL and WAFL.
Most of the rest are designed for embedded use cases, are research toys, or are embedded research toys.
The variability is caused by the "incompatible" NAND flash interface (read, write and erase), while the IO interface to the host system is read/write (and occasional trim to let the device know of unused pages). Therefore, another interface, other than a simple read/write is the holy grail. This interface might be one that give various guarantees for the user, e.g. atomic operations, etc. It doesn't need to only be an object / page store.
SATA is not going to work since it is a block interface only and not easily extensible in a sane way.
"For many years SSDs were almost exclusively built to seamlessly replace hard drives; they not only supported the same block-device interface"
The point of storage is to be able to put anything you want on it. That contract is the block interface, and includes the ability to change the filesystem. A file with internal structures is also a filesystem. The interfaces are fine. Change for change's sake should be avoided. (Providing a bypass, SSD-optimized interface is fine, but, ahem: "put down the crack pipes"... https://news.ycombinator.com/item?id=5541063 )
The article argues that we should change this contract.
> The interfaces are fine.
The article argues that they are not.
> Change for change's sake should be avoided.
The article argues that we should change them for performance's sake.
I also said providing a bypass was reasonable, but the article gives the impression that the block-level interface is yesterday's jam, and something less than the starting point. It is the starting point and will continue to be because block-level storage is the major use-case. Block-level access accounts for basically all bulk storage in /dev. Extra performance and features (ioctl calls, or a management interface) are gravy, but without the block-level interface, it is not accessible to 99.9% of all software and will not serve for general storage, including pre-existing filesystems, from FAT12 to Btrfs. There are applications for which those are interfaces. There's no reason to make it difficult to apply those layers. Looking forward, if it can't store current and unforseen filesystems out of the box (the block interface), even if that means a 50% reduction in speed, it doesn't deserve the name "mass storage." Noone wants a key-value store even if it is 100% faster. They may be faster, but it doesn't look like storage. And there's no need for it to look different either, because people expect to be able to use it like a block device. Take WD Green drives with their larger allocation sizes. There happens to be an optimal cluster size, but the interface is still that of block IO. If flash works best for a certain allocation size or other tweaks, or even a specific high-level formatting, fine, but it's still going to need to provide the block IO interface as a starting point if it's going to be used as a drive.
It's nice to see at least one punitive downvote though.
Unless your home and binary directories are stored on an NFS share.
Edit: Network block device, iSCSI, AoE, etc... the block interface is the lifeblood of the storage area network.
There is quite a bit of a chicken-and-egg problem here though, all the current filesystems basically assume that the underlying media never has any faults and if it does than the media problems are static and do not develop over time. This is obviously incorrect for flash, but even for rotating media it wasn't true. Since every OS requires a fault-free media the SSD vendors are working hard to make sure they provide a semblance of such fault-free media they make it harder to provide the best possible performance or a different trade-off than what they have taken.
If a lossy interface is acceptable, why couldn't SSDs simply expose a faster albeit lossy block device and, if necessary, an extended SMART or custom inspection method, and let the user take responsibility for wear-leveling, ECC, etc? It would be backwards compatible with other things by virtue of presenting the standard block interface. A common ECC+wear-leveling middle layer could evolve allowing use of standard filesystems and a common codebase for all flash storage, relieving the apparent burden on SSD vendors who would love to sell fast unreliable storage rather than reliable storage.
I think the chicken-egg problem is just an egg problem though, because even though I'd love lower level access to the unreliable bits, I have to expect that the market for unreliable storage is very small, not unlike the high-efficiency-but-sometimes-exploding toilet. :)
If reliable storage is the primary use-case, maybe the drives just need to be smarter to keep up. I'd rather an ASIC handle ECC, etc, transparently (for the same reasons I'd rather have a dedicated GPU) than run ECC (or 3D floating point software) on my general processor. If you inevitably want reliable storage and just wind up running ECC, etc, on the CPU, the speed gains disappear and we're back in something similar to a pre-DMA world with the main processor doing something that could be done in parallel by a dedicated chip. If the hard drive is the right place for the offload, I'd rather the economic pressures remain for the SSD vendors to optimize inside that black box, behind the standard reliable interface.
That said, again, I too would love finer-grain control.
The block interface itself is actually matching the flash, you read/write/erase in blocks, they may not be 512 bytes but rather 4k/8k/256k whatever works for the underlying hardware.
Even ignoring the licensing / patent issues with it, its non-journaled and only has a single FAT in most implementations; Its easily corrupted and difficult to repair. It also lacks a number of useful features like pre-allocation, robust meta-data, etc.