Why SATA is now obsolete
zdnet.com
zdnet.com
Because managing flash is hard. For this to work, the amount of exceptions in the filesystem it's self would make it unmanageable.
an example is this: a Fusion-io card up until recently used low quality flash, yet managed to get stella performance. Much higher performance than the equivalent intel ssd card (who have access to much higher quality flash) some of that is down to layout, the rest firmware.
Then you have the flash that lives on a device like your phone, however the firmware of the SoC is tweaked to run that particular flash controller.
At its heart, flash/hdd is just a key:value database (presented with 4k value size) why would we want to complicate that (sure we need to indicate to the hardware that we've deleted a page, but to be honest not much else)
nowadays all of the things (ha) use flash behind higher level of abstraction, be it emmc, UFS, or SD controller. There are no phone soc firmware tweaks.
To avoid an inherent bottleneck?. Perhaps we'd be better served by a larger number of key:value stores with greater parallelism?
It is interesting to note the design trend in mobile devices like phones.
Years ago, it was typical to have a NAND flash controller built into the SoC (like a TI OMAP, Freescale i.MX series or similar).
This is raw Flash memory, and it is up to the SoC plus OS to manage error recovery, remapping bad sectors, etc.
However, in recent years, most mobile devices just use one of their SD interfaces (often 8-bit these days) to access an eMMC chip. This looks just like a SD card, because it has a FTL in it which takes care of a lot of the low level details needed for Flash management.
Some SoCs these days don't even include a NAND Flash controller anymore.
[1]: http://lwn.net/Articles/518988/ [2]: http://lwn.net/Articles/518718/
"... the device still has its freedom to control block-allocation decisions, enabling it to execute critical tasks such as garbage collection and wear leveling ... Next, we present Nameless Writes, a new device interface that removes the need for indirection in flash-based SSDs. Nameless writes allow the device to choose the location of a write; only then is the client informed of the name (i.e., address) where the block now resides."
So it sounds like this approach is actually removing responsibility from the filesystem, not the firmware.
[1] http://research.cs.wisc.edu/adsl/Publications/yiying-thesis1...
They add significant latency to write operations, though, which is a high price to pay for such a small gain.
This abstractions is what have allowed us to transistion from a traditional spinning data store to SSD without much effort (save for delete flags to help the device performance GC and improve performance).
to use your example: imagine if you could use the fusion-io FS on the intel flash.
also remember that several SSD companies that even had their own flash fab are now gone because their firmware was crap.
The obvious way to implement this is to expose the raw flash to the OS and put an abstraction layer into the OS that does what the drive firmware currently does so it can be used with existing filesystems. Just that would have its own benefits because it would remove the black box from around that code and let people improve it. It would also make the drives less complicated/expensive and remove the burden from smaller manufacturers to maintain solid firmware (which they've often failed to do), and make it a lot easier and more reliable to secure erase drives.
Once you have that you can start looking at improving particular filesystems by taking advantage of the additional information now available to the OS.
It may or may not work better than what we have now. Doing so would require a lot more knowledge for the developers involved. Actually "developer" is a fancy term for a job that could not be called "engineering" for being almost completely insulated from that class of problems that an engineering job has to consider as a norm. Expose the raw flash (or other raw access to electrical components or what not) and you get yourself with real engineering problems on your hand. No more (software) "development", but engineering! And engineering is hard.
And there's one more thing - if we think the current state of software fragmentation is bad, wait until when this unified physical computing interface (that the firmwares more or less adhere to and which we're taking for granted today) is taken out.
It isn't about removing abstraction layers, only moving them. It should be part of the OS rather than the hardware. The people who write the abstraction layer have to deal with hard problems, but they do that now. What does it matter if they work for Microsoft and RedHat instead of Samsung and Intel?
The point is to publish and standardize how that abstraction layer works so the people who write filesystems and filesystem tools have better information and can suggest or provide improvements. And to stop forcing every SDD manufacturer to duplicate the software engineering efforts of the others instead of focusing on hardware.
The actual memory contents behind the controller are dependent on physical characteristics. Using the FTL firmware, the flash+controller (such as eMMC) vendor is free to do all kinds of tricks, for instance, depending on the quality of a particular batch of NAND. Bad batch? Use more of the spare for error correcting code. Particular memory pattern that generates interference? Tweak the scrambler. Slow? Interleave between a couple of NAND chips. (These examples are hypothetical)
Tying filesystems to the physical layer makes no sense. It would mean I wouldn't be able to use 'dd' to copy a partition to some other device, since it would have different physical characteristics which the filesystem would need to take into account. It would mean that I wouldn't be able to take an iSCSI volume and write it out to disk to 'de-virtualize' virtualized storage.
"Flash-based solid-state drives (SSDs) have revolutionized storage with their high performance. Modern flash-based SSDs virtualize their physical resources with indirection to provide the traditional block interface and hide their internal operations and structures. When using a file system on top of a flash-based SSD, the device indirection layer becomes redundant. Moreover, such indirection comes with a cost both in memory space and in performance. Given that flash-based devices are likely to continue to grow in their sizes and in their markets, we are faced with a terrific challenge: How can we remove the excess indirection and its cost in flash-based SSDs?
We propose the technique of de-indirection to remove the indirection in flashbased SSDs. With de-indirection, the need for device address mappings is removed and physical addresses are stored directly in file system metadata. By doing so the need for large and costly indirect tables is removed, while the device still has its freedom to control block-allocation decisions, enabling it to execute critical tasks such as garbage collection and wear leveling.
In this dissertation, we first discuss our efforts to build an accurate SSD emulator. The emulator works as a Linux pseudo block device and can be used to run real system workloads. The major challenge we found in building the SSD emulator is to accurately model SSDs with parallel planes. We leveraged several techniques to reduce the computational overhead of the emulator. Our evaluation results show that the emulator can accurately model important metrics for common types of SSDs, which is sufficient for the evaluation of various designs in this dissertation and in SSD-related research.
Next, we present Nameless Writes, a new device interface that removes the need for indirection in flash-based SSDs. Nameless writes allow the device to choose the location of a write; only then is the client informed of the name (i.e., address) where the block now resides. We demonstrate the effectiveness of nameless writes by porting the Linux ext3 file system to use an emulated nameless-writing device and show that doing so both reduces space and time overheads, thus making for simpler, less costly, and higher-performance SSD-based storage.
We then describe our efforts to implement nameless writes on real hardware. Most research on flash-based SSDs including our initial evaluation of nameless writes rely on simulation or emulation. However, nameless writes require fundamental changes in the internal workings of the device, its interface to the host operating system, and the host OS. Without implementation in real devices, it can be difficult to judge the true benefit of the nameless writes design. Using the OpenSSD Jasmine board, we develop a prototype of the Nameless Write SSD. While the flash-translation layer changes were straightforward, we discovered unexpected complexities in implementing extensions to the storage interface.
Finally, we discuss a new solution to perform de-indirection, the File System De-Virtualizer (FSDV), which can dynamically remove the cost of indirection in flash-based SSDs. FSDV is a light-weight tool that de-virtualizes data by changing file system pointers to use device physical addresses. Our evaluation results show that FSDV can dynamically reduce indirection mapping table space with only small performance overhead. We also demonstrate that with our design of FSDV, the changes needed in file system, flash devices, and device interface are small."
The question is, what non-file abstraction should a flash device offer? Slow persistent RAM? A key/value store? An SQL database? Computing has had "RAM" and "disks" as its storage models for so long that those concepts are nailed into software at a very low level.
The NoSQL crowd could probably find uses for devices that implemented big, fast key/value stores.
It can essentially do three operations: read block, write block, erase bunch of blocks at once. For the purposes of filesystem, this is actually an workable storage interface that does not need any additional abstraction layers.
I'm not an architecture expert, so please correct me if I'm wrong about this.
Though I suppose a byte is still not a bit! I'm not sure I've ever seen a bus interface that addressed individual bits…
Certainly not everything out there is byte addressable, but neither can one say unequivocally that all RAMs have large atomic word sizes.
Storage presents block devices to operating system. Spinning Rust and all manner of SSD use lots of clever engineering to appear as a block device.
Block device is the most basic abstraction of all storage. You really want MySQL to manage NAND? ;)
But if you think about it for even a little bit you realize that you can't just get rid of the abstraction, you're just moving it out of the drive's firmware and into the file system. And that gets you nothing because now the file system has to be compatible with every manufacturer's NAND drives. And the file system has to be updated any time that system needs to be changed. You can't just pop a new NAND drive in and have it work, you have to download new NAND drivers for the file system, etc.
I'm also skeptical of just how much performance gain you'd see. The article doesn't include these figures, which makes me skeptical that they are significant.
Delegation of duties is important for good system design. Drives necessarily need to present themselves in a standard way to the system. Having SSDs appear like normal drives is fine. They aren't memory and they shouldn't be used as memory.
Moving it from firmware to file system does not improve performance by itself, but it enables optimizations. And FTLs in firmware are far from optimal.
There are perhaps 10 more address spaces and translations. It the way OS and hardware are build.
> Why not have the file system directly manage the flash?
Too slow innovation, it takes decade to replace operating system. Also main CPU is very power hungry and slow...
More likely the SSD controller will become part of CPU and SOC.
Hear hear! The basic storage model that arose in the 50's and 60's being the best fit for recently developed hardware seems as likely as the security models from the 50's and 60's being the best fit for the Internet.
(If the obnoxious popup adverts from ZDNet strike me as archaic, they are not doing well.)
In fact in the thesis' abstract, the proposed solution is still a layer of indirection. Instead of the FTL giving an abstracted address space, the "device" is given the responsibility of controlling wear, gc, and parallelism by dynamically reallocating blocks itself, and the filesystem picking up some more of the management stuff.
I am certainly not qualified enough to read this dissertation critically enough to know if it is fundamentally sound theoretically, but as a scientist I know that it is close to meaningless until it has been experimentally tested and reviewed..
[1] http://forum.crucial.com/t5/Crucial-SSDs/Why-do-i-need-AHCI-...
I liked the idea the first time around so I like it again. But it's harder than it sounds.
I find something lovely and poetic in the fact that Lisp is over 50 years old and Multics is almost that old, and both may be having a bit of a renaissance.
"hard disk" will not be the ram. its just is made of the same chips, but still logically separated. and there is no need for a sata controller to address it really.
That's all. You still need a file system if you want to address the files. Just like address blocks of ram or anything else. If you give free reign to the app, there is no security, reliability, etc. Even ram is actually indexed as well.
On top of that you need a representation for the human behind it - no matter how hard they try to remove that (because its much easier to lock you in when there is no generic way to browser a memory device), its not very convenient compared to a file system.
the "nameless writes" idea is OK - a FS person would just call it even lazier allocation, not anything really different.
The article is of course correct, and some manufacturers are already offering SSD-like storage which is connected like a stick of RAM.
I'm sure the gap between storage and the main system pipeline (CPU-GPU-RAM-etc) will only shrink as time moves forward. However as an interim solution things that acted like hard drives were convenient.
"All problems in computer science can be solved by another level of indirection, except for the problem of too many layers of indirection."
Hence this level of indirection, the discrete "drive", is what he thinks should go poof.
First time I've heard that quote, but I love it. Reminds me of this one:
"There are only two hard problems in Computer Science: cache invalidation, naming things, and off-by-one errors."
Binary: as easy as 1, 10, 11
Also the commas and binary mixed together are a little odd. Least of all because you didn't make all three two binary digits (i.e. 11, 10, 01).
Regardless that joke is going to cause more arguments than laughs.
You do not "read" binary from right to left. At least in English, the most significant digit is on the left, just like any other base (e.g. base 10, base 8, or base 89432890432).
Also, for your leading 0, I would write: Bob has 23 apples, and I have 5. I wouldn't write: Bob has 23 apples, and I have 05. Base 2 is not special, so I wouldn't write 01.
You can also turn your order argument on its head: since the order of digits in base 10 is exactly the same as base 2 (most significant on the left), you could say that "As easy as 1, 2, 3" should actually be read as "As easy as 3, 2, 1", which matches your binary reading "As easy as 3, 2, 1".
There are only 10 hard problems in Computer Science: cache invalidation, naming things, and off-by-one errors.
There are 10 kind of people in the world: Those who know binary, those who don't, and those who start counting from zero.
That kills the joke. 'off-by-one' is subtler.
There are 10 kinds of people in the world: Those who know hexadecimal, and F the rest...
(I think it's funny though.)
- What goes "Pieces of seven! Pieces of seven!"
- Parity error
Look up "recursion" in the dictionary and it says: See Recursion.
You know why Complex Jokes aren't funny? Because the joke part is imaginary.
What's the best thing about telling UDP jokes? if don't You they get care it.
Don't forget that UDP allows repeats and drops!
A: Yes.
Turns out, the second one is originally from one Phil Karlton (as reported by Tim Bray [1] , that's enough authority for me), without the "off-by-one" part.
The first one, which I was goddamned sure was Dijkstra, is actually by David (not A.) Wheeler (via Butler Lampson) [2], inventor of the subroutine.
That rabbit-hole took me longer than expected. Hopefully I'm saving some time for someone else ;)
[1] https://twitter.com/timbray/status/506146595650699264
[2] http://www.dmst.aueb.gr/dds/pubs/inbook/beautiful_code/html/...
"All problems in computer science can be solved by another level of indirection"
"But that usually will create another problem."
-- David Wheeler
"...except for the problem of too many layers of indirection."
-- Kevlin Henney[0] http://en.wikipedia.org/wiki/NVDIMM
We move ever closer to "Computronium".
I'm sure someone will figure something out but pretending it's an HD does kind of solve that issue for now.
A pity we didn't see this until now, but we've adopted your suggestion.
All: you can get this kind of thing taken care of sooner by notifying us at hn@ycombinator.com.