Coding for SSDs
codecapsule.com
codecapsule.com
1. Most of the recommendations are things that will have little to no benefit on any modern SSD because all of this is handled for you in the firmware. When SSD firmware is not perfectly optimizing your performance, it is doing so to extend the life of the drive. Take this with a grain of salt though since many SSD manufacturers use really bad code provided by the controller manufacturer (SandForce notably provided awful stock firmware).
2. SSD benchmarks commentary is mostly accurate - since performance is very motherboard chipset-dependent, we chose the best ones to publish. Can't really blame anyone, but it means that if you bought the drive you are incredibly unlikely to achieve identical performance.
3. "Benchmarks are hard". No, not really. I wrote most of the benchmarks we used for linux testing. The reason benchmarks could be considered difficult is because all of the internal benchmarks that are used were written in-house and not released to the public. Reviewers never really had the right tools.
4. Enterprise SSDs and Consumer SSDs are drastically different. If you want consistent high performance, get an enterprise PCI-X drive. NAND quality is the major difference... and good NAND is hard to find. Generally an enterprise SSD will also have a significantly faster controller, and on larger drives, will contain more than one controller to avoid straining resources.
Where do we find real advice about how to best utilize SSD?
I work in an enterprise setting and get to talk to the ssd vendors and fish for details on how to best use their ssds and I still find it hard to get full information. Sometimes I wonder if there can even be a single compendium of knowledge about ssds. There are just so many moving parts in there that I doubt the ssd vendors themselves know what is really happening. Many times there are strange behaviors that we hit and then the vendors get to debug their ssds and explain new things about what just happened.
One effect that I can attest to is that I know quite a bit about the ssd inner-workings (or at least how it shows to the outside world, never saw an ssd firmware source line) but it would be very hard for me to sit down and write something about all aspects. I can however answer questions when needed.
I found this amusing in the context of SSDs versus spinning rust. :-D
If you really want to be pedantic, there are electrons moving in there so it is solid state but it has moving parts :-)
The high-level optimizations are important: * Choose a good SSD * Read and Write in "page" multiples and "page" aligned * Use lots of parallel IOs (high queue depth) * Do not put unrelated data in the same "page"
A page used to be 4KB, SSDs are now switching to 8KB and will switch to 16KB later on. Just pick a reasonable size around that (16KB if you can do it will last you a while). Don't sweat the page multiples too much, the SSDs will most likely have to handle 4KB pages for a long while due to databases and such so they will keep some optimization around that size anyhow, it will make it easier for them if you use a larger size.
I wouldn't heed any of the advice on single-threading, the biggest performance boost comes from parallelism and writes are anyway buffered by the SSD (a good SSD has a super-cap to have a good sized write cache).
Totally agree with your view of the advice given.
For example, TRIM has surprisingly little effect because modern SSDs (Micron) will move cold data to worn cells so that all cells participate in wear leveling. TRIM still saves you some, just nowhere near what you might expect.
It's nice to know that all that work on high-performance IO in databases doesn't need to be thrown away just yet.
But then you need to handle stuff like wear leveling, transaction management, ECC and other forms of recovery. And a lot of the stuff you need to do is probably flash-part specific (e.g., read disturbance, probably stuff around channel management and throughput, etc.).
I actually proposed allowing the firmware of a recent consumer product have such access to the flash (because I didn't trust the flash vendor's translation layer), but got shot down. I don't know how that turned out; probably they spent a bunch of time doing qualification (code for: "Fix your damned FTL bugs or we find another vendor. Wait. We don't have time for that. Fix as many as you can, or we'll be mad ... or something. Here, have some money.").
i'm sure that for any given controller's FTL (and this article claims that there really are only a couple on the market), i could tweak my algorithm to work reasonably well. but that's a sign of a leaky abstraction
i'd also like access to the small SLC portion of the drive, though i'm working around that for now with journaling
i'm not an expert in flash memory. my model is basically a block device with larger block erasure, and that the number of erasures each block can handle is limited
Everything on a computer approximates a block device. Even RAM is treated as a type of complex block device in sophisticated, high-performance databases, because it is. Scheduling operations to optimally match the characteristics of block devices is a (the?) primary optimization mechanism.
It all depends though if you want throughput or latency. If you really really care about latency you should balance the queue depth and probably use it between 16 and 32. The reason is that with higher queues you get more collisions on the same die and then latency suffers. There are read-read, write-read, erase-write and all the other combinations but those three are the interesting ones.
There's a lot of complexity going on behind the scenes of modern SSDs. Simplistic benchmarks as posted on most review sites often don't address real-world performance. The ones I pay most attention to are things like BootRacer; it's quite remarkable how often a drive will be slower at booting Windows than a competitor, even when it beats them in simple metrics like sequential / random reads / writes with different queue depths.
I can see the potential for firmware being optimized for modalities that crop up often in benchmarks but less often elsewhere: the aforementioned sequential / random reads and writes with different block sizes and queue depths. Doing well on benchmarks probably drives a bunch of sales. But detecting and switching modality may not be latency-free, and shortcuts taken to improve absolute performance in those modalities may harm global performance in a real-world IO mix.
http://www.anandtech.com/show/9144/crucial-bx100-120gb-250gb...
I was even told that an SSD was made to work consistently over the life span but drop some of these consistency work (i.e. slow downs at start of life) and only kick them in after so many power-on-hours so that it will not look bad in benchmarks.
It's a fact of life that customers rely on benchmarks and that vendors cannot educate all customers so the masses out there need to see a good benchmark so that's what gets optimized for.
sigh