NVMe, the fast future for SSDs
pcworld.com
pcworld.com
6 Gbps, 600 MBps (the capital B is supposed to indicate 'Bytes' versus 'bits') the encoding is 8b/10b which is 10 bauds per 8 bit byte.
PCIe 2.0 has 2.5Gbps "lanes" PCIe 3.0 has 5Gbps "lanes" they can be ganged together for additional bandwidth. (x1, x2, x4, x8, x16) it is also 8b/10b so you divide by 10 to get Bytes per second (250MBps/500MBps).
Both SATA and PCIe have a 'transaction limit' which is a function of the controller, which limits the total number of operations per second (IOPs). The product of the IOPs and the size of the transaction can never exceed the bandwidth of the channel. But it often is under. For example a typical SATA disk control (prior to the popularity of SSDs) would do about 25,000 IOPs, and if you had 512 byte (.5K) block reads and writes, you could read and write 25,000 * .5 or 12,500K or 12 MBps (which was much lower than the theoretical bandwidth of 200MBps on 2Gbps SATA II channels. Optimizing channel utilization requires that you figure out how many IOPs your OS/Controller can initiate and then sizing the payload to consume the max bandwidth. Large payloads and you'll push IOPs down, smaller payloads and you won't use all the bandwidth.
One of the nicer aspects of ATM was that it was designed and specified for full channel utilization with 64 byte packets which made it possible to reason about the performance and latency of an arbitrary number of streams of data moving through it.
> For example a typical SATA disk control (prior to the
> popularity of SSDs) would do about 25,000 IOPs, and if
> you had 512 byte (.5K) block reads and writes, you
> could read and write 25,000 * .5 or 12,500K or 12 MBps
> (which was much lower than the theoretical bandwidth
> of 200MBps on 2Gbps SATA II channels.
That doesn't seem to jive with reality. What am I missing here? SATA HDDs would regularly hit 100MBps in sequential transfers.To pick a 2009-era HDD review/benchmark at random that illustrates this: http://www.storagereview.com/western_digital_scorpio_black_5...
So when characterizing a typical SATA drive you would start with 4K sequential reads and work up until your bandwidth hit either the channel bandwidth or stopped going up (which would be the disk bandwidth). Unless you ran across a reallocated sector many SATA drives could return data at a rate of 100MBps with 1MB reads. Or even smaller read sizes if you had command caching available. Random r/w was an issue of course because of head movement (burns your IOPs rate while waiting for the heads to change tracks)
You can do these experiments with iometer[1], there was a great paper out of CMU which talked about illuminating the inner workings of a drive by varying the workload[2]. Well worth playing with if you're ever trying to get the absolute most I/O out of a disk drive.
[2] http://repository.cmu.edu/cgi/viewcontent.cgi?article=1136&c...
"While a single SATA port is limited to 600Gbps, combining four makes for 2.4GBps of bandwidth."
600Gbps * 4 = 2400Gbps = 3GBps
Maybe he thinks a byte is 10 bits or something?
It's also really odd that he's using Bps at all. I've never seen MBps anywhere other than this article. Usually it's Mbps and MB/s.
going from "b" to "B" is either 8 or 10 fold. As some of the other comments have noted. SATA uses 10 bits per byte.
so 2400 Gbps = 300 GBps (8 bit)
or 2400 Gbps = 240 GBps (10 bit)
I don't really know why it's become standard to publish raw bit rates in bit/s and data rates in bytes/s with error correction taken into account but not protocol overhead, but those are the two kinds of numbers you almost always see quoted nowadays. Raw bit rates at least map pretty directly to clock speed, and I guess protocol overhead must be too variable and too complicated for most people to bother explaining.
I'm currently replacing all our SANs with storage servers filled with NVMe SSDs (as the tier 1 storage, commodity SATA SSDs for second tier). I've posted the link to my first blog post about it in the comments on another NVMe post in the past: http://smcleod.net/building-a-high-performance-ssd-san/
I'm close to writing the next post around the actual build, my findings, benchmarks etc... Hopefully I'll have that done next week - but the system comes first.
I'm a little disappointed with this article as I think it could do with a) some technical review and b) some more detailed information.
we need an interface that allows us to bypass the FTL and access the underlying erase blocks
[0] https://en.wikipedia.org/wiki/Log-structured_file_system
I'm not sure that I'd want the protocol to get involved in the intricacies of backing store housekeeping as you propose. Newer generations of SSDs may not even have erase blocks or translation layers; do we want to have yet another protocol when the technology changes?
The hiding the happens currently in the block interfaces (HDD, SSD & NVMe) definitely allow for easy integration and let things work pretty good for most cases but prevent getting full performance from the device and also obstruct diagnostics when things fail.
Is the api different, or are we still reading/writing disk files?
Should we do memory mapping of the files or not?
Should we parallelize access to different sections of big files? Or write a ton of small files?
How does this affect database design? Current big data apps emphasize large append-only writes and large sequential reads (think LSM trees). Does this make sense any more?
What does disk caching mean in the context of these new drives?
BTW, does anyone know if NVMe uses the ATA command set for side channel stuff like configuring encryption and the like?
Shameless plug: if you want to work with this stuff, check out purestorage.com/jobs
They also had RAM drives you could buy. The point then as now is that increasing the "high performance" working set space of a program, increases the amount of transactional data that can be "in flight" during an operation, and that increases the overall size of the data set you can work with.
I've been waiting for these boards to come down in price for about 4 years now. I started talking with Intel about them early on (we used their XM-25 SSDs because it was a price point for flash that was "enough" better than spinning rust that it made sense) and they insisted on trying to sell us the same flash chips on a PCIe card for 10x the dollars, I (and many others apparently) refused to pay that. Sure if you have a 'cost is no object' data base or something but for a large internet working set where revenue differences are measured in cents per thousand transactions? Not so much. I know one company that went so far as to design and build their own PCIe Flash card. I have heard it did great stuff for them.
[1] LIM - Lotus-Intel-Microsoft spec for extended memory on IBM PC compatible machines.
http://en.wikipedia.org/wiki/Expanded_memory vs
https://books.google.com/books?id=KjwEAAAAMBAJ&pg=PA61&lpg=P...
I was thinking of the plug-in expanded memory cards rather than the plug in hard drive cards.
In comparison, PCIe is a much more sophisticated serial protocol. It's an entire damned packet network with addresses, subnets (so to speak), routing, retry, a credit system for bandwidth sharing. Unlike ethernet, which typically tops out at achieving ~75% of its theoretical bandwidth, PCIe typically tops out at ~95% of its theoretical bandwidth. Crazy stuff.
I'd bet good money that the PCIe PHY can beat the crap out of the SATA PHY on energy/bit and that the disparity will only widen with time.
http://www.samsung.com/global/business/semiconductor/product... http://www.pcworld.com/article/2866912/samsungs-ludicrously-...
SATA drives connect to the host system over a SATA PHY link to a SATA HBA that itself is connected to the host via PCIe. The OS uses AHCI to talk to the HBA to pass ATA commands to the drive(s).
PCIe SSDs that don't use NVMe exist, and work by unifying the drive and the HBA. This removes the speed limitation of the SATA PHY, but doesn't change anything else. The OS can't even directly know that there's no SATA link behind the HBA; it can only observe the higher speeds and 1:1 mapping of HBAs to drives. Some PCIe SSDs have been implemented using a RAID HBA, so the speed limitation has been circumvented by having multiple SATA links internally, presented to the OS as a single drive.
NVMe standardizes a new protocol that operates over PCIe, where the HBA is permanently part of the drive, and there's a new command set to replace ATA. New drivers are needed, and NVMe removes many bottlenecks and limitations of the AHCI+ATA protocol stack.
PCIe is packet level, NVMe is a block device interface.
From [0]:
"The controller is the same 18-channel behemoth running at 400MHz that is found inside the SSD DC P3700. Nearly all client-grade controllers today are 8-channel designs, so with over twice the number of channels Intel has a clear NAND bandwidth advantage over the more client-oriented designs. That said, the controller is also much more power hungry and the 1.2TB SSD 750 consumes over 20W under load, so you won't be seeing an M.2 variant with this controller."
[0]: http://anandtech.com/show/9090/intel-ssd-750-pcie-ssd-review...
Is there something I'm missing?
Plus this would have the advantage of driving down prices for thunderbolt and increasing adoption.