Intel Storage Performance Development Kit
github.com
github.com
Yeah, that's great if all you're running is a benchmark. As soon as you need to combine this with a network stack, also polling its own devices at fast as it can, it becomes a lot harder to avoid those context switches. If you're running an actual application it becomes even harder. Likewise if you have more devices than you have cores to spin waiting for them.
In an extremely latency-sensitive and resource-rich environment this kind of thing can yield great results, but otherwise it's almost a form of cheating. Yes, that's what I was accused of when I wrote network drivers at Dolphin and again at SiCortex that some people felt polled too much. Oddly enough, users felt that their CPUs should be running their applications, and were more interested in maximum throughput per cycle consumed. Hardware designers don't build interrupt-based interfaces just for fun, and it's worth remembering that Intel makes most of its money from selling CPUs. Think about it.
It's definitely not a drop-in replacement for a kernel storage stack in the general case, but rather an optimization for specific applications (e.g. storage appliances) that can be structured to take advantage of the polled/no-interrupts model.
You will then, of course, have contention in transitioning data between the two, but there are existing and useful models for eliding much of the locking load there.
Does it make sense for every application? Unlikely. Does it make sense for an application that runs on 10k nodes, and where moving all IO to poll mode user-space doubles the number of requests you can serve per unit of time? Saving 5k machines worth of capex buys a lot of engineering complexity.
And since we know that systems like this have been done, 2x should not be an unreasonable number. See for example https://www.usenix.org/system/files/conference/osdi14/osdi14... and https://www.cs.cmu.edu/~hl/papers/mica-nsdi2014.pdf
As I also said before, this kind of thing does have its place. It all depends on how many cycles you're likely to spin before you actually find anything to do, and whether you have another need for those cycles. If it's not many cycles because you really are pushing a lot of I/O, that's great. If you don't need those cycles for something else because you're an appliance and this is your only job, that's great too. If it's a lot of cycles and you do need those cycles for other things - which is the most common case - then burning lots of cycles busy-waiting for events that haven't happened yet only decreases real hardware utilization and increases either capex or time to completion.
If you have long-lived work that shouldn't block processing more packets, it would be typical to offload that to separate thread/process from the one doing the packet processing (e.g., for control plane work).
This Intel library means that the network and storage can use the same event loop. It will integrate beautifully with Omni-Path's user-level library.
Currently SPDK consists of a usermode NVMe (PCIe-attached SSD) driver. We will soon be releasing a usermode driver for the Intel I/OAT DMA engine (copy offload hardware) that is available on some server platforms.
And can this be combined with PCIe-attached accelerators (e.g. Xeon Phi or GPUs)?
I am not familiar enough with the Xeon Phi or GPU programming model to say for sure, but they could possibly be used to offload tasks like hashing/dedup or other storage-related functions.
Sorry, I was not referring to accelerating storage-related functions, I was wondering about efficient DMA copy from one PCIe device (Intel NVM storage) to another (Xeon Phi accelerator) which would be for useful many different functions, if the NVM storage device capacity is much larger than the accelerator device memory.
Some of the straightforward use cases would be inside network-attached storage appliances (ideally in conjunction with a user-mode network stack) or in a database (database systems already typically want to avoid any OS interference with storage access). In general, the NVMe driver can be dropped in fairly easily when existing code is using something like Linux AIO with O_DIRECT on a raw block device; the AIO programming model maps quite directly to the NVMe driver programming model (create a queue, submit I/Os, and poll for completions).
Intro here https://software.intel.com/en-us/articles/introduction-to-th...
I wish I had an NVMe device to play with this.