Caching Beyond RAM: The Case for NVMe
memcached.org
memcached.org
Dormando mentions the test was done with the help of Accelerate With Optane, a collaboration we have with Packet to provide access to servers with Intel Optane SSDs. Check out https://www.acceleratewithoptane.com/ for more info, and you can find me and the Packet team over at slack.packet.net. We're especially interested in open source projects that want to test what they can do with the tech and are interested in sharing what they learned with the broader community. Thanks to Dormando for going first!
Assuming using 32GB dimms to use all six lanes (if Skylake)? This is something that has bit me a few times - most Skylake cpus have 6 lanes, so balanced is 192, 384, etc. Not 4 lanes like we are used to!
Would also recommend trying different ram configs anyways, on our side we have seen better thru put on 768 than on 384. Even 512 performs better in some cases than 192.
Edit: Missed this the first time through:
> Calculating this is done by monitoring an SSD's tolernace of "Drive Writes Per Day". If a 1TB device could survive 5 years with 2TB of writes per 24 hours it has a tolerance of 2 DWPD. Optane has a high tolerance at 30DWPD, while a high end flash drive is 3-6DWPD.
Notably, CLOCK keeps items that are accessed atleast once during a round while LRU will kick out the least accessed item.
The benefit of using CLOCK is that you don't have to maintain a list but only a ringbuffer. Removing an item for a CLOCK's buffer can be almost free if you use a single bit to indicate presence. A LRU will have to maintain some form of list, array or linked. In practise, LRU is expensive to implement while CLOCK is simple. CAR offers LRU performance with less complexity.
https://en.wikipedia.org/wiki/Cache_replacement_policies#Clo...
https://dbs.uni-leipzig.de/file/ARC.pdf
[p.s. there is also the matter of the (patterns in the) various trace runs. Does anyone know where these traces can be obtained?]
Of course there are more schemes than just LRU and ARC, and one can try to employ lock-free schemes more than I'm willing to do. This is just my experience.
You can mitigate the exclusive lock using a write-ahead log approach [1] [2]. Then you record events into ring buffers, replay in batches, and have an exclusive tryLock. This works really well in practice and lets you do much more complex policy work without much less worry about concurrency.
[1] http://highscalability.com/blog/2016/1/25/design-of-a-modern...
[2] http://web.cse.ohio-state.edu/hpcs/WWW/HTML/publications/pap...
The whole thing is a mess.
I imagine there are quite a couple of variants of ARC you can use without violating the patent.
Hard Drives fail because of vibration, broken motors, and things like that. MTBF is the typical metric for hard drives. There are also errors that pop up if data sits still too long (on the order of years), because the magnetic field loses its charge over time.
TL;DR: there's a lot to it and I'll be going into it in future posts. The full extstore docs explain in a lot of detail too.
If the same can be done with Optane SSDs, the lower latency will at higher queue depth will certainly help.
I assume you don't need licenses for just making a dumb PCIe card, if you don't name it with trademarks? Or are there patents you need to license to sell PCIe-compatible, non-electronic cards?
The relative scarcity of NVMe ports/bandwidth per server may make that as unattractive as doing the (RAM) caching on the database server itself, but it's not obvious, if one could only spend the money in one place, where it would be best spent.
Historically, I've had trouble convincing management to get anywhere close to reaching the capacity of a server's PCIe lanes, possibly because of a general ignorance around direct attached (but external) storage options, but the popularity of NVMe may change that.
Now the challenge may become convincing people that 128 lanes per CPU is better than 40.
There's also the (extension of the perpetual challenge) that the mere existence of such a price-competitive option to Intel means one doesn't have to panic at the slightest hint of scale and invest immense engineering effort and non-linear infastructure cost into a distributed database to replace a "single"/central [1] RDMBS.
[1] With all the usual replicas and caches for performance (and failover), as is the point of the article.
it ensures that a lot of operations can't touch secondary (like miss, touch, delete, sets of new items, etc), which reduces load on the IO by quite a lot.
edit: Also extstore itself will support NVM + SSD sort-of-layers soon enough. I'll be retesting that on the same optane+ssd machine in a couple weeks.
It has already been established that for most consumer workloads, the latency differences between memory and Optane is negligible. This article shows that heavy-duty workloads (ie: high-traffic memcached clusters) can be accommodated by Optane and NVMe too. Clusters of 500K IOPS drives can take us most of the way there.
I don't want to get all /r/hailcorporate but Optane drives are great products and (more importantly) you can run them on AMD platforms (ie: SP3) too. Granted, NVMe drives bring the fight and they're much more cost competitive at scale, but that will change soon.
Wait, you can? I thought Optane required specific motherboards for the DIMM versions? From what I understand, the M.2/nvme devices just show up as a drive in Linux right?
But I thought the actual DDR4/DIMMS only work in certain Xeon boards.
[0] https://www.intel.com/content/www/us/en/products/memory-stor...
A dollar per gigabyte doesn't seem half bad at all for the top end of performance.
There's always going to be the top tier storage that costs an arm and a leg - that just means two steps down gets affordable. Optane pricing will drive down NVME which will drive down SATA SSD.
Why pay 2x more for NVMe NAND if your video game load times aren't any better?
My first HD in 1991 cost like $100 and was 40MB. So, there's that.
Besides "a dollar per gigabyte" is $1000 for 1TB.
512MB SSDs used to cost ~ $1000 just 5-8 years back...
Of course, as with previous SSD tech, the price will come down fairly rapidly.