Intel's New Optane 905P Is the Fastest SSD
tomshardware.com
tomshardware.com
I guess that modern SSD are just so fast, that going faster is not that big deal. But if you want the best SSD for non-server workloads, Optane is a way to go.
Anyway, the big difference with Optane is that its performance is more similar to having more system memory. I have not benchmarked this, but I believe a system with just 8GB ram + Optane will run better than a system with 16GB ram and Samsung 970 pro for workloads using up to 16GB of ram. There is a reason Intel announced its memory drive technology for Optane DC4800x https://www.intel.com/content/www/us/en/software/intel-memor...
EDIT: adding a link to lwn.net article explaing the use of swap/memory overcommit https://lwn.net/Articles/704478/
No, not even close. PCIe is still way slower than DDR4. Once your working set grows beyond available RAM, the Optane SSD will be a better swap device than any flash-based SSD, but you still notice your system slowing down drastically from all the swap activity.
Intel's Memory Drive Technology is really intended for situations where you have multiple Optane SSDs, and a workload that wants a large amount of total RAM but seldom has a true working set larger than actual DRAM capacity.
Traversing the root complex can take upto 100 cpu cycles each way.
PS: DDR4-3200W in quad channel mode is crazy fast and should break 100GB/s though sill a long way from a 1080Ti 484 GB/s.
On the other hand, Optane does have a big advantage at QD=1, which is the point people are making here. Looking at NVMe as an indication of Optane performance is the wrong approach since it has very different performance characteristics.
https://img.purch.com/r/711x457/aHR0cDovL21lZGlhLmJlc3RvZm1p...
(of course you're not incorrect that games aren't really using access patterns that are optimal for superfast SSDs)
Typical database load is an example of big queue depths. You have dozens of queries executing simultaneously and you're interested in throughput. So Samsung SSD is fine there (and raids are even better), unless your database is not typical and serves only a single client with low latency requirement.
Also because that was a bottleneck. Now the CPU is often the bottleneck, so I doubt you'd notice Optane (most people can't even notice NVMe vs. SATA).
Streaming video from the internet does not use the SSD at all, and a high-quality 1080p video file is maybe 5-10 MB/s of bitrate, you can easily pull that off a spinning HDD that was manufactured 20 years ago.
Video editing at 4K or 8K is one of the few use-cases where NVMe's sequential performance does provide a big benefit... assuming you are not editing using proxies.
Optane's QD=1 random-4K performance does present an opportunity for big speedups on consumer use-cases. But Intel really has to get the prices down if they want to see consumer adoption, right now there is an obvious benefit to cheaper SATA SSDs that allow you to get more data off spinning-rust drives vs a smaller, massively expensive Optane drive (even if it is incredibly fast).
Also, on PS3 it runs in less than 256 MB of RAM... where is it loading these assets to? Most games have RAM consumption on the order of 4-8 GB in most situations, not all of that is assets, and not all of that is read sequentially.
Funny enough, I skipped over SSD entirely, and upgraded to Optane from HDD.
It was like trading in a Model T for a Model X.
[1] https://www.treehugger.com/gadgets/plasma-tvs-suck-electrici...
Watt is a measure of energy/time. It doesn't make sense to say watts per hour.
I believe this is what the GP was referring to - the sustained usage of 300-600W over the course of an hour. Thus being very expensive to run for long periods of time.
In my area I’m charged at 16.56 pence per kilowatt hour, meaning at an average consumption rate of 450w - that television would cost me in the region of £13.41 a month to run, assuming 6hours usage per night. Which is quite expensive.
Perhaps, but that doesn't change the fact that Plasma TVs are generally more expensive to run than modern LCD or OLED TVs. [1]
Pioneer's 9th generation Kuro KRP-600A had an operating power consumption of 478 watts. [2]
[1] https://www.rtings.com/tv/learn/led-oled-power-consumption-a...
That isn't intel only either[0], although IIRC their hardware implementation is intel only.
Having the ability to handle a large volume of random writes, relatively cheaply, makes building reliable, fully restartable software much simpler.
Many storage engines that I make use of are B trees or LSMs these days. These engines are usually selected because they provide very good read performance, and acceptable write performance.
Improved random access writes only make these engines more attractive and performant. For instance, boltdb, which is similar to LMDB, requires 2 iops to perform a durable write. In the benchmark for 4K random writes, Optane achieved 180,000 iops. This gives us 90,000 writes per second, sustained. Possibly more at peak, since presumably that's how they get the claimed write iops of 550,000.
That's a phenomenal number of writes. It makes in-memory stores like redis irrelevant for a wide range of applications that would have required it not long ago.
Full disclosure: I work for Redis Labs.
My biggest issue is that it's only a 512, the biggest available at time of purchase (mid 2016; I only replace my systems every 5-6 years or so; 6850k/32gb mem was my limit then).
Also, doing benchmarking stresses loads up the thermals and when it hits 70c after a little while perf is throttled down somewhat; this doesn't happen in normal use, but nevertheless I've ordered a NVME heatsink which I hope will be good for it.
I see the Optane is a PCIE solution. Seeing as my m2 slot is filled, and noone makes u2 drives (the only other spare slot for such things my motherboard has), that'd make it the perfect upgrade extension for myself if I can fit it in between the GPU and SB card - hopefully it'll come down in price a bit over time.
Here is a real world figure. I can compile my FreeBSD NanoBSD build in 30 minutes compared to 2.5 hours on SATA3 SSD's
Using a Toshiba NVMe XG3 mounted on a paddle card.
I see zero heat issues with the Toshiba.
The NVDIMM stuff intel is working on helps alot with latencies, but, as others have noticed in this thread, NAND bandwidth is comparable to crosspoint in practice. So, if you are transferring many KB at a time, you don’t see the 1000-100x in practice. This breaks the performance assumptions made by all existing storage systems, so nothing but research prototypes are seeing anything close to 100x.
Most of these points are moot for NVMe drives, because PCIe latencies are so high. The reviewed model, gets about 10x better latencies, for 4x the price, and that’s still significant.
The only Optane-related thing that requires a particular CPU platform is their Optane Memory caching software for Windows. That's the very bottom of the Optane product stack, while the 905P is the top.
Intel's Optane Memory caching for Windows is built on top of their consumer platform's NVMe software RAID functionality. Intel's consumer chipsets have a mode where they hide any NVMe devices connected through the chipset, making them no longer appear in regular PCIe device enumeration. Instead, the drives can be accessed through non-standard interfaces on the chipset's SATA controller. This remapping ensures that only Intel's software RAID drivers can bind to the drives, which makes things a bit simpler for Intel.
This NVMe remapping feature was first added to Intel's Skylake generation chipsets, and Optane Memory was introduced with the next generation (Kaby Lake). Skylake systems didn't get the firmware updates necessary to add Optane Memory caching support (for booting off a cached volume) even though they had all the hardware capabilities and their firmware already had NVMe RAID boot support.
All Optane products released so far are standard NVMe SSDs and can be treated as such. If you want to use SSD caching software or software RAID from somebody other than Intel, it won't care whether or not your drive is an Optane product.
While it can clearly vary widely depending upon the nature of what you are building, it can in some cases be worth using faster storage to shave many tens of minutes off the build time.
https://www.joelonsoftware.com/2009/03/27/solid-state-disks/
For DRAM-equivalent latency and endurance you need STT-MRAM, which has already been available in DDR3-compatible DIMMs. Both STT-MRAM and 3D Xpoint will be available in DDR4-compatible DIMMS too but only STT-MRAM will run fast enough and long enough to replace DRAM.
Samsung has bigger issues with writes though, almost all of their current drives are TLC NAND which has extremely poor write performance so while short bursty writes can be fast (due to there being a chunk of SLC NAND cells as a write buffer) sustained writes are terribly slow. Optane doesn't typically meet the same peak write performance but it will beat most (all?) TLC drives on the market in sustained workloads.
Anyone using more advanced configurations (e.g. RAID) for these type of devices and what/how..
e.g. how to take advantage of things like this for large RDBMS installs which require N>1 devices for capacity alone (if not more throughput)...
Like: Taking GPU/PCI-slot optimized server and threw 8 of this style of device in there as a soft RAID50, etc.
would love to perform some experiments, but I don't have a spare $20k lying around 'for fun', so real world battle tales would be great...
I don't believe anyone has actually made one of those. MiniPCIe is a different form factor than M.2; the latter is newer and has almost completely replaced miniPCIe. MiniPCIe only provides a single lane of PCI Express, while most M.2 variants allow for two or four lanes, and almost all NVMe SSDs support at least two lanes.
Thanks for catching that! :-)
https://www.win-raid.com/t871f50-Guide-How-to-get-full-NVMe-...
I doubt that's true in practice, especially on a desktop workload.
It would be nice if there was a resident manager that watched for certain processes to be launched and then pinned their directories (or certain files) into memory, then unpinned them when they were shut down.
https://channel9.msdn.com/Shows/The-Defrag-Show/Defrag-Disab...
I'm thinking a shared resource attached to a hadoop cluster with R & Python workloads across 100's Gb's of data - so offload to Hadoop for embarrassingly parallel over 20+ TB but go as quick as you can for large somewhat parallel or sequential loads. About 10 users...
We have a couple of GPU servers for DNN's so that's a separate workload for us.
Start with a fat machine and see how far you can go: https://www.hetzner.com/dedicated-rootserver/matrix-ax
You don't need Hadoop if you crunch through 100GB.
My impression is that Microsoft has been falling down on the job for some time. in the '90s and aughts a computer felt slow after three years, and was nigh unusable after five, even for just word processing and web browsing.
These days? the almost new MacBook I'm typing this on only comes with 8gb ram; that was a decent (but not great) loadout in 2011. And my thinkpad? the other laptop I use? It is from 2011, also has 8gb ram, and runs just fine. As far as I can tell, the big difference between my ancient laptop and my new one is that my new one is way thinner (and has a keyboard that is dramatically more vulnerable to foreign matter)
I mean, I do own an occulus rift, and it does very much require a modern computer to run, but as far as I can tell, that's pretty niche. Hardware requirements just aren't going up the way they used to.
http://www.storagereview.com/corsair_vengeance_ddr3_ram_disk...