Rest in Peace, Optane
specbranch.com
specbranch.com
You had single device (optane stick) which served as both - ram and disk.
After decades of current approach we finally could physically get rid of disks - hdds, ssds or even nvme sticks.
https://www.theregister.com/2022/08/01/optane_intel_cancella...
>Optane presented a radical, transformative technology but because of this legacy view, this technical debt, few in the industry realized just how radical Optane was. And so it bombed.
I'm no expert on costs of these things and especially wouldn't dare to predict how these costs could've behaved in the future, but my guess would be not that the technology had no use, rather that it was too expensive for what it was good at, and the second best was cheap enough to warrant choosing it over Optane.
>Intel Optane persistent memory is blurring the line between DRAM and persistent storage for in-memory computing. Unlike DRAM, Intel Optane persistent memory retains its data if power to the server is lost or the server reboots, but it still provides *near-DRAM performance*. In SAP HANA operation this results in tangible benefits
>Speeds up shutdowns, starts and restarts many times over – significantly reduce system downtime and lower operational costs
Process more data in real-time with increased memory capacity
Lower total cost of ownership by transforming the data storage hierarchy
Improve business continuity with persistent memory and fast data loads at startup
https://sp.ts.fujitsu.com/dmsp/Publications/public/wp-perfor...
If Intel really wanted to redesign the Von Neumann architecture, they would have had to be prepared to absorb losses for much longer, way north of a decade.
The alternative might have been to focus exclusively on providing SSDs using the new technology and maybe try to segue into this new memory architecture 10 years later. Like Itanium should have initially focused on beating competing x86_32 chips of the era in benchmarks and ship the new ISA as an afterthought.
Thank you, Intel, for trying to push the envelope, though.
People might have bought Optane if the pitch was “Postgres/MySQL runs twice as fast” rather than hoping someone else would make your purchase cost-effective later.
Anyone trying this needs to figure out a decade-long schedule with points where something would be worth using for some reason so they don’t have to run the whole thing in a vacuum hoping it’ll be worth it at the end.
And in the later Xeon architecture (Xeon Scalable 3rd gen, I think), intel expanded the persistence domain to CPU caches too. So, you didn't even have to bother with CLFLUSH and CLWB instructions to manually ensure that some cache lines (not 512B blocks, but 64B cache lines) get persisted. You could operate in the CPU cache and in the event of power loss, the CPU/mem controllers/and the capacitors on Optane DCPMMs ensured that the dirty cache lines got persisted to Optane before the CPU lights went off. But all this coolness a bit too late...
Another note: Intel's marketing had terrible naming for Optane stuff. Optane DCPMMs are the ones that go into DIMM slots and have all the cool features. Optane Memory SSDs (like Optane H10) are just NAND SSDs with some Optane cache in front of them. These are flash disks, installed in PCIe slots but Intel decided to call these disks "Optane Memory" ...
Yes, the real case for Optane memory is that, supposedly, you don't have to fsync(). And insisting on proper fsync() tends to tank the performance of even the fastest NVMe SSD's. So the argument for a real, transformative performance improvement is there.
Do you mean the latency of ensuring fsync safety is lower?
No they don't. A fence only imposes ordering. It's instant. It can increase the chance of a stall when it forbids certain optimizations, but it won't cause a stall by itself.
CLWB is a small flush, but as tanelpoder explained the more recent CPUs did not need CLWB.
It's better to design for unexpected restarts than design for a golden in-memory image which needs to be carefully ported around, have all its connections wired back up, and so on.
You're going to get unexpected restarts anyway. The faster and more reliable you can make recovery from that, it benefits you in the moving use case. The kinds of things you might want to do to enable reliable restart - like retry mechanisms for incoming requests - make migration work too.
Transient & disposable memory structures like keeping track who's logged in or compiled SQL execution plans that facilitate access to the persistent "business data", much of that stuff will need to be in RAM/HBM/CPU cache anyway, for performance reasons and as these things do not necessarily need to persist across a crash/reboot. The data (and likely indexes, etc) need to. But you won't need a buffer cache manager that copies entire blocks around from storage to different places in memory and vice versa. Your giant index or graph could rely just on direct memory pointers instead of physical disk block addresses that need to get read to somewhere in memory and then are accessed via various hashtable lookups & indirect pointers. And you don't have to ship entire 512B-8kB blocks around just to access the next index/graph pointer, just access only the relevant cache line, etc.
With proper design, you'd still have layers of code that take care of coherency, consistency and recovery...
It wold be a great use for that technology to decrease function startup time .
Consoles have been using unified memory since the 8th gen (PS4/XB1/Switch, kinda sorta even WiiU).
And NVidia has CUDA unified-memory slide decks going back at least 5 years.
That might be the promise, but that premise is fundamentally flawed. I mean, what drives the need to classify memory devices is performance characteristics, and even if your starting point is this idealized world where Optane was ubiquitous then all it took was a ephemeral memory technology to significantly outperform Optane to create need support the performance-oriented non-persistent memory type.
On one hand, you have a large market of people that are not sensitive enough to storage performance for the extra cost to be worth it. Ordinary storage is perfectly adequate for their needs and they would see little benefit in transparently making the storage their software was designed for faster.
On the other side, there are people who care about storage performance a lot. So much so, that they will design and modify their software to take full advantage of the storage and system characteristics. They use the storage so well that the gains by putting something in the middle will be marginal or non-existent, and certainly not worth the extra cost —- they would gain just as much by adding more storage devices.
The effective target market always ends up being “people who care a lot about storage performance but use their storage in the most naive way possible”, which isn’t that big of a market in practice. Using good software design to achieve similar results on commodity hardware is almost always the better option.
The transition from HDD to SSD technology is a counter example to your claim. It was a drop-in replacement (same SATA interface) and the tech significantly improved performance without any other software modification needed.
Octane was an incremental improvement in performance when compared to SSD, and in cost when compared to RAM. An awkward spot to be in.
Optane PM was byte-addressable and had latency of ~300 ns. It rendered the entire block storage abstraction obsolete. It was so radical that using it effectively would require throwing out the assumptions all OSes have made about IO for the past fifty years.
The fact that it offered an incremental improvement even when its unique capabilities were completely ignored shows how phenomenally capable the technology was.
For a superior but new tech to stick, it has to find a viable market to be first self-sustaining, only then can it attack larger and more lucrative markets (Back in those days, PC was probably more lucrative than mobile). Only when SSDs matured enough, did they start eating up the entire storage market.
The difference between Optane and an SSD was never so striking.
They came in 2.5 inch form factor from the very beginning, the laptop shape... which fits pretty much anything. Desktops and servers use it.
The packaging was trivial. They were just cost prohibitive for anything more than an operating system drive.
This worked out to our benefit, the OS is where random access performance shines
It might be fair to say flash used in SSDs started in other applications... but it's important to remember that's also different flash.
It wasn't until the flash changed significantly that we got SSDs (durability), changing everything we once knew. Price, applications, etc.
I don't think so at all. The value of Optane was that it was dramatically different, offering byte-addressable storage with latency of hundreds of nanoseconds. Getting the most out of this requires radically redesigning how modern systems store data, including ditching the idea of block storage.
Optane was the fastest traditional SSD on the market by a pretty wide margin, but that's not all it could have been.
>They use the storage so well that the gains by putting something in the middle will be marginal or non-existent, and certainly not worth the extra cost.
I don't agree with this either. How do you propose to get latency down to match Optane using ordinary SSDs? It doesn't matter how much hardware you throw at the problem; parallelism cannot reduce latency. You need another storage system to hit those targets. Maybe that can be a more traditional DRAM cache with battery backup, but that's only adequate in bursty workloads.
Putting "something in the middle" isn't worth it if the rest of the architecture stays the same. But it could bring about tremendous speedups if the new capabilities were used.
It says here an SLC drive has the same latency for sequential read (random read is 5x higher though). [2]
[1]https://www.intel.com/content/www/us/en/products/docs/memory...
[2]https://www.solidigm.com/products/data-center/d7/p5810.html
Obviously it failed as a product but I am not so sure that the cause is as you lay it out
It was a massive headache for me. One fine day my laptop ran into the common Windows issue of 100% disk utilization. I tried all the common fixes to no avail, and at some point I remembered my disk had some funky new tech called Optane. I disabled Optane through its software and was able to directly access the underlying HDD. I checked the fragmentation level for the HDD, lo and behold it was fragmented to oblivion.
Turns out because Windows treats Optane disks as SSDs even though I actually had an underlying HDD, my HDD was simply not defragmented by the OS. After a few rounds of installing and uninstalling large games, the HDD was in an unusable state with regards to fragmentation.
I did a short write-up PSA on r/Windows10, and apparently the issue was widespread enough that my post helped about 10 people in the comments. Thinking back, this whole series of events is partially the reason why I moved from being a non-technical person to a (somewhat) technical one. Good times.
Absolutely horrible product. And not even cheap before having to turf it out in favour of a proper SSD.
Edit - wrong prefix sorry!
This is a 905p vs SK Hynix Platinum for reference.
I lobbied hard for some samples, but they simply would not provide them. I think that if we were representative of how they acted this may be the reason they never got buyin. If we had had those parts maybe we would have implemented the POC's we were planning, and those POC's may have made it to production and the portfolio of the time. Perhaps there would have been a good market ready for 2015.
I remember sitting through those pitches and being like "Ok, you don't have any idea how and where to apply this new thing of yours, which is understandable, but then you are unwilling to let anybody have a shot at it. What I'm doing here then?".
IDK what they were thinking, but the general mood was "oh, another Itanic".
In the end optane was a “cool, nice to know and play around” product with a real market way smaller than Intel was prepared to support.
From a consumer perspective, it was a fast small SSD that doesn't wear out. That was enough to justify it's value at the prices it went for(Or at least it seems like it, I never personally used it). It didn't need any radical rethink of computing to be worth it.
I'm not even a fan of the nonvolatile RAM idea. Rebooting is often the first thing we do to fix stuff. Modern computing trends towards regenerating from descriptions, not adding more persistent state. Persistent state can get messed up, better to be able to wipe and start over.
So there was significant hurdle going beyond 16GB per machine, and sales pitch for Optane was DRAM substitute at near-NAND density and cost. Which, if taken at face value, implied >64GB RAM without committing to server/workstation platforms. That was appealing, at least to me back then.
Nowadays you can buy 4 sticks of 32GB DDR4 or DDR5 modules for ~$350, which makes a hypothetical 64GB Optane DDR3 module a rather moot point.
Depends on the environment. That used to be the common approach for windows, less so for *nix systems.
Outside of some specialized database applications, having an entire program running right in persistent memory seems like a bad idea.
And even then, I'm assuming databases would still want strong separation between persistent stuff and stuff that can be restarted any moment, to minimize the surface area for problems to stay persistent in.
You wouldn't have to necessarily worry about DB dumps since the backing storage would still be very fast but could survive reboots.
If I had a database that is smaller than 400GiB in size, I could have made it screaming fast while being safe with an Optane drive.
Indeed one of a kind technology.
Another issue may be that Optane DC persistent memory simply was not fast enough to replace non persistent RAM.
Still, I hope that another technology will arise, which is byte-addressable and persistent/durable. I think it could radically change the design of database systems again. You wouldn't have to have a page cache / buffer manager, which retrieves same sized blocks (or multiples of blocks) for instance. You probably wouldn't even need serialization/deserialization to disk.
It would be great if it'd be possible to for instance read 512byte blocks from disk with current SSDs, but I guess the block overhead for meta data might be too big.
PCIE 3.0 Optane recently sold so well it got backordered and the price popped. (Ok maybe it was LevelOneTechs Wendell? Maybe it was organic? https://www.newegg.com/intel-optane-905p-1-5tb/p/N82E1682016... ) Today PCIE 4.0 prices have fallen so fast that you can RAID0 a few M.2s and get about the same IOPs from a single Optane drive for a similar $ per GB, though not exactly the same durability / DWPD. Chia is no longer a driving force for drive durability, but then again Optane was designed for databases / SAN and not proof-of-space (e.g. PNY https://www.pny.com/lx3030-m2-nvme-ssd ).
Optane as a consumer product might have been a good play when laptops mostly had spinning disks, but market timing was too late. Today those consumer Optane drives were, yeah, a mistake, but a cool part for a homelab.
Optane as an enterprise / workstation product though is king for: * databases / ml datasets that don't fit in memory * large SANs that host critical VM persistent storage
It's a small niche but when IOPS and durability matter, there's nothing close. However if Optane really ends at PCIE 4.0, then top-tier PCIE 5.0 nvme will probably meet or beat Optane in 2024.
I can't speak to these specific devices because I haven't been doing infrastructure work since before 2015, but I recall the price of certain enterprise products would routinely skyrocket a few years after it became hard to find them. I recall having to source a replacement drive for a server at my Dad's company and discovering that finding any drive at least 4.5GB in capacity for this 15 year old server was challenging. I ended up having to buy a much larger drive at about twice the cost of a modern (at the time) enterprise-grade SAS drive.
I doubt my specific example is all that similar to what might be happening. In my case, it was a confluence of the usual suspects: RAID controllers that are very particular about the specific drive models/firmware versions that they work reliably with (which everyone else needs to replace their drives, too!) coupled with the server using a no-longer-common interface (SCA) and a drive tray specific to the vendor which was no longer a popular (and went Bankrupt)[0].
I suspect it's too soon for it to be that simple but I wonder if there are specific circumstances where "My Optane drive failed and replacing this drive at any cost is the only fix that gets me access to my data, again".
[0] Not important, but it was a Gateway brand server if you can believe it. And I ended up kludging the tray situation, anyway.
So as a result Optane didn't look on the marketing or even in typical reviews any different then competing NAND-based solutions, or heck would even look slower. And lore developed on the internet even amongst tech folks that it "only made a difference on servers". But I was fortunate enough to grab a few to use for core storage and wow is it noticeable, I've built lots of regular SSD big arrays and they look great at simple patterns with large blocks and then one gets into regular workloads and they absolutely tank, with occasional noticeable blips when garbage collecting or the like gets hit. Better than spinning rust overall, but surprisingly not by much sometimes. Whereas Optane is rock solid consistent no matter what.
Intel and Micron were really dumb in how they tried to use it and push it, but I think it's too bad as well it never really got much popular recognition in terms of differences with NAND. A lot focuses on "closer to RAM" but it was also a better SSD.
A different form factor are the NVMe sticks. Given the relatively small capacity, the 118GB NVMe SSDs were not very expensive, and make ideal system drives for server applications.
But brute force seems to have defeated Optane: Intel's enterprise flash NAND SSDs are just super over-provisioned, retaining gobs of spare capacity that result in 8TB devices with one complete drive write per day durability, every day for five years guaranteed.
This reminds me of running Unix on the Cray X-MP. It worked, but was very slow since calling a function stalled these superpipelined machines designed to run FORTRAN really fast.
Seymour Cray's CDC machines were fast computers, but IMHO those Cray Corp machines were essentially high performance vector units that happened also to be able to run a bit of control code. I guess it wasn't sexy enough to build a coprocessor, or maybe by then there was enough choice in mainframes that interconnect (which was far from standardized in those days) would have been a barrier.
I mean, I do understand (but disagree with) the specific use case at NASA where I encountered it: the CFD would run all night (or longer) and when done the files went to some Irises for visualization, which also took forever. So running the same source code (basically rsync iirc) sounds easier in theory. But at what cost? I think it would have been better to just write the transfer code to run under Cray OS.
And indeed, once unicos was on the machine people did want to run it interactively.
It means it would be possible to read/write data in much more fine granular chunks (potentially saving a lot of storage space in some cases).
Intel discontinued/deprecated Optane, before I could do anything really cool with it. But Intel can probably still reuse lots of their cache coherency logic for external CXL.mem device access.
One serious early adaptor and enterprise user of Optane tech was Oracle. More specifically Oracle's Exadata clusters (where database compute nodes are disaggregated from storage nodes that contained Optane) and connected via InfiniBand or RoCE. And since Optane is memory (memory addressable), they could skip OS involvement and some of the interrupt handling when doing RDMA ops directly to/from Optane memory located inside different nodes of the Exadata cluster. I think they could do 19 microsecond 8kB block reads from storage cells and WAL sync/commit times were also measurable in microseconds (if you didn't hit other bottlenecks). They could send concurrent RDMA write ops (to remote Optane memory) for WAL writes into multiple different storage nodes, so you could get very short commit times (of course when you need sync commits for disaster recovery, you'd have to pay higher latency).
With my Optane kit I tested out Oracle's Optane use in a single-node DB mode for local commit latency (measured in a handful of microseconds, where the actual log file write "I/O" writes were sometimes even sub-microsecond). But again, if you need sync commits across buildings/regions for disaster recovery and have to pay 1+ ms latency anyway, then the local commit latency advantage somewhat diminishes. I have written a couple of articles about this, if anyone cares: [1][2].
Fast forward a few years, Oracle is not selling Exadata clusters with PMEM anymore, but with XMEM, which is just RAM in remote cluster nodes and by using syscall-less RDMA ops, you can access it with even lower latency.
[1] - https://tanelpoder.com/posts/testing-oracles-use-of-optane-p... [2] - https://tanelpoder.com/posts/testing-oracles-use-of-optane-p...
Holy cow, you nailed it: Intel can't get Optane fabrication cost down fast enough, and everyone is moving to CXL anyway, which presents its own latency challenges that tend to hide Optane performance advantage.
I wonder if you could leverage Optane's bit-level addressability in a shared memory pool scenario.
Sorry to see this tech go...
I eventually pulled it to upgrade to a Terabyte ssd
Optane had great potential for big databases, but you're not running a big database.
It also had very interesting potential for supplementing memory, as a very low latency place to store bulk data. But then the Optane DIMMs ended up being just as expensive if not more expensive than actual RAM. So outside of certain niches, especially because of how narrow the compatibility was, just spend the money on more RAM.
It actually was a net decrease in performance! The reason is that first writes would saturate the cache SSD. Then the octane driver software must manually flush the cache to the main SSD.
What happened was that the laptop CPU actually heated up faster doing this, and reached throttle temps sooner.
I disabled the optane function and the performance increased to something typical.
I was so sick of empty suit sales reps looking at us funny saying "why for you not want optane!?" every other week (read that like an Idiocracy character, please).
I don't give a shit how much Intel is incentivizing you to sell it to me. I. Don't. Want. It.