Intel Announces Optane DIMMs
anandtech.com
anandtech.com
Of course it also makes it possible to rapidly reboot into an OS if all you need in the boot loader is to copy a gigabyte of 'ram' from one place to another and then jump to it. What that enables is powering down servers that aren't being used with the ability to power them up in milliseconds rather than seconds. That has been of Google and other cloud providers to have 'power proportionate computing' clusters.
Very cool stuff indeed.
I suppose you don't even have to copy it before you jump to it. You could map the pages and then lazily copy when you get page faults.
> powering down servers that aren't being used with the ability to power them up in milliseconds
Still, if this is your goal, can't you get 90% of the way there by keeping some DRAM powered up while everything else is off? Refreshing DRAM requires some power, but not that much from what I understand (as a software guy).
The memory bus on these processors is 3 - 30 GBs so copying a full gig would be a 3 - 30mS sort of deal.A pretty short start time. And yes, you tell the BIOS to skip POST (its a common start up option in even off the shelf BIOS packages these days) in order to get a faster boot time.
This should say 30 - 300ms I think
You could also just map it and not even fault on access. For data that's mostly read and seldom written, and not too much of a hot spot, there would be no reason to move it to DRAM. Depending on the read latency of these DIMMs, that might include a lot of executable code.
STR suspend to ram, available in off the shelf computers since ~1995
[1] https://newsroom.intel.com/news-releases/intel-and-micron-pr...
Once you introduce volatile memory as highspeed buffer, the coolness factor drops quite a bit as you end up in the same volatile/non-volatile tiered storage split that we have always had. It's just back on a parrallel bus like in the good ol' days of PATA.
Likewise, quad-channel DDR4 setups can push around 60GB/s, while Optane PCIe SSD's have only been able to push around 2GB/s. The PCIe x4 interfaces that they use should have been able to push 4GB/s, and if they started hitting a wall there, they could just have used a full PCIe x16 interface which could push around 15GB/s. This indicates that the throughput was not interface bound.
Nothing seems to signal that Intel have have anything up their sleeve to provide the 25x latency improvement and 30x throughput improvement needed to just be even with DRAM, although placing the chips on the memory bus will of course provide some speedup.
I would not be surprised if early users of these devices were seeing 10GBs+ of data throughput.
I can say that the PCIe x16 cards we develop at work have no problem reaching the theoretical maximum, transmitting around 15GB/s worth of payload data (e.g. our 2x100Gb/s NIC with one port pushing 100Gb/s in both directions, and another a few tens of Gb/s in both directions). We don't make cards smaller than x8, so I can't give measured numbers there, but Intel should have no problem transmitting ~4GB/s on that x4.
NVMe surely has an overhead, but this overhead is far from interesting, unless you are implying that the overhead eats a whopping 35% of the available bandwidth. Likewise, if Intel had hit a PCIe/NVMe bottle-neck, going to up x8 would not not have been difficult in any way (x16 is more annoying, though).
The performance has either been restricted by the flash itself, or by the controller—not by the interface. The numbers are also small enough that the software side shouldn't be a problem yet. Attaching to the memory bus minimizes penalties that are unlikely to be a bottle-neck.
The latter is probably a firmware bug, but the best Dell ProSupport has been able to do to solve it is "flea drain" (unplug, hold the power button for 60 seconds, plug back in).
It's quite a pain in the arse, and one of the reasons we prefer to use Supermicro machines for development at work, using only Dell for testing.
Given your typical motherboard which you tell to bypass POST, it can go from power applied to memory controllers alive in just a few microseconds. The code then to go from there to running is also quite small, especially if there aren't things like graphics controllers on the PCI bus that have to be initialized first.
No matter how fast a disk is, using it means either some expensive serialization/deserialization step (and also the associated memory access to create the 'working' object that my logic actually works on) or writing my algorithms to forego in memory objects (and the associated features offered by my programming language, e.g. classes / objects or whatever) and working from the raw byte values.
What I really want, and would be a game changer as to how we use things, would be that my programming languages heap can be made persistent (or at least a part of it). In this case instead of:
var mything = new Thing();
load_thing_from_disk(mything);
I might have: persistent var mything = new Thing();
Done. However this also introduces more questions, like transactional commits to memory etc (as few apps are coded to ensure consistency of memory across reboots).However I cant help thinking that some way to harness persistent fast memory without needed some complex disk->logic mapping would be a game changer.
Edited: spelling and wording
It is the game changer that you wish for, since the marshaling logic that you mention is gone. Persistent Memory can be accessed directly through memory mapped file, bypassing the traditional read()/write() I/O paths. Recent file systems have also been modified to a) skip the page cache layer and b) forgo the msync() call that would be otherwise required to synchronize the modified pages. This is what's called DAX (Direct Access [0]). In the place of msync() you can now just use CPU cache flush instructions. These two file system changes entirely eliminate kernel code from the I/O path (apart from the initial page faults).
Persistent Memory Development Kit contains libpmemobj [1], which is almost exactly what you are imagining ;) It's a persistent heap, with transactions for durability. It's not as nice (yet) as your code snippet, but here's C++ example [2] of a persistent queue push:
obj::transaction::exec_tx(pool, [this, &value] {
auto n = obj::make_persistent<Node>(value, nullptr);
if (head == nullptr) {
head = tail = n;
} else {
tail->next = n;
tail = n;
}
});
`make_persistent` is, akin to `make_unique`, a memory allocation of a "Node" class. Once allocated, we can just assign the newly allocated object to a different persistent variable. No kernel code executing, no serialization ;)[0]- https://www.kernel.org/doc/Documentation/filesystems/dax.txt
[1] - https://github.com/pmem/pmdk
[2] - https://github.com/pmem/pmdk/blob/master/src/examples/libpme...
[0] - https://www.boost.org/doc/libs/1_63_0/doc/html/interprocess/...
[1] - http://pmem.io/pmdk/cpp_obj/master/cpp_html/classpmem_1_1obj...
The actual hard part of persistent heaps isn't the persistence part. It's transactionality and upgrade management.
It's still neat, though.
So its still a serialize/deserialize cycle, but the access libs built on top of the persistent memory look interesting.
There's been a lot of interesting research around file systems for persistent memory.
One that shows a lot of promise is NOVA [0]. Its focus is on making full us on this new type of memory. And it's not just pure research, they are attempting to get NOVA included into Linux kernel [1, 2].
And while talking about file systems, we shouldn't forget about the effort that was put into modifying the existing ones to support DAX (Direct Access) [3, 4].
[0] - http://nvsl.ucsd.edu/index.php?path=projects/nova
[1] - https://lwn.net/Articles/729812/
[2] - https://lkml.org/lkml/2017/8/5/188
It didn’t even occur to me that the file systems will need to change to fully take advantage of NVRAM. I wonder at what point the abstraction will stop leaking and require another higher layer to account for differences in performance. I’m sure the OS will need tuning, but applications might not unless they’re pretty bare metal.
These are byte addressable - they will look like RAM to the OS. (If you have a motherboard that supports them, there is a slight change to the spec).
To the application developer, the interface will be http://pmem.io/pmdk/ . Last I looked, there were several ways to do things, but the most commonly used would basically be allocating a chunk of memory and assigning it a file name so you could re-open it next time.
This is exciting because it could open up exciting possibilities like zero-CPU IO with DMA straight to persistent memory, "sleep" mode that is essentially free, re-thinking paging in modern operating systems, and generally re-thinking everything we've assumed since core memory.
On the same day we are talking about WASM microkernels in another thread. Things are getting fun again :)
That would be https://news.ycombinator.com/item?id=17187384
So the PC platform finally caught up with Amiga's RAM disk ;-)
Still, I do fondly remember my Amiga days!
But the speed and density seems to be a problem. E.g. it's still 2x slower than DRAM, but theoretically should be faster. It could also have 3 bits, not just 2.
It will be interesting to see the real benchmarks on this DIMMS. We all want to know if they're comparable to real DDR4 memory.
The current market is terribly overpriced (there's some debate on if there's price fixing with the big three or if it's a genuine shortage/supply problem with the Note recalls and new phone releases). DDR4 is nearly double the price it was the last time I did a build over a year ago. :-/
EDIT: Looks like these chips will be specialized for certain server boards/CPUs and only share the DIMM interface and not protocol.
How does my OS/app see this? Is it accessed like regular DRAM memory... except slower and persistent?
Or would my OS see it as a "normal" drive... except one that's really fast and happens to be connected via a DIMM slot instead of PCIe/SATA/whatever?
It is very likely that generic kernel support will come for use in Linux and Windows directly, building on top of the existing DAX systems in those operating systems (DAX - direct access - APIs being used for IO to memory-like devices, bypassing cache layers which are useful for more traditional storage types). This would allow a user to create a regular old storage volume in their NVDIMM for general use.
https://www.kernel.org/doc/Documentation/filesystems/dax.txt
Do note that NVDIMMs aren't a drop-in replacement for regular DRAM DIMMS, despite using the same bus and electrical subsystem. You'll need proper hardware support on your motherboard and CPU, since memory controllers are on CPU these days.
Do you mean it will present them to userland as storage, rather than see them as storage?
Seeing them as storage implies to me that the DIMM emulates an AHCI, which i don't think is the case.
128GB, 256GB and 512GB per module is sadly too much for consumer motherboards. Why not a 16GB version, didn't Intel even launch Optane with those small sizes?
Its cheaper and slower than DDR4, but persistent and likely to be way denser.
> 128GB, 256GB and 512GB per module is sadly too much for consumer motherboards.
NVDIMMs are for database applications. If you were running 1TB in-memory databases, but are willing to lose a bit of performance to severely reduce costs, you're in the market for an NVDIMM.
While I don't have the figures to back it up, I believe the differences in latency (and by extension random access perf) are even more dramatic, and where the real performance advantages come from.
16GB is a single 3D XPoint memory die. The DIMMs need to use more than a few dies to support the throughput that people expect from their memory bus. The same is actually true of DRAM; if your current memory modules only had one DRAM die each you not only would have 1/8th the memory capacity, but your memory bandwidth would be annoyingly small as well.
New CPUs will be required, and if there's a technical justification for that it will probably be that accessing an Optane DIMM requires timings that are far outside the normal range for DRAM modules that the existing memory controllers were designed to accommodate.
https://software.intel.com/en-us/blogs/2016/09/12/deprecate-...
NVRAM solutions exist today that are just DRAM with a supercap and/or battery plus a flash SSD. Or PCIe cards. Or ordinary RAM plus specialized firmware that kicks in on brown out and copies RAM out to SSD.
Optane looks neat, but it's kind of expensive and has relatively low write endurance compared to ordinary DRAM (from an enterprise filesystem journal perspective).
Here's some documentation about the version of it that Sun Microsystems licensed and sold:
https://docs.oracle.com/cd/E19957-01/801-4896-11/Presto.chp1...
It says:
"Each NVSIMM contains memory, a battery, and power controller circuitry, which ensure that the memory is not lost when the system is shut-down or halts because of an abnormal condition."
"Synchronous write requests to disk are intercepted, and the data is stored in non-volatile memory"
I sincerely hope no storage company acknowledges a write after just writing to volatile main memory. That's a recipe for disaster when a node goes down.
Sorry to break this to you, but this is exactly what every major storage product does. Data will be mirrored to separate DDR behind and then good status is given to the host. The data will be destaged at some point later when cache space is required for some other operation.
The data itself is safe as long as the battery backup (or capacitor for smaller systems) is charged enough to handle a power outage. The storage system knows the battery levels and may not allow a write cache if there isn't enough supplemental power to destage the full write cache in the event of a power loss.
https://en.wikipedia.org/wiki/NVDIMM
Fortunately XPoint should solve this.
For some SSDs, it's under a week.
"Sequential Write (up to) 2200 MB/s"
41e15 / 2.2e9 /60/60/24 = 215 days sequential write to failure
"Mean Time Between Failures (MTBF) 2 million hours"
2e6 / 24/365.25 = 228 years MTBF
So it seems the MTBF is being stated at 0.26% average write utilization.
[corrected math]
A 750 GB drive assuming you want to store the data for 48 hours can only write (750 GB/24/60/60*1000) = 4.34 MB/s on average. Dropping that to even 1 hour still gives reasonable lifetimes.
For a MTBF of 2 million hours, that means that on average, if you have one thousand drives, then you should expect one drive failure every 2000 hours, or 83 days (1k * 2k hours = 2M hours)
Of course this breaks down at an MTBF of over 50 years, as thosre rarely mention exotic failure modes, and don't actually have an MTBF of 50+ years over the life, but an annualized failure rate corresponding to 50+ years MTBF, measured over the first couple of years or even the first year. For non-wet-electrolytic-capacitor-using computing, one can calculate a temperature-dependent MTBF in the 5-25 years range, mostly depending on how bad the chips are hit by electromigration and similar aging in the semiconductors. This is incidentally a reason why I miss clock speeds for different processors as reported by overclockers to at least in some cases extrapolate the life due to electromigration, as there is a formula with like iirc 2 parameters, which gives a temperature (and maybe voltage) dependent lifetime/MTBF for this semiconductor device. I'd likely sttrive for about 5 years MTBF on the processor, if speed is of concern and reliability/uptime not in the foreground.
That assumes your drives fail with a constant independent probability, like nuclear decay events (Poisson distribution). The reality is more like https://en.wikipedia.org/wiki/Bathtub_curve .
MTBF is not a good metric for complex systems used under wildly-varying load conditions, but ... it's a metric.
It seems quite low for any practical purpose. I don't doubt that there probably are some tiny shitty drives that will conk out after a week like that but are there any reasonably popular drives like that?
https://www.anandtech.com/show/12674/samsung-announces-970-p...
This was a drive by a top brand NAND manufacturer.
[1] https://www.cs.cmu.edu/~jarulraj/papers/2015.storage.sigmod....
Maybe something like Aerospike? (don't know much about it but I've heard it's "for" that)
But to fully benefit from persistent memory, the DBMS will need to modified. To see how that might look, read this [0] post by Microsoft about their efforts in SQL Server.
There's also an interesting research paper from CMU [1] that talks about challenges associated with pmem in the context of databases.
[0] - https://blogs.msdn.microsoft.com/bobsql/2016/11/08/how-it-wo...
[1] - https://www.cs.cmu.edu/~jarulraj/papers/2015.storage.sigmod....
[0] - https://news.ycombinator.com/item?id=17195018
[1] - https://www.openfabrics.org/images/2018workshop/presentation...
The "Full procedure" of reading a DRAM cell is:
1. Row-Address -- Load a "row" (usually 1k to 8k. DDR4 is 8k IIRC) to the sense amplifiers. Sense-amplifiers can indefinitely hold data, but there's relatively few of them.
2. Column Address -- Once loaded, you talk to the sense-amplifiers.
3. Precharge -- You begin to move the data from the sense-amplifiers back to the DRAM cells. Again, step #1 obliterated the data, you have to write it back regardless.
4. Row-Address -- After the old data is loaded, you send it back.
So regardless, you have to Read-then-write EVERY time. In fact, DDR4 has faster write-speeds because you don't have to do the read step if you are only writing.
Intel doesn't publicly share the full specifications documents for their SSDs any more, just the 2-page product briefs. And the news article contains all the official information that's public so far about the Optane DIMMs.
They mention write cycles in the article...
> The existing enterprise Optane SSD DC P4800X initially launched with a write endurance rating of 30 drive writes per day (DWPD) for three years, and when it hit widespread availability Intel extended that to 30 DWPD for five years. Intel is now preparing to introduce new Optane SSDs with a 60 DWPD rating
For example, in the x86 world, pairing NVMe drives with putting portions of the application writing to NVDIMM drives core performance, say, in SQL, from 40% util to 100% util.
Even a single 8Gb DIMM can dramatically increase utilization and performance.
And it might even be slightly better than halving the bandwidth needs, since swapping DRAM banks isn't free, so you might be saving on mildly thrashing your DRAM controller when you're using DRAM and the NVMe drive is trying to read at the same time.
The write cycles are covered in the article, though without much detail as to how wear leveling works with this kind of setup.
Basically you run everything you can in memory and then just mmap() in the files you want to use?
How will user-mode applications refer to their persistent data if not by a filesystem path? You gotta put access permissions on some kind of object that humans can copy-and-paste into their backup scripts.
https://www.usenix.org/system/files/login/articles/login_sum...
https://en.wikipedia.org/wiki/Phase-change_memory#Timeline
Theoretically it should be 1000x more durable than flash, and also 1000x faster, but they are not there yet. But it looks like they solved the packing problem. And Micron insists that it is chalcogenide based, but not "phase-change memory", the one they started in 2012 and took back in 2014.
3D XPoint is said to be able to stack the storage medium on the die and requires no access transistor -- which "PCM" has. So it is based on PC material but has some new kind of selector.
The 3D/stacking part is important for scaling/density -- even NAND-Flash has hit 2d limits and gone vertical.
This interview from 2013 is still worth a listen imo!
http://www.se-radio.net/2013/12/episode-199-michael-stonebra...
That is pretty impressive.
e.g. malloc owner PID x is now PID y
To summarize, this product is shot, and is just hype.
I'll check the technical questions/gaps and answer or fill them in tomorrow.
Considering Intel's history, I wouldn't bet they direct their efforts based on a YouTube reviewer.
Also, considering Intels very anti-competitive behavior (e.g [0], [1], [2]), I am wary of stating Intel entering the DDR industry will make it any less corrupt.
[0]: https://www.theverge.com/2014/6/12/5803442/intel-nearly-1-an... [1]: https://www.wired.com/2009/12/ftc-sues-intel-for-anti-compet... [2]: https://www.youtube.com/watch?v=osSMJRyxG0k
Eh. https://en.wikipedia.org/wiki/Intel_1103
RAM was their 1xxx product line, EPROM was 2xxx, microprocessors 4xxx (later 8xxx, what with 8 glorious bits of data bus ...)
How so? Point to the quote?
Persistent memory is something that Intel has been working on for at least 5 years (https://github.com/pmem), and given that that's the public face of the software side of it, they were likely starting to develop the hardware even earlier.