They're a nightmare to program because OSes do not have a good abstraction for them (at least not yet). Accessing them through the file-system seems sub-optimal (this is byte-addressable memory and not a block device). Accessing them through virtual memory is also pretty bad because they're much slower than DRAM.
Both Windows and Linux implement DAX, which, as @the8472 explained, allows bypassing page cache in memory mapped I/O. Additionally, DAX optionally allows you to flush your data directly from user-space instead of calling msync.
And that's the gist of NVM programming model [0], its entire point is to allow applications to avoid the now hugely excessive abstraction layer of traditional storage.
And I will freely admit that programming to raw memory mapped files can be difficult, but there is ongoing work on making it easier. An example of that is, excuse the shameless plug, Persistent Memory Development Kit [1], which makes writing new software for this new type of memory much simpler.
Performance of an NVDIMM is obviously hardware dependent, but the now widely accepted programming model works with the assumption that persistent memory is fast enough so that it is reasonable to stall a CPU while an instruction is accessing it. I'm not sure on what hardware evaluations you are basing your claims on, but let me assure you that the HW solution being described in the blog post does not violate that assumption.
[0] - https://www.snia.org/tech_activities/standards/curr_standard...
[1] - http://pmem.io/
If you need dynamic mutable state however, as great as these libraries are, you will need a more complex solution with memory allocation and transactions.
Think of persistent-memory data as more like data resident in the memory of a runtime which can experience a "hot code upgrade", like the Erlang runtime.
In Erlang, when you hot-upgrade your running code, you usually do so through a managed system of "relups" (RELease UPdates), which are sort of a cross between an RDBMS migration, and a traditional installer-package full of newer versions of code and assets.
The Erlang runtime takes this package, unpacks it, and then runs a master relup script, which can been authored to do arbitrary things (including, if ultimately necessary, fully rebooting the node, throwing away all that in-memory state.) Mostly, though, a relup script calls into individual "appup" scripts for each Erlang application. Those applications then specify how their corresponding running processes are to be updated—which can sometimes be fraught (if e.g. the new release requires that you add new service-processes or remove old ones, migrating in-memory state into a new architecture), but usually just means calling a "code_change" callback on all the service-processes.
This "code_change" callback is the thing that's most like an RDBMS migration: it is called from the event-loop running in the old version of the code of the service-process, and passes in the old in-memory state; and when it returns, it's returning the new in-memory state, to resume the event loop in the new version of the code of the service-process.
This is basically how I'd picture dealing with code updates (including ones due to build-setting changes) in software that deals with pmem: you'd architect your code such that the library that touches the pmem can have multiple versions of it dynamically loaded (though not running) at the same time; and then you'd stage a migration from the old code's pmem state encoding, to the new version's, by
1. dlopen(2)ing the new version of the lib;
2. telling the old version of the lib to stop any ongoing work;
3. handing off the toplevel pmem state-handle that the old version of the lib was using, to a "migrate" function in the new version of the lib;
4. replacing the old version's pmem state-handle with a dummy one;
5. telling the old version of the lib to terminate (and so do the trivial cleanup to the world it sees through the dummy handle);
6. tell the new version of the lib to initialize, using the handle to the now-migrated-in-format pmem;
7. dlclose(2) the old version of the lib.
Basically, picture what something like Photoshop would have to do to enable you to upgrade its plugins without restarting it or closing your working document, and you'll have the right architecture.
If 1. the new format is just like the old format except for one little difference to one struct, and 2. structs point to other structs, rather than containing them; then it's just a matter of calling your within-mmap(2)ed-arena malloc(2)-equivalent function to get a new chunk of the pmem arena of the right size for the new version of the struct; and then rewriting the pointer in the other struct to point to it; and then calling your free(2)-equivalent on the old version of the struct.
If you change the structure of some fundamental primitive type like how strings are represented, then you're probably going to have to rewrite your whole pmem arena.
Though, also, you can just make your code deal with both old and new versions of the struct, and only migrate structs when they're getting modified anyway. (This is equivalent to the way you'd avoid an RDBMS migration rewrite an entire table, by instead adding a trigger that makes the migration happen to a row on UPDATE, and then ensuring that your business-layer can deal with both migrated and un-migrated versions of the row.)
That's part of the reason why I was thinking something like Cap'n Proto or Protocol Buffers might make sense for a lot of structures. You pay a bit of cost for writing but get to gracefully handle upgrades to the structure if you do it right. I'd imagine you want to use something higher level just above them to organize the records. But this is all a really new area of thinking about this so I'm probably being a bit obtuse about it.
With DAX[0] linux already has the ability to put a filesystem (currently ext4 and xfs) on NVDIMMS and then let userspace address them through mmap while skipping the page cache indirection. I.e. you're directly byte-addressing them through the memory controller via standard memory-mapped file abstractions. Direct block device mapping of nvdimms without filesystem is also possible.
[0] https://www.kernel.org/doc/Documentation/filesystems/dax.txt
Also a battery would only last so long. IIRC DRAM needs to be constantly refreshed, so, it would be a trade-off between capacity and duration.
Optane seems[1] to be 20~30X slower than DRAM but 4~10X faster than server SSDs
If persistent memory pans out as a technology it will completely upend the way we think about building software and the cost tradeoffs of hardware. (as much as or more so than the transition from spinning disks to ssds)
Because I'll keep reminding people that putting a DRAM cache in front of some flash can very closely approximate a large persistent memory. If people wanted to build software for that kind of system, they could do it today. The hardware is not the blocker.
Note that Optane DIMMs have been delayed by around two years at this point and we still don't know what they will cost.
As for cache-warming, that's also a configuration issue. When you reboot, leave the 'cache' portion of DRAM alone. Then as soon as the service resumes, the cache is already hot. When you shut down a node for an extended period, consider spending five minutes writing the cache to disc. And the article is about cloud servers anyway, where a shutdown typically implies losing all local storage whether it's persistent or not.