UltraRAM: 'Universal Memory' That Brings RAM-Like Speed to Non-Volatile Storage
tomshardware.com
tomshardware.com
This is almost the same structure as traditional flash. That is, it has a “floating gate” which sits between the control gate and the channel between source and drain. When a large “write” voltage is applied to the control gate, charge tunnels into the floating gate, and stays there. Then a smaller “read” voltage on the control gate causes conduction in the channel. The main difference here is that the floating gate is made of many thin layers of heterogeneous (III-V) highly-valent semiconductor so that a smaller voltage is needed to tunnel charge into it. Other than the floating gate, this is a very traditional SLC flash structure, and it’s done in a very large area. They use the word “quantum” a lot, but you can substitute that word with “very thin”. I believe the utility of that structure versus bulk III-V is that it creates multiple barriers that allow precise linear control of the tunneling depth over a larger voltage range, and many places for charge to get ‘stuck’, versus one potential well that would just tunnel straight to the drain.
Holy crow, I hate this.
Having too much bandwidth or capacity is a problem in IT that often solves itself more often than not.
The question for me would be whether this is truly a drop-in replacement for either disk or RAM, or if it's something that needs to be treated as a separate peripheral while new applications are written to really take advantage of it— for example, a database that DMAs directly from permanent storage to the network with no concept of loading or caching anything in RAM.
a single device would not be a drop-in replacement of both RAM and SSD for today's computers. fundamental assumptions are made at the architecture level of today's computers which preclude RAM and bulk storage being anything but separate.
a replacement for one or another could be possible for today's machines, though. or both, if you have two devices with the correct connectors. but not both if they share address space.
what does it mean for an operating system, or even an application, to "start up" when all that really means is moving data to RAM so it can be accessed quickly enough?
when storage and RAM do finally converge, there will be no "boot" process in the conventional sense. you will power the device up and it will be ready instantly; everything is already "in RAM" the instant power is applied.
I don't think we would want "program on hard drive" and "program in memory" to be isomorphic for a bunch of reasons, one being security.
just like an MMU can restrict what process(es) can access a particular byte or block of RAM, something similar could restrict access to a common storage medium.
> Well, more than that happens when a program starts up.
of course more than that happens. I am trying to make a point and not document the minutia of the entire boot process of a modern x86_64 computer.
> I don't think we would want "program on hard drive" and "program in memory" to be isomorphic
that is one of the biggest potential reasons to use the technology in question, and one of the biggest reasons that it is being worked on.
Sure, that'd be cool, like segmentation registers.
> of course more than that happens. I am trying to make a point and not document the minutia of the entire boot process of a modern x86_64 computer.
That's fair but it did seem like a central part of your point was that "starting up" was equivalent to existing on disk - though I suspect I was misreading your post.
Heh. Everything old is new again. I think the VM/370 was the first to do this in 1972.
The AS/400 was like that - a single address space for everything. Its descendants still do that to this day as IBM's midrange server family (right below the mainframe line).
We've kind of been there with Smalltalk and image-based Lisp machines - from within the image, everything seems persistent. It's not as cool as it seemed to be back then - being able to reset to a known clean state is useful and having no distinction between what is persistent and what shouldn't be creates some interesting problems.
Worth noting though that computers that had ferrite core memories worked like that. RAM contents were intact after a power loss and the original internet routers, based on the Honeywell H316, shipped with their software loaded into the core memory.
edit: I did some quick searches and it does look like RAM is orders of magnitude lower latency
But certainly ultimately to most benefit from this kind of thing, you'd need a new architecture for it, with a new, RAM-like interface between the CPU and the unified memory backend.
I'd guess we see this kind of thing first either in either niche use cases like microcontrollers, or from somewhere like Apple, since they'd have the vertical integration required to actually pull it off.
Then again, maybe it's even simpler than that— no fancy copy-RAM-to-storage is needed if the RAM itself is persistent. Literally just power everything down, and then when you power on later, pick up from where you left off with nothing more to say about it.
Compare that to 2002 Hynix Net DDR II with a read speed of about 100MB/s and clearly you can see how wrong I am.
My bad, I was confusing Flash that had a DRAM-like interface, and Flash that could read as fast as DRAM. But I cannot edit my response now.
RAM has always been perceived as fast, but storage was always perceived as slow, for decades, and didn't begin to get fast until 20 years ago. I think if there were 15MB/s HDDs matching RAM bandwidth in 1989 it would seem a lot more impressive than a really really fast, near-RAM bandwidth-fast SSD today, even though 12GB/s is so crazy fast it is beyond comprehension.
Thus that mean it also is almost as easy to manufacture?
It's been a research topic to decide how to best use persistent byte addressable memory. Performance disparities are one problem, but the larger issue is actually one of correctness. Data structures are written to protect against race conditions in critical sections, but generally they are not written to maintain correctness in the event of a power failure and subsequent bring-up with some indeterminate persistent state (anything that went through cache is persistent, nothing else is. what was happening when the machine went down)?
It's an interesting problem, lots of solutions. Most use a transactional interface--now doing this efficiently is the harder problem separate from but informed by the actual performance of the hardware.
Take your typical in-memory database: The primary scalability constraint is how much "fast enough" memory you can park close to your CPU cores. Persistence lets you make some faster / broader failure & recovery guarantees, but if those are major application concerns, you probably wouldn't be using an in-memory database in the first place.
Even at peak NAND and DRAM price ( at ~3x of normal ), where Optane had a slight chance of competing, Intel was still selling it at cost, with a volume that could not even fulfil their minimum purchase agreement with Micron. With a technical roadmap that doesn't fulfil their initial promise, and questionable cost reduction plan even at optimal volume.
Compared that to DRAM ( DDR5 / DDR6 and HMB ) and NAND ( SLC / Z-NAND ), as much as I love the XPoint technology the cost benefits just didn't make any sense once DRAM and NAND falls back to normal pricing. Micron already know this all the way back in 2018, but the agreement with Intel kept them from writing it out clearly.
Another obvious operating mode is to make the NVRAM be the first line of swap and have things spill from DRAM into NVRAM when there is memory pressure. I think Intel also offered a mode like this, where it made everything look like DRAM and it would spill onto NVRAM as it needed to. Of course, this would be better rolled into the OS or application level so smarter decisions can be made, but given very high NVRAM densities it’s a great application for the technology.
We should probably start rethinking everything we do from the OS layer to the way we write applications and exchange data. Truly performant non-volatile RAM would be (will be?) a revolution that rivals some of the earliest groundbreaking work in computer science.
In the research world, due to various things about the kind of work that is feasible to do and get published in a reasonable timeframe, they just aren’t looking at these bigger questions as deeply as I think they ought to. People do more minor iterations on the existing approaches, which is a pity.
And if it's not cheaper as RAM it's not interesting anyways. Chances are if you could afford 1:1 paging plenty of applications would happily do and not mind the separation between a fast copy and a persistent copy at all.
That's honestly a worst-case scenario. Imagine an application that is permanently in RAM -- and no matter how often you kill it, it will never "reload" from a clean state.
> its fast switching speed and program-erase cycling endurance is "one hundred to one thousand times better than flash."
I know flash can get some impressive endurance now, but if you start treating it like RAM, is 1000x that endurance even enough?
Anything with limited endurance will need some kind of controller in front of it that makes sure that it wears out evenly. You'll end up with "copy-on-write" RAM that will be inherently slow/space inefficient, either due to block size being way too large for typical RAM access pattern or the data structure managing it becoming too large.
And of course you incur extra latency from having to do look-ups in the first place.
This[1] Open Source data store I'm maintaining is storing a huge tree of tries eventually on durable storage, but it's essentially a huge persistent index stored in an append-only file. The last layer of "indirect index pages" stores list of pointers to page fragments, which store the actual records (the list however is bound to a predefinded size). A single read-write trx per resource syncs the new page fragments, which mainly only include the updated or new records plus the paths to the root page to a durable device during a postorder traversal of the in-memory log. The last thing is an atomic update, which sets a new offset to point to the new UberPage, the root of the storage. All pages are only word aligned except for RevisionRootPages, which are aligned to a multiple of 256 bytes. All pages are compressed and might in the future optionally be encrypted.
I guess these byte addressable persistent memory devices are an ideal fit for small sized parallel random reads of potentially tiny page fragments to reconstruct a full page in-memory.
However, I hope they'll eventually achieve a better price in comparison to DRAM and SSDs than Intel Optane DC PMEM current price range.
And Optane failed that test.
EDIT: But oh, the huge latency!
You might be able to salvage some performance by having a large write queue and running all cache as write-through, but that will make writes, on average, far more expensive than reads.
Perhaps a more reasonable option would be to use only static memory for the caches and have a battery-backed coprocessor flush the caches on power-loss.
There is work to make this transparent to existing data structures, but it (of course) imposes its own overhead.
No longer will we need to "load" things, there'd be no reason to copy an executable or resource from the filesystem into memory to use it, it can be executed, used and modified directly where it is. This is wonderful for a great many things.
A system native to this type of memory may be divided into "permanent" and "scratch" storage. So if you create a new file, you do something of a named allocation, and operate directly in that buffer, and allocate another unnamed (scratch) buffer for the undo/redo stuff. Now there's no longer a save button, since there is only the one instance of your file. Sure, reallocation and stuff will need to be convenient, probably you could have a convenient, but that's a small issue to fix.
Then there's boot times, powering off? save PC and registers to some predetermined place, powering on: load them, and you're ready to continue, literally no boot time.. No reason to put the system to "sleep" you can power it on and off within a few clock cycles. Add a framebuffer on the display, or something like eink, together with "power up on input" and you get enourmous power-savings and battery life
Assuming you do want to keep using virtual memory, you can not-simply¹ use huge pages and avoid the log N TLB factor. Which is a big thing, because TLB misses are very costly, so it can be worth it, even though you are giving up a lot of the paged-memory benefits.
¹ Lots of limitations and caveats. Your MMU needs big enough pages with a reasonable amount of TLB entries. Managing memory becomes more difficult. Your CPU likely has few large TLB entries, so only certain size ranges fit, etc.
For example:
Once upon a time, back around the turn of the century, I worked in animation. Primarily with Flash. After upgrading to Flash 5 we discovered some killer bugs in it related to huge files; when working with the files created by putting everyone's scenes together into one file to do final polish and prep for outputting a .swf, after a few days of that, it would reliably crash within a few minutes of work. And it would crash so hard that it left something behind that made it impossible to just swear, launch Flash again, and get back to work; these crashes could only be remedied by a full reboot. I was the Flash director of the episode we discovered this and I had a very unpleasant week trying to beat that damn thing into a shippable shape.
We'd already worked around some other bugs where Flash slowly trashed its working copies of files; it was standard procedure to regularly save as a new file, close, and re-open, as this seemed to make a lot of cleanup happen that wouldn't happen with a normal save. This had the bonus of giving us a lot of versioned backups to dig through in the event of file corruption or accidental deletion of some important part of it. Which is a thing that can easily happen when you have about a dozen people trying to crank out several minutes of animation on a weekly schedule.
I no longer work with Flash but I have certainly experienced my share of data loss due to my big, complex tools with massive codebases that are older than half the people now working on them getting into weird states. The habits I built back then of saving regularly mean that I rarely lose more than a few minutes of work, instead of potentially losing every hour I spent on a complex file because there is no more distinction between a file 'in memory' and 'on disc'.
So let's wait and see.
A system where RAM is non-volatile can have great many computations running concurrently. No need to "start" a program, its running state just stays since the last time. You may need to sometimes restart a program from a known good state, if things go wrong.
Also, all the undo info (subject to space limits) is persisted and available whenever needed. All the auxiliary structures for a document / picture / code you work on just live next to them, are them. To send a document, you export it to an accepted format, detaching from the in-program representation.
I’m down for that.
tbqh that's a problem with software design.
It's better to just power down devices and not let software react. That would force software to handle unexpected failures much more gracefully and we'd all be better for it.
When all of your disk becomes RAM disk, only persistent, things already change in interesting ways.
But hardware based on a persistent-RAM-only design would definitely need a very different OS. (With a Linux compatibility layer, inevitably, though.)
A file system executing at RAM speed would fundamentally alter perceptions of distribution and thus decentralization.
The real issue is that software and APIs aren't keeping up with the hardware. If you do IO with memory mappings or anything POSIX, you will not be able to saturate a modern NVME SSD. And I'm not even talking about Optane here.
This means that there is no market for these crazy fast devices. Look at the storage-optimized instances of every major cloud provider, and you won't find them. They would be too expensive, and the software people would run on them wouldn't be any faster than the older, cheaper SSDs.
It'll take time for the software to evolve and catch up. Initiatives like DAX, eBPF, and io_uring in the kernel are a big step forward in the right direction.
And even before that there was Linux AIO etc.
This is not a home-use-case, but it certainly exists today in larger Enterprises.
So it's not (necessarily) about having a very fast filesystem, it's about having persistent memory.
Is the goal to say "okay, here is a persistent chunk of memory with a file or program in it, keep it preserved while the OS reboots, then if it's a program reconnect the files and sockets it was using"?
Because you can just remove the word 'persistent' and do that today. Most hardware shouldn't wipe memory on a reboot, and if it does talk to your vendor for a BIOS option. And all the support code you need in your OS should be roughly the same.
Part of me hopes someone in here proves you wrong :-) Us techies do mad things sometimes...
If this becomes a real product, I expect that market will be created. My guess is that it would show up in Apple hardware very early on, if not first. It probably would increase battery life (if I understand this tech correctly, you won’t have to power the RAM while the system is sleeping), and they’ve shown they’re willing to completely redesign their hardware.
For instance you can perform an on-disk b-tree lookup in the kernel.
These devices are so fast that crossing the kernel boundary becomes the limiting factor.
That is still not enough write operations to let anyone treat this like a bigger version of RAM, is it?
SRAM needs 10^20
DRAM 10^17
Any memory exploiting variable dielectric properties is subject to dielectric failure.
The interesting thing here is the non-volatility for this type of structure on Si instead of GaAs.
So, I was confused by the title of the article.
DRAM does have a retention limit, only a few milliseconds. That's a drawback.
"Memristors can potentially be fashioned into non-volatile solid-state memory, which could allow greater data density than hard drives with access times similar to DRAM, replacing both components."
So it mainly stores the whole index in one file (append-only) plus an offset file to map timestamps to revision offsets (besides the 32bit revision numbers / revisions stored in the main file). This means that we can store checksums of the child db pages in parent pages as in ZFS and to store the checksum of the UberPage in the page itself to validate the whole resource and we don't need a seperate transaction log despite an in memory cache for the single read-write trx.
https://github.com/sirixdb/sirix and a query processor called Brackit for document stores based on a JSONiq like query language: https://github.com/sirixdb/brackit
https://www.theregister.com/2015/07/28/intel_micron_3d_xpoin...
Always thought 3DXpoint (later released as optane) seemed cool, but had an almost impossible hill to climb by being completely oversold before it was released such that it was doomed to disappoint.
No need to worry boys, Windows will grow in size accordingly. Remember, if Windows is not eating 20% of your storage, you don't have the latest according to Microsoft evangelists. (/s)
You especially don't want to be building a 3d memory array with dozens of layers of cells if each layer requires a few EUV steps. That's never going to lead to economical high-volume manufacturing.
Oh. Okay.