UltraRAM
tomshardware.com
tomshardware.com
> 10 million write/erase cycles
this is not going to compete with DRAM, which needs to endure trillions of write/erase cycles in its lifetime.
Unless they grossly underestimated its durability, a name like UltraFlash would seem more appropriate?!
Maybe I'm confusing something, but to reach a trillion cycles in, say, a year, would take overwriting all your memory 30 times a millisecond. That doesn't sound right?
Or is that trillions of any writes and erases?
Still if you only changed the state of the memory once per frame, you would do it in RAM, not in cache. At 1000 FPS (we should consider the worst scenario even if rare) that's 3 hours of playing a game to reach 10 800 000 reads/writes.
Now question is what happens if that bit gets damaged, perhaps the memory just disables it as damaged, and uses another bit for this memory address from now on. Perhaps it makes the ultra ram slower over time as more bits (sectors) get damaged?
I agree that some regions risk being R/W more than others, so memory controllers should indeed perform some kind of wear levelling, but otherwise I find it hard to imagine trillions of overwrites across GBs (or TBs) of data. 1e6 cycles is definitely doable, and on the low side, even for flash devices. 1e9 is pretty good for general-purpose memories, and few applications require 1e12. Not even SRAM or DRAM have unlimited endurance, due to physical wear. It's hard to find a source on this, but I would probably hand wave it at around 1e15 cycles for DRAM? This would be 30 years of operation for one access every microsecond.
I think a trillion, or at very least 10s-100s billion is the correct amount of cycles for RAM.
Lots of non-pathological workloads might write to a memory location every millisecond, such as a game with a 4-pass renderer running at 240Hz.
Until recently a billion was a trillion, or vice versa, depending on whether you're from the UK or the US.
A GHz is a GHz no matter where you are. :-)
Optane did inspire a lot of R&D into persistent data structures, databases and file systems that started to challenge the traditional model of local memory and persistent storage. IMHO, a few of those projects were a little bit overoptimistic, and used NVRAM as DRAM without many restrictions. For NVRAM to be viable, I think it still needs to have overprovisioning, wear levelling, memory-protection and transactions, provided by hardware and/or an OS but not necessarily with traditional interfaces. It is mostly a matter of mapping it CoW via a paging scheme instead of directly, and it will still be at near-DRAM speed.
The performance is so high that the assumptions that had led to the old file system interfaces don't apply any more. There is opportunity for something better.
begin()
try {
work on
several
in memory
data structures
commit()
}
catch {
rollback()
}
And all those data structures are either collectively updated or not.When writes are persistent and cause wear, the consequences of e.g. a common buffer overflow or use-after-free bug can be much higher than if they were not. Even an unoptimised loop that writes to NVRAM could be bad.
You would probably go for some approach where most memory addresses are direct-mapped, and then the few that have been written most are redirected to new addresses.
The reading of the direct-mapped addresses would be super fast, since you can do the read in parallel with the lookup in the remapping table (just to check that this is a direct-mapped address). Reads of non-direct mapped addresses might take a couple of extra cycles, but that doesn't matter because they are very rare.
To do any of that, CPU memory controllers need to be able to handle per-request variable-latency RAM, which to my knowledge today they do not, although it would not be a big redesign to add.
That's not true, a shared counter (i.e., an atomic integer) is cached – in fact, there's no guarantee that its value is ever written back to system RAM.
You're probably thinking of non-cacheable memory: the kernel can set the MMU attributes of a memory page such that the CPU will avoid the cache when it accesses addresses in that page. This is completely independent of atomic accesses on memory locations [1].
[1] At least typically – there may well be CPUs which disallow atomic accesses on non-cacheable memory.
RAM is not easily removable in most of today's electronics. So replacing RAM once a year actually means replacing all your devices once a year.
no one complains about not being able to replace the processor in their phone because it 'never' breaks. batteries on the other hand do, and to varying degrees are replaceable.
https://www.youtube.com/watch?v=X7C_hdJsY4Y
I think for my kids I may have them skip traditional through wire soldering for SMD with hot air and toasters.
https://hackaday.io/project/27900-reflowduino-wireless-reflo...
Assuming a typical 5-year lifecycle, 10 million writes means 1 write every 15 seconds. That's more than enough for executable code, CDN content, or a database index. I can definitely see systems with 75% UltraRAM for read-heavy data and 25% traditional RAM for write-heavy pages acting basically as L4 cache.
> The process was repeated five times, resulting in a little over 10^7 program/erase cycles applied to the device. As can be clearly seen in Figure 4d, there is no degradation of the ∆IS-D window throughout these tests, meaning that the endurance is at least 10^7.
https://onlinelibrary.wiley.com/doi/epdf/10.1002/aelm.202101...
Why stop at 10M? Is the erase operation really slow?
Ten trillion cycles would take over 150 years.
I'm guessing a silicon lab doesn't have "the rest of the computer" that would allow them to run this ram at full speed constantly. This UltraRAM isn't something they can just slot into their motherboard.
10 million is 14 hours. It takes longer than that to prepare your documentation. Something is rotten in Denmark. A skeptic could very, very reasonably assume that cherry-picking is going on here, and that 10m to degradation isn't far off from the truth.
> Assuming ideal capacitive scaling[33] down to state-of-the-art feature sizes, the switching performance would be faster than DRAM, although testing on smaller feature size devices is required to confirm this.
So, they have no idea of its performance. Yet.
The current set up is based on separating volatile and non volatile memory and adding caches to paper over the slowness. Caches are getting bigger and bigger because of the huge speed disparity. I think you underestimate how much of a game changer this could be.
This is persistent and fast.
If this takes off, and it does only last 10s of millions of cycles, just use cache for fast changing things and ultraram for everything else.
If it lasts trillions of cycles, it potentially would completely change pc architecture. It was the 80s when we had ram/rom that could keep up with the processors of the day. This potentially gets you an instant on computer, no need for caches, no need for memory for the graphics card, no separate hard drives. Just one big simple bucket of bytes for everything.
If they've done that, awesome. Make it, show that it works, licence how to make the thing to semiconductor companies and retire wealthy. Or maybe the university owns the IP.
I'll leave it to the experts.
That's a common package for testing ICs -- notice the array of dies inside and the haphazardly placed bonding wires. It isn't the final form factor.
I doubt they've actually demonstrated the speed/power claims practically; that's what the new test kit, and potential fab partnerships are for.
If this actually pans out, it will be worthwhile to stack a lot of it on the same package as the CPU. The reason memory is so far in current systems is mainly that having it closer wouldn't actually meaningfully help, because almost all the latency is reading data from the DRAM array anyway. If they suddenly get an economical new memory type that has an access latency of tenth of what DRAM does, they are going to figure out how to get it close enough that the signal travel will not be a meaningful part of the total latency.
The OS can just always be loaded and ready to go; when power is restored it checks to see if the hardware has changed and just loads up the 64 MB of CPU cache. It could take just a few milliseconds. It takes on the order of a millisecond to charge the capacitors in a desktop PSU. "Restarting" becomes basically the same thing as reloading, and takes >100s of times longer than actually restarting the device. That's crazy to consider.
If boot time is 0, stuff will just unplug itself after its been idle for a few seconds. I'd expect the hardware in phones/laptops to become more distributed, with basic vital functions handled by a separate processor. Probably the screen gets taken over by a very simple processor that can only display the time, battery %, cell info (or the current screen buffer, for a laptop) and user input causes the main cpu to wake up in between frames.
1ns write operations suggest fast read too.
Shifts like this are so impactful it’s hard to predict exactly what good designs will look like until we’ve had 5-10 years hands on for the industry to shake out how the Hw topology will looks like (maybe more since HW dev cycles prevent fast iteration and testing of ideas)
> In all of the above tests, the program and erase states were set using between 1 and 10 ms voltage pulses, two times longer than the switching times used in our recent report of ULTRARAM on GaAs substrates.[15] In both cases, the devices operate at a remarkably high speed for their large (20 μm) feature size. Assuming ideal capacitive scaling[33] down to state-of-the-art feature sizes, the switching performance would be faster than DRAM, although testing on smaller feature size devices is required to confirm this.
> Why do you even need a cpu cache?
Cell read time is entirely different from latency and throughput. This stuff still reads in rows like RAM and can't just be accessed freely like registers.
This is why CPUs have multi-level caches, even though the transistors in L1 cache and L2 cache are typically the same -- the difference in access latency is not because L2 is made of slower memory, but because L1 is a very small pool very close to the CPU with the load/store units built into it, and L2 is a bit further away.
However, if main memory latency is suddenly a lot lower, it might change what is the most efficient cache level layout. The currently ubiquitous large L3 cache might go away. That would of course require very high bandwidth to the memory chips, because L3 does bandwidth amplification too.
...The OS can just always be loaded and ready to go; when power is restored it checks to see if the hardware has changed and just loads up the 64 MB of CPU cache.
The idea, called “Orthogonal Persistence” way back when, has been around quite awhile. Here’s my (probably spotty) idea of the history:
Researchers wanted instant-on for their early visions of tablets. To make sure security and networking would still work properly, there was an idea to use Capabilities (which were around since the 1960’s) to support this and solve the chicken and egg problems that were thought to arise.
Capabilities later became widely adopted just for better security, but Orthogonal Persistence never took off, because never rebooting would have required much higher levels of reliability, which would have been expensive to achieve. So today’s devices still reboot, but also have a fast "wake from sleep."
So I’m not sure if we will ever have true “Orthogonal Persistence.” We might have much slicker “wake from sleep” instead.
I'd expect the hardware in phones/laptops to become more distributed, with basic vital functions handled by a separate processor.
This is already the case!
> The technology has been integrated into Xilinx's FPGAs and Zynq UltraScale+ family of multiprocessor system-on-chips (MPSoC).[7]
Referencing a paper https://www.eejournal.com/chalk_talks/2016033002-xilinx-ultr... FROM 2016 !
The tech looks super cool if it does get commercialized.
Capacity and price killed it, no word about these in the article
Cost permitting, this stuff would replace RAM, not the drive. No more loading into ram; now the bottleneck is loading into cache and that will always be trivially fast just because cache is so small.
Even if its too expensive to replace RAM, if it can fit the minimum bits of an OS then I think cold boot time still goes to 10s of milliseconds. Might take a couple years, but interactivity doesnt need to wait on ram to be filled.
Barely so. NVMe sequential throughput is measured in gigabytes per second. So you can get this under 300ms. And you can optimize the order in which things are loaded so that the important ones arrive first, not all in-memory data is hot.
What makes booting take time are serial dependencies between boot stages, timers (boot prompts for humans, but also for hardware to power up), careful device enumeration and initialization and stuff like that.
If you've been unpowered all day, or if your hardware has changed, or you're worried about security, then you can choose boot from scratch. The only other reason, IMO, is because the computer has just been put together. If all those parts can restore their previous configuration, all they have to do is signal "yep, I'm still in the same configuration" and we should be able to pick up where we left off (again, except for disk drives/ram/networking etc).
So, will it scale down? Will it be cheap to manufacture?
This seems misnamed either way.
Where is the catch? Price? Throughput?
The smaller the capacitor, the faster it can charge/discharge. This tech has only been tested at sizes ~1000x large than the state of the art, and the speed advantage assumes it scales perfectly with the scaling laws. Reality is never that kind, but it might be mostly that kind.
It's still theoretical, though. There might be some manufacturing quirk that makes it not work as well at small sizes. Defects that don't matter now might be huge at that scale. If power requirements creep up, they may kill longevity, which may require them to sacrifice speed... everything has to go right, or it can become a balancing act.
Assuming everything goes great, it's still somewhat more complex than DRAM- more layers. It will certainly cost more than conventional RAM, but with ICs in particular it's very hard to know if that will be 10x more or .1% more.
Or rather, the silicon "in mice" equivalent: in a test sample 1000x the scale, with only hopes and wishes that things won't change too much when they scale down.
All the cool mice these days are running around with memristor-based brain implants. This will be a huge upgrade for them. They'll be able to spend a small fraction of their usual daily time running in the hamster wheel, charging up their symbiote brains.
Also, considering big tech corp. tendency to lockdown stuff, will we need a hack just to do a system reboot?
People have been thinking about that for over 5 decades! This is a part of the history of Capabilities.
Just google, “capabilities computer science"
However, after the Violin Systems boondoggle one may find it significantly harder to find growth capital.
Good luck =)
That'll be really nice if they can get it into production...
The world is not ready for the diseconomies of scale of a second memory type. DRAM and SSD fabs are struggling so why shouldn't they?
Goodbye promising technology!
They can cater to a niche business market to which UltraRAM can add ultra-high value (pun intended) for particular data processing or persistence needs.
One idea that comes to my mind is stock markets. Automated traders took over it and their fight in on the sub-milisecond scale.
Imagine how much an investment bank would pay for UltraRAM if it allows them to process real time data much faster and make ultra-money with it? (again, intended and not sorry about that!)
Conveniently that's also a market where Apple and Google have enough control over the software side to make things work with a new, weird memory scheme (RAM slower than permanent storage, but still needed because of durability).
I don't think you understand your own point. The querying, indexing, and performance bit are tied to the data structures used internally by the database, not the technology used to persist data.
It's not just that people "can't write priginal applications" but that in fact people shouldn't always write their own bespoke single-purpose databases for each application. Getting ACID, MVCC, efficient storage, indexing and backups etc. at the same time is hard, really damn hard. You might get over some of them, e.g., efficient storage, with hardware but there's no free lunch on those topics.
A database is like using a library: You can always write it from scratch (and sometimes you should even) but in 99% percent of cases you should rely on the tried, battle tested existing solution.
Persistence is not yet solved.
Our programming models are currently heavily influenced by the way we store and query data and the underlying registers/memory/cache/storage HW. With PRAM, we can simplify programming models using Persistence Ignorance (PI).
This makes no sense at all. Databases are much more about what data structures are used internally and the high-level interfaces provided to access said data than the underlying technology used to persist data.
There is also the matter of how much data can/needs to be persisted, which is not addressed at all.
I'm sure the company has some tests to prove those claims. /s
Anyone did a HALT test with the M-DISC ?
P.S.: we switched to the metric system.
Endurance figures come from actual testing with their 20um version. Retention is based on looking at the decay over 14h. Since it decays to begin with then plateaus they look at fitting a line to it from some time before the plateau (otherwise the answer is "infinite years") which gives 10^7 hours.
https://onlinelibrary.wiley.com/doi/epdf/10.1002/aelm.202101...
It is unclear from the article if these plots are tests of individual memory cells or a large collection of cells. Any serious attempts would involve an array of cells so a million such graphs can be plotted together etc.