3D Xpoint memory: Faster-than-flash storage unveiled
bbc.com
bbc.com
Here's how to spot BS: first, the density claims are immaterial until they prove yield at a technology node same as dram's today. There are multiple billions investment between 180nm and high yield (and low mask count) 2xnm, no matter how cool is the new memory technology.
Second, speed claims must be explicitly about latency: flash bandwidth is as large as you want it, latency is ~100us (read). Even so, the moment anybody claims that latency is faster than Dram, you know they're feeding the hype and lying to you: dram latency does not depend on the memory technology, rather it depends on the array size. So an 8Gb chip of any memory technology that is fast enough, is likely to be just as fast as dram.
Third: power consumption. Dram's active power is as low as it gets, the memory cell in particular stores information with really little energy. The array interconnect and circuitry consume most energy and a different memory technology won't change that.
The various sort of stacked memory coming into commercial use have the potential to reduce interconnect power and potentially make active RAM power at least slightly relevant. But I think the interesting thing about reducing RAM power consumption is decreasing passive power. Not that I expect anything to come of it for at least another half decade.
This isn't something that appeared out of nowhere. Take a look at this presentation:
https://www.micron.com/~/media/documents/products/presentati...
Page 20 starts the section on "3-D Cross-Point Memory". They even have a 64Mb demonstrator on page 23, fabbed sometime before the 2011 Flash Memory Summit.
Additional info on 64Mb demonstrator: http://investors.micron.com/releasedetail.cfm?releaseid=4672...
Additionally, in early 2014, they stopped selling PCM modules, stating that "Micron's previous two generations of PCM process technologies are not available for new designs or technology evaluation, as the company is focused on developing a follow-on process to achieve lower cost per bit, lower power and higher performance."
Job description details mentioning PCM, chalcogenides, and "cross point technology":
https://www.linkedin.com/jobs2/view/12292797?trk=job_view_si...
And as far as process, as early as 2013 (2013 fall analyst conference handouts), Micron had PCM on a 45nm process and was listing 2xnm as next node. If they've gotten together with Intel and announced this, they have already reached 2xnm and/or beyond or are certain in their capability to do so.
Because of the way flash memory is organized into pages and blocks, latency is very workload specific. Can you point out where anyone claimed it was better than DRAM latency?, the BBC article says "DRAM chips are still faster than 3D Xpoint, but the difference is much smaller than when compared with flash".
And power consumption. This is PCM. Power consumption by the array when not being read/written is zero.
>While they did not specifically state it, it looks to be phase change memory (edit at the Q&A Intel stated this is not Phase Change).
Source: http://www.pcper.com/news/Storage/Breaking-Intel-and-Micron-...
"Relative to phase change, which has been in the market place before and which micron itself has some experience with in the past again this is a very different architecture in terms of the place it fills in the memory hierarchy because it has these dramatic improvements in speed and volatility and performance"
They didn't say that it wasn't PCM, and they didn't say that the technology was different. They said that the architecture was different than the PCM that Micron produced before. Which would be the cross-point organization and the all important selector element. They mention that the cell works my a "property change of the material", or a "bulk material property change".
Go look up Micron's recent patents and applications. This is PCM.
I do find it very interesting that they are bending over backwards not to call out the memory technology and specifically answering a question directly regarding if it is actually PCM.
Re the latency caveat I was suggesting BS spotting rules. You're right that this article did not claim the memory was faster than dram.
RE power consumption, it's true that it you don't use it an NVM will not consume power, yet that is not what people refer to when talking about power consumption. Can you quote a technology paper that shows lower than dram, or lower than flash for that matters, active power consumption? (The proper unit is W/GBps here)
And yes, array size affects latency but so does the underlying process technology (much more so across technologies).
-refresh power is on the order of milliwats per Gigabyte
-memory access is on the order of 0.1W per GB per second
For refresh power, let's face it: it's puny. Regarding the active power, I have yet to see anything (off chip and this size) that is better than that.
P.S. I did say active in my original post :)
watt != energy
-refresh period is more like 32-64ms
-refreshing is way (way) cheaper than streaming data out since you refresh an entire word line, and you do not pay io cost, which is a significant fraction of the cost of accessing dram.
Wait, what? You mean they are making us wait 10ns for row open and 10ns for row close just for shits and giggles? Or did you mean to say "bandwidth"? I'm confused (genuinely, not sarcastically).
I'm only half kidding. As a complete industry outsider it doesn't seem ridiculous to think that SRAM could do to DRAM what SSDs did to HDDs. Alas, I have no idea what the economics + trends are and I never did more than dabble in semiconductor engineering/physics so I will only ever find out after the fact.
SRAM needs 6 transistors per bit, DRAM 1transisitor+1capacitior. SRAM just doesn't scale and it's very expensive.
http://www.kitguru.net/components/ssd-drives/anton-shilov/sa...
If the only thing standing between SRAM and DRAM were a constant factor of <6, DRAM would already be history.
The most convincing explanation I've heard is that caches are so damn good at hiding latency that getting rid of row open/close just doesn't matter. A few minutes of googling suggests that they often run at a 95% hit rate on typical workloads and a 99% hit rate on compute workloads. You would still need a cache even with main memory as SRAM to hide transit-time, permission checking, and address translation latency, so SRAM main memory wouldn't actually free up much die space, it would just make your handful of misses a bit faster (well, it would free up the scheduler / aggregator, but not the cache itself). The reason why I called this one "most convincing" rather than "convincing" is that even with a 99% hit rate a single miss has such atrocious latency that it would seem to matter.
And yet I cannot purchase even a 1GB SRAM stick.
I don't see why that equates to "just doesn't scale". Can you elaborate?
BS: "going into production this year"
Possibly not BS: "going into production this year at 2x nam technology node with a y Gb part"
Many nvm memory technologies, including mram and PCM have been "in production" for quite many years, with << 1Gb parts on old technology nodes, that is.
How? Because its bit addressable and persistent. Together this makes it much simpler to implement some durable storage . We don't need the log structure good for NAND block erase issue. We don't need to worry about the flush cost compared to HDD (and this ones even faster than NAND). It would be simple to batch write the data to a slower storage if the Xpoint memory fills up.
You can design databases that keep the hot data in memory and merge the result with older disk storage and this would allow a lot of batching for efficient processing and storage. But given its a database/transaction you need durability and that makes thing that much complicated. There are still lots of problem to solve when you cross the limit of a single machine but the single machine limit can get a lot larger for a lot of problems.
I do wonder how much it would be slowed down by the kinds of sophisticated error correction SSDs are now relying on.
Bit addressability has absolutely nothing to do with endurance. NOR flash is bit addressable but suffers from the same endurance limitations as NAND, because they're fundamentally the same kind of memory cell, just connected differently.
In practice, if you burn out any one bit, you need to retire a chunk of the array at least as large as a cache line. And it's not likely that you'll actually be able to directly hammer a single bit, because the endurance is still low enough to require ECC.
Only in some bizarro world where "three orders of magnitude increase in performance" also means "we'll write three orders of magnitude more data into it".
Loads are about use cases, not about how fast you can fill a disk. If my company produces 1TB analytics info per day it wont suddenly produce 1000TB just because I can write to the disks we buy faster.
Of course being able to fill it faster also opens up some new, more heavy, use cases. But for any existing use cases, we'd be writing the SAME data volumes we do now, just 1000 times as fast and with 1000 times the endurace.
And even if we write 100 the data we do now, we still get 10 times the endurance.
We have two models of storage - volatile working storage, and things that simulate disk drives. It's not clear what to do with persistent randomly addressable storage at DRAM speed. Having to go through an OS and a file system to access a few bytes kills the performance advantage of such devices. Making the device look like RAM makes it too easy to mess up. We need something in between, probably with processor support to allow controlled access without going through the OS for each access.
The great thing about RAM being volatile is that you can reboot and clean up your mess. With persistent storage, things can go gradually downhill.
That probably is not the best one can do, but it is simple, and may be fairly easy to get ‘right’ if you put the code handling the file system part into a secondary kernel address space. That keeps the part that can mess up the file system’s metadata small.
I can see a scenario where opening a file for writing works just like mmap[0] with MAP_PRIVATE, i.e. you get blockwise copy-on-write, except everything will persist as if everything on your filesystem was under a VCS like git.
I reckon just like volume management, encryption and snapshotting has steadily been folded in to the filesystem, so will the VFS and page cache.
"BTT is a library that converts a byte-accessible namespace into a disk with atomic sector update semantics (prevents sector tearing on crash or power loss). ... BLK is a driver for NVDIMMs that provide sliding mmio windows to access persistent memory." https://lwn.net/Articles/649588/
>1000X faster, 1000X cheaper!
Then "No, you can't buy one right now, but you will be able to do it "soon". And, no it won't actually be 1000X faster neither 1000X cheaper because blah blah blah..."
I think that's sufficient to not put it into the vaporware category.
Performance questions are more valid.
Then it mentions 10x more performance with a PCIe/NVMe interface.
Still all good things, but 1000x is probably more marketing than reality.
Source: http://www.intelsalestraining.com/infographics/memory/3DXPoi...
clearly, something, somewhere is causing progress to happen, despite your inexplicable inability to see it.
the real problem is people keep making software that gobbles up all these gains.
Second, those 1000x faster and 1000x cheaper. I've been around a few decades, and we DO have 1000x faster and 1000x cheaper stuff now.
CPUs are 1000x the speed of 1980 CPUs.
1GB of RAM would cost you a house back in 1990.
A 1TB disk would cost you half a skyscrapper plus take 2-3 houses to house back in the day.
MRAM has similar performance to SRAM, similar density to DRAM but much lower power consumption than DRAM, and is much faster and suffers no degradation over time in comparison to flash memory. It is this combination of features that some suggest makes it the “universal memory”, able to replace SRAM, DRAM, EEPROM, and flash.
https://en.wikipedia.org/wiki/Magnetoresistive_random-access...
All that being said, there are people who make the exact same claims about RRAM as you quote for MRAM, which 3d-xpoint appears to be.
Also note that you can buy MRAM parts right now, which are replacements for battery-backed SRAM, and are more radiation resistant than SRAM. Densities are fairly low though.
Where does that come from? From everything I've read is that its structure is fairly analogous to DRAM, "simply" replacing the capacitor with the magnetic tunnel junction, which has its component layers stacked vertically, thus not really taking up any extra space.
> scaling it down to small process nodes is still problematic, even with spin-torque transfer.
yeah, that's the main gist I'm getting too from following the news.
> people who make the exact same claims about RRAM as you quote for MRAM, which 3d-xpoint appears to be.
At least the xpoint incarnation still seems to be slower than DRAM according to that article, while MRAM is being offered as SRAM/battery-backed DRAM drop-in replacement.
> Where does that come from? From everything I've read is that its structure is fairly analogous to DRAM, "simply" replacing the capacitor with the magnetic tunnel junction, which has its component layers stacked vertically, thus not really taking up any extra space.
I did some searching; older references show an 8-12F2 size, for e.g. the Everspin parts. Grandis claims a 6F2 size which is indeed comparable to DRAM.
>> people who make the exact same claims about RRAM as you quote for MRAM, which 3d-xpoint appears to be.
> At least the xpoint incarnation still seems to be slower than DRAM according to that article, while MRAM is being offered as SRAM/battery-backed DRAM drop-in replacement.
Right, the product they are claiming they will manufacture next year is slower than DRAM and less dense than flash. (Frustratingly I couldn't find a reference for if they are talking about latency or throughput when they say "slower"; it makes a big difference for which applications will be hurt by the performance mismatch).
However, there doesn't appear (yet) to be a fundamental reason why resistive ram must always be slower than DRAM, nor a fundamental reason why they couldn't do MLC tricks with it, so you can't say all RRAM will be slower than DRAM and less dense than NAND.
The largest devices I've seen are 4Mb (512kB), which isn't a lot by PC standards, but is dead handy for embedded.
browses
Oho, they make one which is pin-compatible with SRAM chips!
http://www.fujitsu.com/global/products/devices/semiconductor...
Price looks like $15 a single part to $12 for ten. Ouch. If we assume that it's about $20 a megabyte in bulk, a gigabyte would cost about $20k. This is, price-wise, equivalent to:
- RAM, in 1996
- Spinning disk, in about 1988
- Flash, in about 1998 (extrapolation, my chart doesn't have nay data before 2004).
Ref: http://www.jcmit.com/memoryprice.htm
Today's statistics have been brought to you by Late, and Tired. Enjoy.
[1] http://newsroom.intel.com/community/intel_newsroom/blog/2015...
[2] http://ark.intel.com/products/84685/Intel-Xeon-Processor-E7-...
crossbar is a startup in the field, their site gives plenty of details about the tech.
Lets just forget about this "memristor" PR-nonesense. HP did as well. (Their activities seem to be dead)
[1] http://hardware.slashdot.org/story/08/04/30/211228/memristor... [2] http://hardware.slashdot.org/story/12/07/25/2127229/the-hp-m... [3] http://hardware.slashdot.org/story/14/11/25/2027220/how-inte... [4] http://hardware.slashdot.org/story/15/03/26/2037221/micron-a...
https://www.google.com/search?q=site%3Amicron.com+%22cross-p...
>By contrast, 3D XPoint works by changing the properties of the material that makes up its memory cells to either having a high resistance to electricity to represent a one or a low resistance to represent a zero.
Which sounds a lot like memristors: https://en.wikipedia.org/wiki/Memristor to me. If it's memristor memory, the main advantages are that it's cheap, uses hugely less power than flash, takes up less space, and has far lower latencies.
We're probably a long ways from replacing DRAM with memristors because it's still much higher latency, but if this stuff scales up well you could do something like put the whole page file on it and get even faster loads than current top of the line SSDs.
If anything I'm hoping the introduction of a competing new technology will make higher capacity SSDs cheaper.
"3D XPoint works by changing the properties of the material that makes up its memory cells to either having a high resistance to electricity to represent a one or a low resistance to represent a zero."
How does it do that?
Ignore this person. They don't know what they are talking about.
This is phase-change memory. It's been on everyone's radar for the last decade. HP is pushing their titanium oxide memristors. That's a completely different technology.
>While they did not specifically state it, it looks to be phase change memory (edit at the Q&A Intel stated this is not Phase Change).
Source: http://www.pcper.com/news/Storage/Breaking-Intel-and-Micron-...
"Each megabyte of 3D Xpoint will certainly be significantly cheaper than the equivalent amount of Ram. And the new technology has the added advantage of being non-volatile, meaning it does not "forget" information when the power is switched off. But, unfortunately it is still not quite as fast as Ram, and some - but not all - applications need the extra speed the older tech provides."
What can be read from the announcement: it's not as rewritable as DRAM: only up to 1000 more writes than NAND. It's also slower than DRAM. But it seems better than NAND in all aspects.
Err, they mention several times the density of NAND.
For a lot of users 16GB might not be worthwhile, 64GB might actually allow some people to not use a SSD at all, so it I think they will have to provide 32GB of storage. That seems obvious enough that the fact it's not stated in the article seems somewhat concerning.
If they release it next year and it turns out well, I can certainly see Apple pulling such a move for the MacBook (Air).