Crossbar Resistive RAM stores a terabyte on a chip
venturebeat.com
venturebeat.com
The write performance sounds good, but the read performance seems very low. Also I don't know where they get 7 megs per second for flash. Sounds like the picked the worst performer to compare to.
I'd be more impressed if they had something on the market.
Flash chips are generally multiplexed and behind a smart controller that compresses and and caches data, which can greatly effect the throughput numbers.
Crossbar can do the same thing, in which case the same number of chips provide higher throughput. Read latency is not parallelisable, of course, so there'll be no way for flash to catch up there without making fundamental advances.
That is, if all these numbers are correct.
no, reducing RAM power 20x does not extend battery life of devices to weeks. Anything with a screen or active antenna will likely not see much of a difference
Battery life is affected by such a multitude of components that it's hard to extrapolate savings in one component to the battery life of the complete device.
> could also perform its storage functions at 20 times lower power
> non-volatile
> 10 times the endurance of NAND flash chips that it could replace
If this is true, wow, but I have a hard time believing it.
I remember an article stating that embedded flash controllers typically have a 20MB/s speed limit. This in turn meant that wireless speeds beyond 200Mb/s were mostly worthless, because the devices simply could not utilise any more. If someone can find the article again, I would be delighted.
However: if the reality with RRAM is even half as good as the claims go, then it would certainly encourage the hardware vendors to invest in somewhat better controllers. After all, what good is a new, hyperspeed storage medium if you can't access it any faster than the old one?
The whole premise of taking NAND from 25nm to 19nm (for instance) is to fit more floating gates in the same area. You can take that as a smaller die, or as more bits on a slightly larger die than the previous generation.
Die size is indeed a major factor on cost. For a given technology (litho node + process, e.g. # and type of steps), the cost to process a wafer is fairly constant regardless of die size.
If you shrink a die size, you fit more die on a wafer. Additionally, yield goes up (given an independent manufacturing defect density), and especially for large die, the tessellation around the edges has a major impact.
Indirectly, even testing is related to die size, in that there is a limit to tester parallelism, and more gates means more time and more combinatorial patterns to test, e.g. for stuck-at testing.
There are of course non-linear costs in packaging and package-level testing and elsewhere.
[1] Which dominate with memory since production runs tend to be large and the regular patterns make for less design investment than, say, a CPU.
"[Crossbar's CEO] could not estimate the price of a 1TB RRAM module, but said it will cheaper than NAND flash partly because RRAM is less expensive to manufacture."
http://www.pcworld.com/article/2045926/startup-crossbar-pits...
If we're talking about rendering beautiful complex worlds with post-processed special effects, I disagree.
The only reason to prefer the GPU is because it confers an advantage over the traditional CPU+RAM combination. There's nothing inherently special about a modern GPU. The GPU is a sequence of actions and abilities encoded into hardware, e.g. the ability to automatically perform various kinds of texture filtering transparently to the game developer.
Since the GPU is hardware, and since hardware is less flexible than software, a graphics programmer would always prefer a software-based pipeline to a hardware-based one. The reason hardware pipelines are preferred is strictly because their advantages outweigh their disadvantages. Typically, using a GPU enables graphics programmers to create renderers which are 10-100x more efficient than software-based renderers, so the added flexibility of a software rasterizer tends to be forgotten in the face of massive efficiency enabled by the GPU.
The GPU primarily became popular because (a) it offloaded part of the computation from the CPU to dedicated hardware, freeing up the CPU for other tasks like game logic, AI, and more recently physics computations (though nVidia is trying hard to convince developers that hardware-accelerated physics is a viable concept), (b) GPUs increased the amount of available memory, and (c) GPUs dramatically increased the throughput (memory operations per second) of graphics memory.
Memory latency plays a key role in many modern graphics algorithms, such as voxel-based renderers. It's often the case that an algorithm needs to repeatedly cast rays against a voxel structure until hitting some kind of geometry. Therefore, within an individual pixel of the screen to be rendered, this type of algorithm can be hard to parallelize because typically the raycasting can't be broken up into parallelizable steps. It typically looks like, "While not hit: traceAlongRay();" for each pixel, each frame. I.e. this algorithm can only trace one section of the ray at a time before tracing the next.
That raycasting algorithm is memory-latency-bound because it completes only when it finishes looking up enough memory locations that it detects the ray has intersected some 3D geometry. In other words, by reducing memory latency by 2x, and assuming memory bandwidth is sufficient, then this algorithm will complete twice as fast. This means instead of 24 frames per second, you might get 48 frames per second.
So, all that said, if it becomes common to have 1TB of regular RAM with the latency and bandwidth traditionally offered by GPUs, along with a surplus of available CPU cores to offload computations to, then software renderers will once again become preferable to GPU renderers. A software pipeline will always be more flexible and easier to maintain than a hardware pipeline, simply because the featureset of the software pipeline isn't restricted to the capabilities of the videocard hardware it's executing on. It's also easier to debug and maintain.
All of that means that it'll be easier for art pipelines to produce more complex, more immersive visual experiences than at present. But replacing the traditional GPU-based renderer with a CPU-based software renderer will only be practical if there's a major advance of RAM technology in the future, because current RAM tech can't match the memory bandwidth / latency of a modern GPU. Hence, any major developments in the area of volatile memory tech will be extremely interesting to graphics programmers.
AMD Kaveri CPU will support GDDR5 apparently. Furthermore, DDR4 is set to appear soon, with significant improvements in bandwidth / latency / power.
So RAM tech may be lagging behind, but by 2014 or 2015, we'll start to see some interesting advances in RAM.
If that won't change games I don't know what can.
Besides GPUs will obviously also use this technology if it really works.
And the bandwidth between the GPU and the GDDR on the graphics card doesn't have anything to do with the PCI bus either, except to a small extent when synchronizing with the CPU or initially loading textures or whatever.
Plus since we're talking about cutting-edge gaming, integrated graphics is irrelevant (and so are APUs).
When you are on a timescale of nanoseconds even electricity's speed can be slow when being compared to something like intel's l4 cache which is on-die.
For any type of volatile memory that latency will exist until mobo designers move the ram closer to the cpu or adopt optical interfaces between parts.
For some cool calculations that can put things in perspective take the speed of light as 299 792 458 m/s and take the time of a nanosecond as 10^-9 seconds to get ~0.3 m/ns for light. That means that for every third of a meter the dimms are away from the cpu means a constant 1ns delay in terms of latency.
http://www.engadget.com/2013/08/05/samsung-ships-first-3d-ve...
In this case, that means a reputable tech reporter from a reputable publication has to say "this is the real deal" before these guys can be taken seriously.
(I'm not doubting that they have any tech, it's just that they seem to be promising something with no downsides, which in most cases means a marketing team has spun out of control.)
Does it? It just sounds like a normal memristor to me.
http://www.engadget.com/2010/08/31/hp-labs-teams-up-with-hyn...
The last I heard was second quarter 2014 by HP, that's been the release date for the last year or so. Supposedly they could release them now, but they are trying to time the market for business reasons.
The phrasing in the article is a little wishy washy, but it sounds like this company is going to production with a TB-per-IC-density non volatile storage product. If true, that's huge news.
Anyone have an idea if this is related or unrelated.
http://www.ece.rochester.edu/users/friedman/papers/ChipEx_11...
Like like a regular transistor can be used for computation or ram.
For some strange reason, the tech press took this to mean that a memristor array can be dynamically reconfigured to act as logic or ram in the same device. This is mostly false -- while yes, this kind of devices can be built (with transistors, they are typically called field-programmable gate arrays, or FPGAs), this requires you to be able to reconfigure not just the gates, but all the wiring that goes into them, meaning that a reconfigurable array has to be more than an order of magnitude larger, and a few times slower than a non-reconfigurable one.
You never want to use an FPGA for ram because you can get proper, non-reconfigurable ram for (way) less than tenth of the cost. You never want to use a fpga for logic if you can afford an asic, because you can be several times faster with hard-baked logic.
Memristors will mean nice FPGAs that retain state on power down, but they will not mean a revolution of dynamically reconfiguring devices.
Oh, and memristor logic is slower than transistor logic (you can do with less gates, but the gates switch slower) so the ability to use memristors for logic elements will initially mostly be a win in memory devices where the necessary logic to manage the device can be made out of the same structures the device is made of.
(venturebeat sounds like a investor magazine to me)
Also, that "Crossbar Chip Design" graphic made me chuckle. It is nearly meaningless by itself, and is placed so far away from the context it's discussed in, it almost remains that way unless you're trying to connect it.
I'm not saying it won't happen, but on the journey from idea to millions of units, a test chip is just the beginning, and in the meantime, NAND is moving.
Look at how slowly NAND has replaced rotating magnetic storage: it's been around since the 80's. Decades of iteration have brought it to a place where it's compelling for non-niche use cases.
http://thedoghousediaries.com/5275 (Amazing Scientific Breakthroughs)