ComputeDRAM: In-Memory Compute Using Off-the-Shelf DRAMs (2019) [pdf]
parallel.princeton.edu
parallel.princeton.edu
Our novel operations function at the circuit level by forcing the commodity DRAM chip to open multiple rows simultaneously by activating them in rapid succession.
That reminds me of this:
https://www.linusakesson.net/scene/safevsp/
It's worth reading the long explanation there, and the HN discussion at https://news.ycombinator.com/item?id=11845770 , but the critical part is this:
In short, one memory cell gets refreshed with the bit value of a different memory cell.
What I find more amazing is that this effect has probably been noticed for as long as DRAM existed and treated as a bug (like above), and it took several decades for someone to think "that could be useful!"
https://insidebigdata.com/2019/07/26/machine-learning-with-m...
https://semiwiki.com/forum/index.php?threads/survey-paper-on...
I tracked these guys because we were doing automata in s/w (the https://github.com/intel/hyperscan project at Intel). There were some aspects of their methodology I didn't like (notably, they tended to run benchmarks that spewed matches, slowing Hyperscan down, while turning off matches on the h/w card!) but I found them v. interesting and they had some really thought-provoking stuff.
”The Elxsi 6400, built in the early 1980s could do logical operations to memory using the off the shelf drams of the day. Designer was Harold (Mac) McFarland, of PDP-11 fame.”
I googled for more info, but didn’t find info about that. There’s a Wikipedia page (https://en.wikipedia.org/wiki/Elxsi) that confirms the McFarland part, and I found the system architecture (https://amaus.net/static/S100/elxsi/systems/Elxsi%20System%2...), which is interesting (instructions for multi-precision ascii addition and subtraction, for example, I would guess for supporting COBOL), but nothing about those DRAM tricks. Does anybody have more info?
Anyone see anything that contradicts this?
On top of this it would be really awesome if this attached 'ram compute block' was re-programmable like fpgas or a simpler higher level construct. then we could essentially do all the 'parallel & dumb' computations in ram itself so CPU would mostly be for control and branch code.
It would cost next to nothing in terms of chip area to have a simple "controller" CPU in each DRAM chip that can issue vector instructions without needing to have its hand held.
The problem is the instruction set: where do you stop? Do you just have simple logic operations, or a full set of floating point operations? Do you have flow control? Stacks? Etc...
Code would either have to be written in the "full featured" language as well as a cut-down language similar to CUDA, or the embedded little chips would have to be full CPUs.
The other issue is cache coherence. It's possible to have designs where either the little DRAM CPUs participate or they don't. Both have big advantages and big disadvantages, and neither is easy.
I suspect that this is going to start turning up in GPUs before CPUs, or possibly for HPC applications before general purpose computers.
Spread ALUs and simple control units across RAM cells, there's little needed because the RAM is its own registers. Some distant big control unit will send out instructions to the processing units, and orchestrate I/O. A bit like GPU but with a different set of constraints. Likely it could be made compatible with OpenCL or CUDA.
regarding vector instruction explosion, this why I left a remark around programmable fabric (which does have to be super fast reconfigure). this way you could morph a bunch of logic blocks into whichever flavor you want. btw this is also not a first either, companies like Stretch & Mathstar have tried to do similar re-programmable fabrics & more recently altera had done re-programmable fabric using parallel/gpgpu languages like OpenCL. one good thing with programmable fabric in this context is that there is not an immense pressure to fit logic in a cycle budget because you can always claim a certain vector instruction takes X cycles to complete without effecting simpler operations taking Y (<< X) cycles.
cache coherency issues notwithstanding, you are right about it turning up in GPUs first, simply because as the target resolution scales past 4k & 8k, VR etc it would be imperative to do a lot of similar parallel operations on huge chunks of memories and memio b/w would be the biggest bottleneck there. this could mostly alleviate that.
what I am unclear about is how does putting programmable fabrics like this impacts DRAM yields?
Wherein he explains something like (up to) "AVX-32768" at about 36 minutes and 30 seconds ( https://youtu.be/5Z7cmyYakAw?t=1594 ) besides the latest AVX-512 with no downclocking whatsover at 2,5GHz.
This is not exactly "In-Memory" but since it's all in the SOC, connected to ring bus with low latency, and large cache...?
All in all very impressive chip, especially when considering the technology they had made their samples in. You can't buy it yet, unfortunately.
https://centtech.com/ai-technology/
They did an AMA on reddit two months ago https://old.reddit.com/r/hardware/comments/ep0lf4/x86_cpu_de...
https://fuse.wikichip.org/news/3099/centaur-unveils-its-new-...
https://fuse.wikichip.org/news/3256/centaur-new-x86-server-p...
Cha, Cha, Cha!
I don't have the expertise to know...
If they don't -- then what would be needed to be done to the memory controller (in this case a custom FPGA memory controller) to make it Turing-complete?
Also... even if it isn't Turing-complete, then probably whatever functionality is missing could be implemented by the FPGA -- although at the probable cost of a memory round trip for those instructions, right?
In other words, you could probably use this for mixed-mode, hybrid, Turing-completeness via additional FPGA instructions, even if the operations on RAM aren't Turing-complete in and of themselves -- or am I missing something?
All of this sounds very promising!
Great concept, great paper -- hope you get well funded for your next round of research!
The CM-1 could do conditional execution: you tell the 64k processors to add but only those with a give flag set to 1 will actually do it while the others will execute a NOP. It is possible to simulate this on this design but it would be rather awkward, just like their 1 bit addition is awkward compared to the one clock equivalent in the CM-1.
Transpose. (Take a bit-parallel array in row 0, columns 0,1,... and place it (bit-serially) in column 0, rows 0,1,... - and preferably do the same for row 1 to column 1, row 2 to column 2, etc simultaneously.)
Random access (Take a bit-serial index N in rows 1,2,... and load or store row 0 to/from row N.)