Making the Case for Feature-Rich Memory Systems (2016) [pdf]
cs.utah.edu
cs.utah.edu
The GPU is the only recent example where I've seen people really willing to rethink how they work to get (very large) performance gains. But I'm not seeing that here. For the examples I can think of, it seems like it would be easier to install more servers with modest amounts of RAM, rather than having a smaller number of servers with lots more RAM and in-memory processing.
Of course, maybe I'm just not thinking of the right examples.
This is the question to ask people when they say "the von Neumann architecture is a bottleneck".
That being said, I think a strong contender is the dataflow computing model. The advent of the GPU already heavily punishes control flow, and workloads are being increasingly constrained, hence the possibility of a real world success for dataflow machines.
There are a lot of options for in-between options as well. For example, mallacc by Harvard modified to reside inside the memory itself could be very useful: http://www.eecs.harvard.edu/~skanev/papers/asplos17mallacc.p...
We see the gains in productivity and the reduction in cost – at least according to some measures –, but it has inherent inefficiencies. Inefficiencies like simple applications using far more CPU and memory than they're due to.
This perhaps makes people think again what would be possible if all the abstraction (analogous to general-purposeness of hardware) were wiped away and what has been possible in the past using far fewer resources.
Another example is the ongoing efforts to improve their support for value types.
A 20 years delay to catch up with what Common Lisp, Eiffel, Modula-3 and Oberon variants already offered in those days.
I suspect we'll see the same thing for value types: it's not going to be quite the easy win it seems. Even in C++, it's easy to create accidental footprint explosions with templates and lose performance to excessive copying with value types. And C++ allows mutable values, which are out of fashion now so Java won't allow them ...
I was curious if the article would mention memory chips that knew how to do bulk memmoves by themselves. Does anyone know if memory subsystems can already do that? If you issue a memmove() for, say, 3kb of memory, does the CPU still have to read it all into the cache and then immediately write it out again, or is there some way the CPU can signal to the DRAM chips that they should do the copy themselves? Fast copying of memory would be useful for GCs.
The CPU core has more bandwidth than the IMC anyway, so there would be no speed-up from adding this complexity to the IMC (it would not only need to perform the operation, but it would also need a way to maintain cache coherence and communicate with the issuing CPU, none of which is a problem if you just do it in the core). It might not even save power.
No quite, the long term idea is to bootstrap OpenJDK as much as possible, making it a meta-circular VM.
Check Project Metropolis and SubstrateVM projects.
Also the commercial JDKs with AOT support, do it for deployment scenarios where JIT isn't desired, e.g. embedded devices.
The OS gets the exception. There are always (usually priviledged) instructions to read and set the memory check bits. It would be darn hard to write memory diagnostics without them, or sometimes even to boot a machine that powers up with random bits in the memory.
If I understand what the author wants, (and I have my doubts that the author understands what they want), then they are asking for an OS feature, not a CPU feature, and simply want the exception to bubble up to user space.