More efficient memory-management could enable chips with thousands of cores
news.mit.edu
news.mit.edu
The technique kind of reminds me of Jefferson's virtual time (and time warp), which rather exists in a distributed simulation context. Virtualizing time to manage coherence of reads and writes is a very good idea.
But my knowledge is a bit out of date, since GPUs are actually including caches now, and what they can do seems to evolve rapidly.
http://www.nvidia.com/content/pdf/fermi_white_papers/nvidia_fermi_compute_architecture_whitepaper.pdfThe idea looks interesting (I have not yet read the actual paper), but calling it «The first new coherence mechanism in 30 years» is ridiculous.
In fact, you could start by looking at the Related Work section of this paper itself. The version at https://people.csail.mit.edu/devadas/pubs/tardis.pdf is better. It is quite telling that the paper does not make the outrageous claim of the title of MIT's press release.
O(log N) of memory overhead per block is nothing new. There were commercial systems in the 1990s that achieved that (search for SCI coherence). Note that there are other overheads to consider (notably latency and traffic).
This paper is very interesting and looks sound, but MIT's press release makes it look silly.
Excuse me for not even trying to make a summary of the last 30 years of research in this field.
The quote from the article
> MIT researchers unveil the first fundamentally new approach to cache coherence in more than three decades
is more hedged ("fundamentally"). I don't know if anyone in this thread has the expertise to evaluate that judgement.
One thing the APIs depended on was atomic-add. I tried to get to 10 millions msg/second between process/thread withing a SMP CPU group at that time. For 10 millions msg/s, the APIs had 100ns to routed and distributed the msg. The main issue was none-cache memory access latency especially for &atomic-add variables. The none-cache memory latency was 50+ns on DDR2 at that time when I measured that on 1.2GHz Xeon. It was hard to get that performance.
I even considered adding and FPGA on PCI/PCIe which can mmap to a physical/virtual address that will auto increment on every read access to get a very high performance atomic_add.
If that same FPGA is mapped to 128,256,1024 cores, one can easily build a very high speed distributed sync message system. Hopefully for 10+ millions / second for 1024 cores.
That would be cool!
No, here we simply increase the number of upstream servers managed by a frontend like nginx.
In special cases you can get better performance with more, but less beefy cores (graphics is a prime example), but in general a few powerful cores performs better. The main reason is because communication among cores is hard and inefficient, so only embarrassingly parallel programs work well when divided among many cores. Plus the speedup you get from parallelize a program is minor in most cases. See Amdahl's law[1] for more on this topic.
Also, I'm not an expert in this area, but I have some familiarity with it. So hopefully someone with a bit more experience can come and confirm (or refute) what I've written.
Having lots of cores isn't likely to matter much for mobile users, simply because most mobile apps are neither optimized for parallelization nor CPU-hungry in the first place. If 4 cores are good enough, it doesn't matter if you add hundreds more. That said, there may be specific application that might benefit, such as computer vision.
This is good news for actually scaling caches with cores though, given how much die space is actually used for cache compared to cores.
Most likely maybe we will have a few speedy main cores for non parallel code like mobile Arm big little and then the thousands slower cores in parallel. Will there be transparent cloud execution where workloads can migrate from local cpu out to massive cloud cpu and back?
Kind of like GPU but for main CPU?
I suspect that an Erlang-like language running on hardware specifically designed to support it could achieve tremendous scalability.