MIT Takes Multicore in a Different Direction
top500.org
top500.org
Poster here: http://livinglab.mit.edu/wp-content/uploads/2015/12/poster.p...
Paper here: http://livinglab.mit.edu/wp-content/uploads/2015/12/2015.swa...
Slides: http://livinglab.mit.edu/wp-content/uploads/2015/12/2015.swa...
Take it for what it's worth.
I'm not versed in the topic, but if that's true, then why wouldn't those numbers be worthy?
This is fairly standard for computer architecture research in academia: the results are based on a cycle-accurate model that simulates what the machine would do when running particular code. The field has a standard set of benchmarks (e.g. SPEC CPU2006 for serial code, and a bunch of different ones for parallel code) that are run on these simulators and are well-understood. The same is true for architecture teams in industry: they start with simulators and (a lot of) benchmarks before any silicon ever exists.
The reason for all of that is that actually fabbing a chip is prohibitively expensive and insanely work-intensive. A real chip has probably millions of lines of Verilog RTL, and a lot of custom layout to get anything reasonably performant (then > $1M for a mask set, multiple weeks or months to get the chips back from the fab, etc). In contrast, a good simulation model, worthy of publishable results, of an SoC with out-of-order cores, a cache hierarchy, and DRAM, is somewhere between 10K and 100K LoC of C++; the model for a new proposed feature is maybe a few KLoC of code on top of that. Once the simulator exists, a grad student can try out new features fairly easily. It's also much more analyzable and instrumentable: a chip is mostly a black box, modulo whatever debugging features you build in, while a simulator can easily dump a "pipe trace" of the pipeline state every cycle. The field has invented lots of different tools and ways of visualizing data to get a sense of what goes on inside the machine and what bottlenecks exist.
So basically, it's a software model, and the software model is much more informative, and easier to tweak and iterate on, than silicon while being "good enough" for trustworthy results.
But as you imply, that's mostly done by hand in C++ with frameworks?
This shortcut is possible because of a common technique of splitting functional and timing details apart: the functional emulator simply runs the instructions in a big interpreter loop and tracks the machine state as, say, QEMU or Bochs would, while the timing model is just cycle accounting given the instruction stream. In contrast, when you build a model in RTL, you're actually doing all the work that industry microarchitects do: you need to get right all the details of (say) speculative execution, or cache tag matching, or whatever, because your microarchitecture is implementing the code execution directly. That's a lot harder to do!
People do sometimes write RTL for their proposed microarchitectures, but that's usually done for power or timing (clock speed / critical path) results. And they usually model just whatever new thing (prediction table, synchronization widget, cache eviction logic) they propose, rather than the whole chip.
That being said, it's generally easier and cheaper (good verilog implementations aren't free) to use a general purpose language for what your describe.
But, I agree with your point that a simulator is a much easier step that gets most of the data.
(All that said, seeing a real chip at the end of a project must be a heck of an exhilarating feeling of success...)
and the actual paper introducing Swarm: http://livinglab.mit.edu/wp-content/uploads/2015/12/2015.swa...
Feels like another hardware direction that will require a lot compiler work to truly take advantage of. Anyone else getting Itanic vibes?
It looks like most of that ability is meant to be handed straight to the programmer, not the compiler.
Seems like this would be a more incremental option than requiring a full new architecture to be designed.