Also rings still don't scale, so they're really going to have to get more clever. I wonder what the perf hit for the ring is going to be.
Also rings still don't scale, so they're really going to have to get more clever. I wonder what the perf hit for the ring is going to be.
One option is to move closer (but probably not actually adopt) something like a scalar operand network used on some of the tile architectures.
For what it's worth, most of the many-core chips that do exist keep it fairly simple. Ambric's MPPA had a mesh with configurable routing (http://www.nethra.us.com/technologies_mppa.php), GPUs have a combination of pretty straightforward hierarchical interconnect and rings, and Azul's interconnect is a tree-like structure as well.
One thing to consider is that, while more complex topologies can buy you something in a supercomputer, give you more bisection bandwdith, and better support point to point communication, it has turned out to be pretty hard to actually write correct programs in that style (lots of arbitrary peer-to-peer communication); this partially motivates simpler interconnects, where the programmer can more easily reason about what's going on. Another concern is that topologies that work great in a server room (3D) are nigh-unroutable on a chip (2D). The wraparound links in a torus are a great example of this; Blue Gene/L used a torus to great effect (http://www.google.com/url?sa=t&source=web&ct=res&...), and the wraparound links drastically reduce worst-case and average point to point latency over a mesh, but those links mean giving up a large fraction of a metal layer on a chip, as opposed to a long cable in a machine room.
Also, I would hazard a guess that the workloads these chips are intended for are mostly multiprogrammed, not multithreaded; if they don't share much data, the network's bisection bandwidth is not as much of an issue as the off-chip bandwidth.
Regarding scalar operand networks as found in Raw/Tilera64: The specific idea (register-mapped networks) probably changes the programming model too much for Intel to adopt directly, but the Single-chip Cloud Computer (SCC) linked above, presented at ISSCC 2010, uses something sort of similar for message passing over their 2D mesh. The communication channels are memory-mapped rather than register-mapped, but the mechanism is similar.
I think they'll drop the rings soon. But it will keep them going until they figure out how to solve their interconnect problems.
I agree with you that the SON as on RAW/Tilera is unlikely to make the leap to Intel. I just hope they'll move towards that direction. Though clearly they may also chose to move in a completely different direction, but they will need a more sensible strategy than they've got now and I really doubt it's going to be rings for the type of straight up make no assumptions about your workload general computing Intel must be good at.
We don't agree that on-chip networks aren't important though. While right now no one's got a good programming model for these things, we're going to need one and there's likely going to be some data sharing involved, which means a good on-chip network is going to be a hell of a thing. Also I simply don't envision a cache architecture that makes sense that doesn't have a lot of unfortunate on-chip communication, and that needs to not be annoyingly NUMA. (Though it may have to be to some degree...)
I guess we'll find out. :)
And as an aside, thanks for taking the time to provide one of the more informative and responsive posts attached to this thread. I think HN could use more architecture folks.