For what it's worth, most of the many-core chips that do exist keep it fairly simple. Ambric's MPPA had a mesh with configurable routing (http://www.nethra.us.com/technologies_mppa.php), GPUs have a combination of pretty straightforward hierarchical interconnect and rings, and Azul's interconnect is a tree-like structure as well.
One thing to consider is that, while more complex topologies can buy you something in a supercomputer, give you more bisection bandwdith, and better support point to point communication, it has turned out to be pretty hard to actually write correct programs in that style (lots of arbitrary peer-to-peer communication); this partially motivates simpler interconnects, where the programmer can more easily reason about what's going on. Another concern is that topologies that work great in a server room (3D) are nigh-unroutable on a chip (2D). The wraparound links in a torus are a great example of this; Blue Gene/L used a torus to great effect (http://www.google.com/url?sa=t&source=web&ct=res&...), and the wraparound links drastically reduce worst-case and average point to point latency over a mesh, but those links mean giving up a large fraction of a metal layer on a chip, as opposed to a long cable in a machine room.
Also, I would hazard a guess that the workloads these chips are intended for are mostly multiprogrammed, not multithreaded; if they don't share much data, the network's bisection bandwidth is not as much of an issue as the off-chip bandwidth.
Regarding scalar operand networks as found in Raw/Tilera64: The specific idea (register-mapped networks) probably changes the programming model too much for Intel to adopt directly, but the Single-chip Cloud Computer (SCC) linked above, presented at ISSCC 2010, uses something sort of similar for message passing over their 2D mesh. The communication channels are memory-mapped rather than register-mapped, but the mechanism is similar.