Buckle Up, Intel Preps 8-Core Nehalem-EX Chips for March Launch
hothardware.com
hothardware.com
STTRAM[1] might be capable of this when it comes online and in production, but it seems flash is unlikely to reach this point and STTRAM is still in development.
By all good estimates it's only maybe 4-10 years before we have to worry about straight up running out of pins on the chips. (We can put more cores in 'em, but we actually will have a really hard time scaling down the pins that connect the die to the rest of the computer, so this will be another I/O bottleneck. The die is going to need as much cache as we can give it.)
While I don't think it's technically illegal, I do find it strange that the manufacturers are allowed to collude to insure I personally overspend by at least 20% and never have top-of-the-line hardware for more than a month. I've come to accept that I'll never know how they picked me, but I'd still really like to know how they get my email passwords.
I think it'll work the other way too: A cheap way for hardware manufacturers to make their deadline is to ship me their current generation a couple weeks before they need to ship the next generation junk.
(Yeah, I know, sometimes you just gotta have the best. But still.)
Especially memory: there are some things that just aren't practical when your working set gets too large (e.g. you start constantly paging), whereas a somewhat slower CPU just requires a bit of patience for what I do (which isn't number crunching or anything else that keeps my CPU pegged).
All that said, the next major machine I build should be in 2011-12 and I'm really looking forward to having bunch of cores with more cache than, say, the 16MB of RAM I bought in 1991 for my 486 Windows machine.
As it is now, wouldn't it be hard to buy "normal" CPU that doesn't have at least as much cache as the address space of a PDP-10 (18 bits of 36 bits words, or a megabyte of 9 bit bytes)? And cache has become the new RAM, RAM the new disk, and disk the new tape.
Wild times for someone who started out with punched card FORTRAN on an IBM 1130 (64KB max, more likely a lot less, 1-10 MB disk, probably closer to the former).
These appear to be still built on the old 45nm process. It should be interesting to see what the 32nm westmere shrink look like. (The ones out so far are mostly only dual cores. They plan to do the rest of the line later in the year...)
There's a lot of interconnect on that chip. Which is not surprising if you're using the same type of on-chip network they've been doing. Clearly they're blocking some things together so it's probably not all to all, but it's not exactly the types of on-chip networks you can expect to scale a whole lot further than where they're at.
Also rings still don't scale, so they're really going to have to get more clever. I wonder what the perf hit for the ring is going to be.
One option is to move closer (but probably not actually adopt) something like a scalar operand network used on some of the tile architectures.
For what it's worth, most of the many-core chips that do exist keep it fairly simple. Ambric's MPPA had a mesh with configurable routing (http://www.nethra.us.com/technologies_mppa.php), GPUs have a combination of pretty straightforward hierarchical interconnect and rings, and Azul's interconnect is a tree-like structure as well.
One thing to consider is that, while more complex topologies can buy you something in a supercomputer, give you more bisection bandwdith, and better support point to point communication, it has turned out to be pretty hard to actually write correct programs in that style (lots of arbitrary peer-to-peer communication); this partially motivates simpler interconnects, where the programmer can more easily reason about what's going on. Another concern is that topologies that work great in a server room (3D) are nigh-unroutable on a chip (2D). The wraparound links in a torus are a great example of this; Blue Gene/L used a torus to great effect (http://www.google.com/url?sa=t&source=web&ct=res&...), and the wraparound links drastically reduce worst-case and average point to point latency over a mesh, but those links mean giving up a large fraction of a metal layer on a chip, as opposed to a long cable in a machine room.
Also, I would hazard a guess that the workloads these chips are intended for are mostly multiprogrammed, not multithreaded; if they don't share much data, the network's bisection bandwidth is not as much of an issue as the off-chip bandwidth.
Regarding scalar operand networks as found in Raw/Tilera64: The specific idea (register-mapped networks) probably changes the programming model too much for Intel to adopt directly, but the Single-chip Cloud Computer (SCC) linked above, presented at ISSCC 2010, uses something sort of similar for message passing over their 2D mesh. The communication channels are memory-mapped rather than register-mapped, but the mechanism is similar.
I think they'll drop the rings soon. But it will keep them going until they figure out how to solve their interconnect problems.
I agree with you that the SON as on RAW/Tilera is unlikely to make the leap to Intel. I just hope they'll move towards that direction. Though clearly they may also chose to move in a completely different direction, but they will need a more sensible strategy than they've got now and I really doubt it's going to be rings for the type of straight up make no assumptions about your workload general computing Intel must be good at.
We don't agree that on-chip networks aren't important though. While right now no one's got a good programming model for these things, we're going to need one and there's likely going to be some data sharing involved, which means a good on-chip network is going to be a hell of a thing. Also I simply don't envision a cache architecture that makes sense that doesn't have a lot of unfortunate on-chip communication, and that needs to not be annoyingly NUMA. (Though it may have to be to some degree...)
I guess we'll find out. :)
And as an aside, thanks for taking the time to provide one of the more informative and responsive posts attached to this thread. I think HN could use more architecture folks.