Analysis: more than 16 cores may well be pointless
arstechnica.com
arstechnica.com
It's brilliant, and seems effective, but it's still just 8 cores per socket.
True, the two groups of cores would probably not be able to share memory (at least not at full speed), but that's something the OS scheduler and memory manager could take care of.
For one, there are plenty of pins. The socket used by Core i7 more than doubled the amount of pins (in the 1300s now), yet a QuickPath Interconnect (Intel's replacement to the FSB) only takes 84 pins. Surely, they can squeeze those in.
Going with multiple mem controllers means that performance would still probably be pretty variable, depending on how the mem controllers are mapped to your physical addresses + how your data is laid out.
That's not true for i7. QPI connects to the I/O bridges and busses, but the on-chip memory controller connects to RAM directly.
And I would guess that the memory controller uses a ton of pins, and is the reason why the triple-channel i7 has such a high pin count. The upcoming consumer-grade dual-channel model will have something like 200 fewer pins.
Absent such workarounds, it's true that there are some problems (large-mesh finite element analysis, maybe) that more cores won't help with, and other problems that exhibit enough locality (large cellular automata) that more cores will help with. This article observes that some problems are in the first category. It would be absurd to claim that no problems are in the second category.
On second thought, it's probably not worth it. It seems modern branch predictors are at least 90% accurate (http://en.wikipedia.org/wiki/Branch_predictor), and only a few percent of instructions are conditional branches anyway (http://bloggablea.wordpress.com/2007/04/27/so-does-anyone-ev...).
So the problem is really that we have:
[lots of cores] <=chokepoint of memory bus=> [lots of ram]
when we could have:
k times: [1/k of our cores] <=single mem bus=> [1/k of our RAM]
That increases our aggregate memory bwidth by a factor of k. This has to be <= number of cores, and the problem being solved needs to be able to be partitioned into k chunks.
This is basically the clustering approach (individual proc+ram working on the problem), with the added advantage that we can leave some of the RAM unsplit so we get 'local' shared RAM for free.
I guess one approach to that is a bigger die size, so you end up with a fixed transistor/cm^2? That would mean die sizes doubling with Moore's law I guess. So you'd want other ways of packing them in (use the flipside of the mobo, try 3d arrays and suffer heat problems).
Or the cores come with attached (non-shared) RAM, which is of course where we are with adding cache.
Other random thoughts: why have memory busses stayed parallel when peripheral busses (scsi, usb) have gone serial? That would reduce pin count for a connection?
Can we avoid going to full 'macro' pins for the CPU-memory bus (and thus pack more pins into the same area for memory connections)? Instead have a smaller, denser collection of pins which are attached as a group to each memory connector?
Sorry for being clueless and thinking aloud, but it's an interesting problem.
FB-DIMM is already a serial-style high-frequency interconnect, but it isn't needed on low-end systems.
I think this is really about business, not technology. What they want is not what the mainstream wants, so they must choose between cheap but memory starved systems or very expensive balanced systems. HPC people have been whining about killer micros for 20 years; this is just another version of it.
Join me in voting up all people who make sensible posts (even if you disagree with them) and who are being voted down by newbies abusing the points system :)
The point made in this article surrounds current architectures and their limitations. The conclusion that more than 16 cores makes no sense might well be a good conclusion for now, but "more than 16 cores may well be pointless" is by no means a conclusion for the long or even mid term.
From the article: "But, to my knowledge, these die-stacking schemes are further from down the road than the production of a mass-market processor with greater than 16 cores."
The article is pretty clear that "more then 16 cores may well be pointless" is for now.
Don't you mean usually? I haven't seen the big break throughs in AI that were expected. I haven't seen a solution to the halting problem, etc, etc.
AI hasn't lived up to the promises made a few decades ago.
* 1958, H. A. Simon and Allen Newell: "within ten years a digital computer will be the world's chess champion" and "within ten years a digital computer will discover and prove an important new mathematical theorem."[53]
* 1965, H. A. Simon: "machines will be capable, within twenty years, of doing any work a man can do."[54]
* 1967, Marvin Minsky: "Within a generation ... the problem of creating 'artificial intelligence' will substantially be solved."[55]
* 1970, Marvin Minsky (in Life Magazine): "In from three to eight years we will have a machine with the general intelligence of an average human being."[56]
PS: A digital computer is the worlds chess champion or would be if we let them compete. Making a useful captia is hard, but computers don't compose poetry so we can still say we don't have AI.
"they give the anecdotal hint that limitations will ALWAYS be beaten."
Limitations will not always be beaten.
In the same way "more than 16 cores is pointless" is a reasonable claim if one presumes that the technical limitations in the article are insurmountable.
Given our past experience, no technological limitation is truly insurmountable.
Therefore, I'm willing to assign as much belief in this article as I am the "five computers" claim.
It's extraordinarily painful to code for, but hey, all performance optimization is an exercise in caching. So it goes.
IMHO, it looks like we'll need some smarter memory bus management. If we're looking at the 1990s, anyone remember the crossbar switches SGI used to put in their short-lived x86 boxes? Thoughts on effectiveness?
It seems to me that each thread already has it's own memory space, just make sure the memory space for the thread is on the same CPU the thread runs on.
It isn't really necessary for each core to be able to access all memory (or at least it'll be way slower to access memory outside it's area).
Perhaps it is time for home desktops to adopt the super computer architecture!
Can somebody explain how mainframe DMA is different from the DMA in your home PC?
I know mainframes have more then 16 CPUs. The memory problem we're talking about here does not apply to them, right?
The real title should have been "With the current architectures and/or memory speeds, more than 16 cores may well be pointless".
There, FTFY.
If the clock speed and number of cores increases at a faster rate than memory speed increases (assuming all the cores share memory) then at some point the memory can't keep up.
Whether or not 16 cores is the magic number, I don't know.