Intel Launches Next Gen Itanium Monster Processor
hothardware.com
hothardware.com
2) There is way more ILP available at run-time than compile time, and what ILP is available in both is much more tractable at run-time. An out-of-order CPU is constantly filling a buffer with instructions (or microcode), and hardware is determining the dependencies dynamically to issue them to the ALU. This is a much more tractable problem than trying to guess the control flow at compile time. A large enough prefetch buffer can overcome a really dumb compiler.
3) Requiring software to be aware of details of your hardware implementation is a really tempting idea, but it has been historically much worse than the opposite. Consider a modern x86 that is nothing like an 80386, but runs the same software, often at higher IPCs than an original 386. Now compare to MIPS which has its delay-slot for branches which often just gets filled with a NOP. Furthermore on modern MIPS cores, which have longer pipelines and branch-prediction, that slot is more-or-less useless!
4) Assuming IA64 doesn't die out, by the time compiler writers figure out how to make code that runs fast on an Itanium of today, Intel will be performing hardware gymnastics to make that code run fast on the hardware of tomorrow.
(But, then again you could always use a background optimization thread...)
AFAI(K|R), Hennesy and Patterson's cannonical text (CA-AQA [1]) reflects this: going from 3rd to 4th edition, we find a new chapter "Limits on ILP", VLIW/EPIC elements have been moved from the main contents to the CD-ROM, too (which probably is not a good indicator, though: the 3rd edition was just too heavy to carry it around a lot ;)
[1]: http://www.amazon.com/Computer-Architecture-Quantitative-App...
haha words for life
That said, looks like an impressive processor.
>Itanium relies on the compiler to optimize code at run-time
Thats sums it up for me :)
But seriously - in case when compiler is able to do parallelization, NVidia GPU seems to be a better - cheaper, more accessible and performant - target.
Itanium seems like it fits nicely in that space - operations complex enough to be very painful on a graphics processor, but parallel enough for you to actually consider using a GPU in the first place.
Itanium differs from other processor architectures is how it handles instruction level parallelism. The processor in the computer in front of you probably uses out-of-order execution (http://en.wikipedia.org/wiki/Out-of-order_execution) to exploit ILP. This happens on the fly, as a program executes. Itanium depends on the compiler to determine where ILP is.
data parallelism is a partial case of instruction level parallelism - N instances of the same instruction run for different pieces of data. A very frequent case in high performance computing or enterprise data crunching tasks supposedly targeted by the Itanic
http://en.wikipedia.org/wiki/Instruction-level_parallelism
and returning to the original specific context of Itanium vs. GPU :
http://http.developer.nvidia.com/GPUGems2/gpugems2_chapter35...
Take your stub at what can be classified as what :)
Correct. The GPU article on NVIDIA's website uses ILP incorrectly. They are describing SIMD operations - single instruction, multiple data - which is data parallelism at the instruction level. This is inherently different than ILP, which is when you extract parallelism from a sequential stream of instructions by executing them out-of-order. Of course, it's possible to exploit ILP on a stream of SIMD instructions.
Back to my main point - GPU vs. Itanic. NVIDIA GPU Tesla have 512 SIMD cores organized into 16 SM ("streaming multiprocessors") with 2 independent instructions (from 2 different threads) issued per SM per clock (each instruction goes to its' half of the SM, ie. to 16 cores wide SIMD group):
http://www.nvidia.com/content/PDF/fermi_white_papers/NVIDIA_...
That gives us [16 SM x 2 instructions x [1-16 cores]] - anywhere between 32 to 512 ops / clock - 1000 GFlops in the best case.
The Itanic - 8 cores x 6 independent instructions - something like 200GFlops.
That depends on the implementation. Lets say there is a program
f3=OP1(f1)
f4=OP1(f2)
OP2(f3)
OP2(f4)
Some possible ILP forms:
OP1(f1)OP1(f2)
OP2(f3)OP2(f4)
or
OP1(f1)
OP1(f2)OP2(f3)
OP2(f4)
A DLP form:
OP1(f1, f2)
OP2(f3, f4)
or may be their DLP implemented as
OP1(f1)OP1(f2)
OP2(f3)OP2(f4)
and if it is the case i don't see why they can't call it an ILP.
ILP means something very specific. This is a discussion of semantics, but semantics are important so that we can communicate easily. If I stick a Hershey's bar in the oven, it is literally hot chocolate, but it is not what we normally mean when we say "hot chocolate." You are talking about parallelism at the instruction level. I'm trying to explain that that is note what people mean when they say "instruction level parallelism."
ok, just point which of the 2 ILP executions i mentioned above isn't really ILP, and which of the 2 DLP executions isn't really a DLP.
OP1(f1)OP1(f2)
OP2(f3)OP2(f4)
Instruction level parallelism because OP1(f1) executes in parallel with OP1(f2) and OP2(f3) executes in parallel with OP2(f4).
OP1(f1)
OP1(f2)OP2(f3)
OP2(f4)
Instruction level parallelism because OP1(f2) executes in parallel with OP2(f3). (But, the schedule isn't as good as the first one.)
OP1(f1, f2)
OP2(f3, f4)
Data level parallelism because OP1 is applied to both f1 and f2. While this has the same result as the first example, a processor would achieve both in different ways. In the first example, the processor would have to fetch two instructions. It just so happens that both of those instructions are OP1. Then it would have to schedule both of those instructions, and it was luckily able to schedule them both at the same time.
The third example is different. In this case, the processor would fetch one instruction, but execute it on both f1 and f2 at the same time. That's why it's called SIMD: single instruction, multiple data. One instruction executes, but it modifies multiple data elements. In the first case, you had to fetch and execute an instruction for each data element.
Why bother distinguishing between them? Because this may also be possible:
OP1(f1, f2)OP2(f3, f4)
That is both data level parallelism and instruction level parallelism.
You could fit ~440 MIPS R10k processors on that thing.
> Given sufficiently intelligent compilers, Itanium could begin to make economic sense in fields that couldn't previously justify the high cost of optimizing for the chip.
As you point out, the SSC has yet to appear. Just look at that layout: it's dominated by cache.
I think the main benefit would be much simpler instruction pipelines, which would include the points you mentioned (branch prediction, prefetch) but also all of the logic needed to keep track of dependencies in an out-of-order processor.
True about data, though I vaguely recall EPIC had advantages there too because without needing to do branch prediction, you didn't need to speculatively fetch multiple memory addresses; meaning the same D-cache went further.