CUDA to x86 compiler, project Ocelot
code.google.com
code.google.com
PTX = Parallel Thread Execution is a pseudo-assembly language used in nVidia's CUDA programming environment. The 'nvcc' compiler translates code written in CUDA, a C-like language, into PTX, and the graphics driver contains a compiler which translates the PTX into something which can be run on the processing cores.
(source - wikipedia)
i am curious whether the analysis stages can improve the code.
> i am curious whether the analysis stages can improve the code.
Ihe conventional wisdom is that any series of analyses that takes you from representation X, through one or more other representations, and back to X can only make things worse, assuming that the JIT from PTX to GPU machine code isn't horrendous. This is because any analyses, optimizations, and transformations need to be conservative to maintain correctness, and high-level semantic information about the parallelism inherent in the application is usually lost in each translation step. In this particular case it might not be so bad, as long as the LLVM IR is rich enough to faithfully represent the Cooperative Thread Array (CTA) semantics in PTX and not flatten them to SPMD code. My intuition, however, is that it's not; LLVM was designed as a fairly generic virtual machine that would faithfully represent most CPU-like execution models, and hardware CTAs (also called 'warps' in Nvidia parliance) are mostly a GPU-only phenomenon. CPUs have SIMD units (e.g. SSE, MMX, Altivec, NEON), but the execution model there is fundamentally different than the GPU.
Of course that will introduce a new level of complexity to the optimization problem.
for me, the big advances in fermi are a unified address space and some kind of cache for the global memory. neither of those change the paradigm, but they may make life significantly simpler when programming the thing.
(and let's hope it goes that far down), that would make things a lot easier as well.
Unified address space I assume you mean across multiple GPUs ? Global memory cache is a double edged sword, that eats in to the transistor budget at a very rapid pace, effectively you already have a cache, you just have to fill it yourself.
GPU programming is definitely a step back in the ease with which you can write programs, but if your problem maps well on to a GPU the speed increases are simply astounding. What would have taken you a cluster with 100 boxes now sits under your desk and consumes 250 watt tops. That's really very impressive.
The way intel seems to edge in to gpu territory and nvidia into cpu territory will make for some interesting stuff happening in the next couple of years.
http://www.nvidia.com/content/PDF/fermi_white_papers/D.Patte...
I'm not sure that's impossible, it just seems very hard.
If nvidia manages to crack that nut then the only thing you'll still need to keep in mind is how big your cache footprint is (as on every other cpu with a cache) in order to maximize throughput.
That would definitely be a good thing.
I've spent in total about 2 months now (spread out over the last year) understanding how this whole GPGPU thing fits in with the rest of computing, it is much like a specialty tool. It is harder to master, more work to get it right once you have mastered it, subject to change on shorter notice than most other solutions (because of the close tie to the hardware) but if you need it, you need it bad and the pay-off is tremendous.
(ie i agree with everything else you say - i just don't understand what that graphic is trying to show).
If you want to read more: http://llvm.org/devmtg/2009-10/Grover_PLANG.pdf