> i am curious whether the analysis stages can improve the code.
Ihe conventional wisdom is that any series of analyses that takes you from representation X, through one or more other representations, and back to X can only make things worse, assuming that the JIT from PTX to GPU machine code isn't horrendous. This is because any analyses, optimizations, and transformations need to be conservative to maintain correctness, and high-level semantic information about the parallelism inherent in the application is usually lost in each translation step. In this particular case it might not be so bad, as long as the LLVM IR is rich enough to faithfully represent the Cooperative Thread Array (CTA) semantics in PTX and not flatten them to SPMD code. My intuition, however, is that it's not; LLVM was designed as a fairly generic virtual machine that would faithfully represent most CPU-like execution models, and hardware CTAs (also called 'warps' in Nvidia parliance) are mostly a GPU-only phenomenon. CPUs have SIMD units (e.g. SSE, MMX, Altivec, NEON), but the execution model there is fundamentally different than the GPU.