Intel SPMD: Compiler for High-Performance SIMD Programming
ispc.github.io
ispc.github.io
IMO these constraints are overstated and OpenCL offers a good abstraction for CPUs too.
Since you seem like you have some experience with testing performance portability with OpenCL and other solutions, I'd be curious to hear if you have any comments about the reference I linked or more general suggestions for alternative means to achieving the same end (performance portability between CPU/GPU architectures at the workstation level).
[1] http://iwocl.org/wp-content/uploads/iwocl-2014-tech-presenta...
[1] https://www.dartlang.org/articles/simd/reference: https://ispc.github.io/ispc.html#the-ispc-parallel-execution...
However a lot of interesting problems are seemingly parallel but highly branching and nonlinear. Take path tracing as an example: it's very little code and highly parallel as each Ray/pixel is independent, yet it's not an easy problem for a GPU: each time a ray bounces it will disperse and not do whatever the Ray next to it was doing in terms of which geometry it will hit etc.
It might seem like today if a problem can benefit from 8 CPU cores then it benefits 100x more from being run on a GPU but this is far from true. A great machine for general computing could do well with a board with 100 x86 CPUs apart from having a big gpu with a thousand cores for brute forcing the "simpler" problems.
Any references to what has "1000 cores"? Nvidia GPUs usually have about 12 or so cores that can be compared to x86 cores, meaning they can independently branch.
For example high end Nvidia 980 GTX GPU has only 16 of such comparable SIMD execution cores. SMXs or whatever Nvidia calls them.
GPU marketing materials confusingly refer as cores to something like x86 CPU SIMD lanes (and that's being very generous to GPUs), that artificially inflates the numbers.
Or put differently, one CUDA core can compute up to 1 FMA per cycle @1196-1300 (?) MHz. One recent Intel X86 core can compute at least up to 16 FMAs per cycle @2800-4000 Mhz.
A WARP is really nothing more than a way to have work for SMXs (and computational units it controls) at as many clock cycles as possible. You need some way for masking FPU pipeline and memory latency.
> All warps are running in parallel (otherwise you won't get the performance numbers) and each has its own control path (actually each has its own code)
It's not that different from x86 hyperthreading, just with more hardware threads. Pipelined execution units are fed each clock cycle by the core. Multiple FP operations are in flight in parallel, otherwise CPUs won't get the performance numbers either.
https://www.nvidia.com/content/PDF/kepler/NVIDIA-Kepler-GK11...
https://www.nvidia.com/content/tesla/pdf/nvidia-tesla-k40-20...
A "warp" is analogous to a hardware thread and you'd have up to 64 of those being scheduled on each SMX or SMM. Each of those SMX/SMMs has four warp schedulers which issue instructions to execution units. In an SMX the schedulers can issue to any of the 192 execution lanes but in an SMM each scheduler has it's own set of execution lanes. If we call a core anything that can independently issue instructions then I guess you'd call an SMX a core but on a SMM each warp scheduler looks like it's own core. But this is all further complicated by the fact that an instruction issued to one lane can be crossed over to a lane that's become idle due to predication. Which is maybe sort of like scheduling but not really.
But yes, you can't compare "CUDA cores" to actual cores and GPUs aren't equivalent to thousands of cores. The GM204 would have 64 core equivalents and most other chips would have less.
There has been a surge in the tractability of massive but simple linear algebra problems lately, such as deep learning, which might have given the impression that GPUs are the answer to any supercomputing.
Which is the idea behind Intel's Xeon Phi "GPU" with 70+ Pentium/Atom cores, which this compiler specifically targets.
[Long before that, Intel showcased an 80 core x86 CPU in 2007 (Polaris/Teraflops Research Chip) – and then promptly shelved it to focus on building programming languages and compilers that can actually make use of it, before introducing the Xeon Phi half a decade later.]
http://sbel.wisc.edu/Courses/ME964/Literature/LeeDebunkGPU20...
Bloomfield era you could have 4 (?) cores per CPU socket. Now Broadwell EP has 22.
Only thing that hasn't scaled much CPU side is memory bandwidth. I think it's only a matter of time until Intel integrates HBM2 or something like it to same package. They've already done that for eDRAM.
4 per socket, and at most 2 sockets per board.
Broadwell-EX, to be released this quarter, has 24 cores and up to 8 sockets per board.
So 64 FP ops per machine and cycle versus… 6144.
In the same time, GPUs went from 900 GFLOPS per card, 2 cards per machine (1800 GFLOPS total vs. 192 on CPU), to 9600 GFLOPS per card, 4 cards per machine (38400 vs. 12000). GPUs are still faster, but the advantage isn't that significant any more.
First with Haswell introducing the fused multiply add, suddenly all the 'peak flops' numbers doubled, which is technically true, but only if everything you do is a fused multiply add (with no cache misses of course).
Even so only the Xeon Phi (and only the unreleased silvermont cores?) has 16 wide vector units, even Skylake still has 8 wide AVX units, which would be 16 fma operations.
Are you saying that AVX instructions are pipelined (or some other technique) and have a throughput greater than their width per cycle?
https://www.google.com/search?q=haswell+32+single+precision+...
> First with Haswell introducing the fused multiply add, suddenly all the 'peak flops' numbers doubled, which is technically true, but only if everything you do is a fused multiply add (with no cache misses of course).
Yeah, FMA (fused multiply-adds).
Better or worse, it's de facto standard to quote one FMA as two FLOPS, because it's a very commonly combined operation.
> Are you saying that AVX instructions are pipelined (or some other technique) and have a throughput greater than their width per cycle?
Yeah, AFAIK, they're (mostly) pipelined and dual issue per clock.
(I say this as a GPU language developer. They are fast, but also a bit of a pain in the ass.)
> temperamental GPUs and GPU drivers
I think because Vulkan won't be nearly as niche GPU driver support and consistency will need to be much better than they have been for OpenCL. Where that stands for flexibility of the compute side I can't say.
I will say this though, a program written well in ISPC with cache locality taken into account together with SIMD can run 100x faster than a naive C program.
See the full paper I link to as a top-level comment for more details.
One example would be the n-body simulation of the computer language benchmarks game. The C++ version uses intrinsics but wouldn't benefit from anything that can do 4 doubles instead of only two at a time.
OTOH, I like the idea of automatic lane width detection.
However, I agree that having another dependency in your build may be more of a problem.