Intel's 50-Core Xeon Phi: The New Era of Inexpensive Supercomputing
drdobbs.com
drdobbs.com
The article states that "All of these [CUDA/OpenCL] problems go away with the Phi. It's a pure x86 programming model that everyone is used to. It's a question of reusing, rather than rewriting, code" but I find it hard to believe I can just drop existing code into it and expect decent performance.
#pragma openmp parallel for
in front of them, and your code transfers through pretty much intact -- it handles the thread wrappers. You can add other pragmas for the times when you need locking.This is a much less intrusive setup than CUDA; you don't have to worry about loading data, or double/float conflicts.
The OpenMP extensions could be a very good fit for scientific programming on this coprocessor.
or if you can't be bothered to wait for such a standard, you should have a look at OpenACC[1], which does exactly this, and exists now. you end up adding code like
#pragma acc kernels for
on top of your for loops, it does the low level work for you.I'm sure many other people are in the same position. I don't mind sacrificing a little performance for a much easier programming environment.
If you want to go all out you can probably get to the required level of knowledge based on a few weeks to a few months of really hard work depending on where you are coming from in terms of experience.
The docs are excellent, there are tons of examples and google will usually turn up a solution to a problem in case you hit a snag.
Compiler support for CUDA is further ahead of Xeon Phi. I haven't seen any evidence that Intel has been successful yet extending the auto-vectorizing capabilities of ICC to massively parallel environments.
For CUDA, one can write code in Python, Haskell, C++ etc.
At this point, Xeon Phi only offers Vectorized C and Fortran. There is a narrow domain of HPC code designed to run on Multisocket Xeon processors that could probably be re targeted trivially to the Xeon PHI.
If you run Linux on the Phi (which Intel ships in their Manycore Platform Support Stack), then anything that runs on 386 linux should run here, which should include Python and Haskell.
If you don't choose to use Linux on the Phi, then your tools options will be limited just like they are if you choose to not use a regular OS on a PC.
You can be sure it won't be any currently popular language, though, because almost none of them have support for the kind of pervasive parallelism needed (why isn't parfor the default kind of for loop?), and of those that do support that kind of parallelism (typically by being purely functional and supporting lazy evaluation), none come equipped with the necessary facilities to optimize the code for a particular GPU (by tweaking how the problem is split up).
The Xeon Phi does have an advantage in that it's easier to get the code running in the first place, but the difficulty of optimizing it for a massively parallel GPU-like architecture is (for now) exactly the same as faced by OpenCL and CUDA users.
I suppose it is strictly less powerful than regular MapReduce, but at least with Hadoop the system administration costs are too much for a lot of people, and machines are getting beefier, so you can get a lot more done on one machine. In another recent thread there was a MS research paper about "ill-conceived" Hadoop clusters processing 14GB of data...
The main benefits I see are:
1) You don't have to write in a specialized language. You should be able to use any language with a good implementation. Scientific code often has Matlab, R, C++, and Python glued together.
2) MapReduce lets you write sequential code, which is easier to learn.
3) You can adapt/port sequential legacy code easily, so you can use a lot of your existing code.
MapReduce is of course similar to "parallel for" but more powerful -- parallel for is essentially the map stage. The reduce stage adds a lot. For some reason most people who haven't programmed MapReduce think of MapReduce as just mapping, and they don't understand reducing.
If you want to do it quick and dirty, don't underestimate "xargs -P" :) That's your "parallel for" that works with any language. You can run that on your Matlab, Python, C++, etc. You need a serialization library but there are a lot of those around. It works well and with a minimum of programming effort.
What the Phi is doing is combining those approaches, running two threads simultaneously and switching threads out on cache-misses. This way you only double rather than quadrupaling your control structures, but you don't have your cores entirely unutalized when you're swapping threads. A really nifty compromise, I think.
Having said that, I've used Unix workstations with less RAM attached than that through much less than 7GBps worth of bus...
Edit: 1 GHz sounds like plenty until you realize it's in-order execution. This would be noticeably sluggish.
Are they even out-of-order? I.e. is it Pentium or Pentium Pro class?
Each core is a simple in order x86 CPU (derived from the original Pentium) with a 512-bit SIMD unit.
Instruction re-ordering is more about taking full advantage of multiple execution units (ALUs, etc.), or not completely stalling the pipeline to wait on a memory fetch.
http://www.tomshardware.com/reviews/xeon-phi-larrabee-stampe...
First, the talent pool for HPC x86 programmers is an order of magnitude larger than for expert GPGPU programmers - Xeon Phi is just a virtual x86 server rack with TCP/IP messaging.
Second, the amount of time and effort to extract useful performance from GPGPUs is quite a lot; if it's for internal use and you're not selling the code to the masses, you're likely to get the same amount of performance with less time on the Phi, unless you're going for "the best, regardless of money & time".
Last, most enterprise customers will want ECC + other compute features. They're sold in the pro-level 3k+ Teslas, which happen to be more expensive than the Phi.
Where GPGPU does make sense: consumer-level hardware using already-written software (workstations and hobbyists in particular) and businesses where performance/watt is crucial at any cost.
With 60 cores reading memory over a common ring bus latency will kill you unless you tile your loops to maximize cache reuse [1], at which point you might as well write a GPU code which preloads blocks of data to local memory and works there.
Also, to beat performance of normal x86 CPU you must use vector instructions, what gives you all the little problems GPU warps are known to cause.
[1] http://software.intel.com/en-us/articles/cache-blocking-tech...
the main optimisation techniques for GPUs aren't difficult to grasp (in my opinion), although not all classes of problem are suited to execution on GPU.
How?
The latest and greatest Nvidia card at your favorite retailer.
That's $44.15 per 1 GHz core.
AMD FX-6300 Six-Core 3.5GHz is $138 = $23 per core (and much faster cores).
Intel Xeon 5148 2.33ghz is $18.
You can get a quad CPU motherboard relatively cheaply:
http://www.ebay.com/itm/Arima-Quad-CPU-16-Core-AMD-Opteron-M...
Also don't forget to add the same costs for the Phi solution.