The Xeon Phi uses a different computing architecture than either a CPU or a GPU so some intuitions about using it will be off. The Phi is essentially a modern barrel processor that uses the AMD64 ISA (it does not understand legacy x86 modes), a really nice vector ALU implementation, and memory bandwidth that is more like a GPU than a CPU. While it will run normal 64-bit software reasonably well with no special considerations, it will not be efficient without tweaking code design. CPU-targeted code attempts to optimize IPC in a single thread; barrel processors are designed to hide latency and
can't drive IPC through single thread optimization.
The reason barrel processors are interesting is that they can be incredibly efficient with their clock cycles. Unlike either CPUs or GPUs, it is relatively easy to get sustained throughput that approaches the theoretical IPC of the silicon for diverse software. The Xeon Phi mentioned in the article has 114 ALUs; it is possible to ensure all of those ALUs are doing useful work every single clock cycle, unlike the much smaller number of ALUs in your CPU. CPUs and GPUs have higher theoretical throughput in some cases but various parts of their ALUs typically spend a significant part of their time idle.
Contrary to marketing, you do not want to program these like an ordinary CPU even though the cores are truly general purpose (unlike a GPU). Thread behavior is unlike CPUs or GPUs. Barrel processors cannot saturate a single core with a single thread! The Xeon Phi has 228 independent threads and you need to use them all the time.
The way barrel processors work is if the hardware supports N threads then each clock cycle you can saturate all the ALUs if some subset M of those threads are not stalled. The M-of-N ratio varies by barrel processor design but is typically 20-50% in my experience. Each clock cycle, a core selects an immediately runnable operation from the basket of threads it can see and executes it; as long as something is runnable in that basket, the core will do real work that clock cycle. Xeon Phi has a 50% M-of-N requirement, so you need a minimum of 114 threads that are not stalled every clock cycle to saturate the processor. The way you ensure that you hit the 114 threshold is to schedule all 228 hardware thread slots with useful work.
For programming, this changes the way you reason about locality and concurrency. If we assume that some percentage of threads can be safely stalled or blocked with no impact on throughput then it changes the way you design your algorithms and data structures. A little additional latency on a subset of threads won't hurt performance, especially if it increases task concurrency. On a CPU stalled threads are expensive, as it leads to idle cores or context switches. For architectures like Phi, you design your data structures and algorithms around relatively small, semi-independently work units so that there is always a large number of tasks that can be assigned to a thread even though this reduces locality. Below a certain threshold, thread concurrency is approximately free because the cores will schedule around stalls due to contention.
I like barrel processors quite a lot. Once you get used to the model, it is an easier architecture with which to achieve efficient massively threaded parallelism and throughput than either CPUs or GPUs. More importantly, they are hard to beat for efficiency for general purpose computing when software is designed for the architecture since so few clock cycles are wasted.