Intel's Xeon Phi Is Being Sold for a Low Price
phoronix.com
phoronix.com
It's hard to compare different devices of course, but in terms of pure FLOPS (and benchmarks) is roughly the same as an AMD r280x, which costs 240$ and allows you to play video-games in between your work.
There is definitely a market for these cards I think, niche but it exists, however the normal price of >$2k is the old-fashioned "industry tax" (c.f. Tesla cards) and not a reflection of the actual performance.
That's what I was thinking too - if Intel just added the needed hardware for some video outputs, they could've produced a version which also works as a GPU, and might receive more interest.
A better buy would be an AMD APU, which for ~150-200 watts (compared to the 225-300 for the Phi, which requires a ~130W CPU to operate) can do roughly the same processing work with a little less onboard memory.
- It doesn't support CMOV, MMX, SSE, AVX, and all other ISA extensions the popped up after the original Pentium.
- It does not use standard System V x86-64 ABI.
- Most of its compute power is due to special vector instruction set, that is not compatible with any previous or future x86 ISA (including AVX-512).
- The only fully supported compiler is Intel Compiler (icc). And in my opinion, its code-generation quality for Xeon Phi is far worse than for other Intel architectures.
- The only supported assembler is GAS. No NASM, no YASM.
- Most debugging and profiling utilities (that run fine on normal x86 cores) are not supported on Xeon Phi. This includes valgrind, Address Sanitizer, Thread Sanitizer, and even memory profiling options in Intel compiler.
Given these limitations, how does it help that Xeon Phi is x86-based?
I can get an FX-9590 for $250, which has 8 cores, but running at 4.7GHz, and higher IPC. In terms of raw compute speed, it seems like the FX-9590 will be at least the same speed. But:
* I can use standard programming tools.
* I only distribute computes 8-ways, not 60-ways. That's easier.
* Anything which is not easily made parallel is much faster. My experience is that there tends to be a lot of that.
* I have less work to transfer data back and forth. For big data, that actually takes a fair bit of time.
When would the Phi be faster? When would I want to use it?
They aren't really comparable. An FX-9590 at 4.7GHz is about 300 GFlops, but the the Xeon Phi 31S1P is slightly over 1 TFlop --- more than 3x as many floating point operations per second. The spec that probably shows the most difference is memory bandwidth. The FX-9590 with 1866 memory achieves 30GB/s. The Phi reaches 320GB/s, about 10 times as much. The presumption is that you might put 4 of these in a server, in addition to the standard processors.
The price is actually quite a bit better, too, as some vendors are offering the Phi at $125 each in quantities of 10: http://www.colfax-intl.com/nd/xeonphi/31s1p-promo.aspx
This article offers some good information about cases where the Phi is a good (or bad) choice: https://software.intel.com/sites/default/files/article/33016...
The short answer would be that the Phi might be the best choice in cases where you need to perform billions of energy efficient floating point operations on a working set small enough to fit in the cards RAM, where the degree of branching is such that a "normal" GPU would be inappropriate, and where the problem justifies considerable programmer time for optimization. Financial and scientific models are the usual examples.
So, ok, the Phi is useful at the current price, but at the 2000USD price other people are talking about in the thread it seems pretty useless.
And if you think $2000 is expensive for 60 x86 cores (~$35/core) when all you care about is the SSE/AVX unit and how many vector ops you can cram through it, you're definitely not Intel's target.
But I'll quibble with the assertion that all GPU's are "integer pansies", although it's a fantastic phrase. Historically, NVidia GPU's are much stronger at floating point operations than integer because of their graphics origins, and NVidia's Tesla is definitely what Intel views as nearest to the Phi's target market.
But AMD/ATI cards have excellent vector integer performance, several times that of NVidia/Tesla, and much better per dollar than Phi until this recent price drop. This is why AMD cards were the best choice for bitcoin mining until the recent arrival of dedicated ASIC's: http://www.extremetech.com/computing/153467-amd-destroys-nvi...
> It doesn't support CMOV, MMX, SSE, AVX, and all other ISA extensions the popped up after the original Pentium.
I'm not saying it's not useful, it's just not quite as useful as having 60 cores that can just run MMX/SSE/AVX.
I'm confused by this. You previously stated that memory bandwidth on the Phi is 10x as high as on the FX-9590. Based on this, I would say that the Phi would be better suited at jobs which require lots of data that doesn't fit in main memory.
So if your working data set fits on the card (less than 8/16 GB), or can be partitioned so that it fits on multiple cards, you can potentially get great performance from Phi. But if it's larger than this, especially if it is less than 1TB and would thus fit in the host's RAM, transfer from the host to the card will likely destroy any performance advantage.
In terms of FLOPS the FX-9590 is still significantly slower. I don't know the exact numbers, but it will be hundreds of MegaFlops, whereas the Phi runs at ~1 TFlop. In part because the cores are basically low-end (in-order) atoms with modules bolted on for HPC (avx-512 instructions, instructions for exponentiation for example).
Data transfer to the card is a lot slower than reading from ram, correct, but on the other hand once it is on the card it is a lot faster (wiki says the fastest Phi has 350 Gb/s, Corsair claims a maximum of 70 Gb/s for DDR4)[2].
So generally, as long as you have a decent ratio of work/memory access and your work is parallel you could use a Phi to speed things a long. You would want to use one in the same situations as were you'd want to use a GPU, and the Phi would be preferable if your computations and memory access patterns are nor perfectly homogeneous (e.g. branching which can't be rewritten), as this kills performance on GPUs.
[1] I'm pretty sure Intel claimed this, and it makes sense when you think about it, but I can't seem to find a source: Could someone confirm this? [2] Do note that it 350Gb/s shared between the cores, not 350 Gb/s each.
I'd be interested to hear about your experiences. By "more complex", do you mean the x64 instruction set itself, or the out-of-order superscalar processors that use it?
GPUs are much stricter in their design, which can lead to some headaches when having to completely change the way you formulate a problem, but once you're past that step, the simplicity of the architecture makes it a lot nicer to optimize in my opinion. (And my personal experience has been that the tools are cheaper (free) and more stable as well, but your mileage may vary).
The reason barrel processors are interesting is that they can be incredibly efficient with their clock cycles. Unlike either CPUs or GPUs, it is relatively easy to get sustained throughput that approaches the theoretical IPC of the silicon for diverse software. The Xeon Phi mentioned in the article has 114 ALUs; it is possible to ensure all of those ALUs are doing useful work every single clock cycle, unlike the much smaller number of ALUs in your CPU. CPUs and GPUs have higher theoretical throughput in some cases but various parts of their ALUs typically spend a significant part of their time idle.
Contrary to marketing, you do not want to program these like an ordinary CPU even though the cores are truly general purpose (unlike a GPU). Thread behavior is unlike CPUs or GPUs. Barrel processors cannot saturate a single core with a single thread! The Xeon Phi has 228 independent threads and you need to use them all the time.
The way barrel processors work is if the hardware supports N threads then each clock cycle you can saturate all the ALUs if some subset M of those threads are not stalled. The M-of-N ratio varies by barrel processor design but is typically 20-50% in my experience. Each clock cycle, a core selects an immediately runnable operation from the basket of threads it can see and executes it; as long as something is runnable in that basket, the core will do real work that clock cycle. Xeon Phi has a 50% M-of-N requirement, so you need a minimum of 114 threads that are not stalled every clock cycle to saturate the processor. The way you ensure that you hit the 114 threshold is to schedule all 228 hardware thread slots with useful work.
For programming, this changes the way you reason about locality and concurrency. If we assume that some percentage of threads can be safely stalled or blocked with no impact on throughput then it changes the way you design your algorithms and data structures. A little additional latency on a subset of threads won't hurt performance, especially if it increases task concurrency. On a CPU stalled threads are expensive, as it leads to idle cores or context switches. For architectures like Phi, you design your data structures and algorithms around relatively small, semi-independently work units so that there is always a large number of tasks that can be assigned to a thread even though this reduces locality. Below a certain threshold, thread concurrency is approximately free because the cores will schedule around stalls due to contention.
I like barrel processors quite a lot. Once you get used to the model, it is an easier architecture with which to achieve efficient massively threaded parallelism and throughput than either CPUs or GPUs. More importantly, they are hard to beat for efficiency for general purpose computing when software is designed for the architecture since so few clock cycles are wasted.
Also, my understanding is that this model is a 270W TDP board with passive cooling. What kind of a machine can you actually install those in?
FWIW, I think I could fit two in the 4-way workstation I left in my parents' basement ;)
Also, think China or India.
As for the cooling question, this is designed for rack servers that have a good front to back forced airflow. No doubt people will try various tricks at fan mounting to run in a desktop PC. See one recipe here: http://openwall.info/wiki/internal/xeon_phi
There are places with heavyweight evaluation/adoption/etc. processes, but there are more agile ones as well. This kind of firesale does pretty well for that model.
[1]: https://software.intel.com/en-us/mic-developer#pid-22612-186...
- Black-Scholes Valuation Computing
- Weather simulation
fairly specialist number crunching I guess
[update]
I see the 'worlds fastest computer', Tianhe-2 uses 48,000 Xeon Phis. Wonder if they got a bulk discount?
The entire thing (with memory and disk) took up 1600 sq feet of floor space.
Here we are 17 years later, and we can get the same processing power in a pci card.
I'm certainly no expert on this stuff but I'm always on the lookout for affordable ways to make my projects easier to work with.
http://www.pugetsystems.com/blog/2013/08/06/Will-your-mother...
Here is a link to Intel's promotion page it has participating vendors in various locations:
https://software.intel.com/en-us/articles/special-promotion-...
:(