No Free Lunch for Intel MIC (or GPU’s)
blogs.nvidia.com
blogs.nvidia.com
I actually sat down with the director of a HPC shop a few months ago and discussed the MIC question.
He was hopeful, but said nearly everything about MIC isn't as mature as it needs to be to warrant adoption. His best case scenario was adopting MIC for his 2016 build.
Furthermore, while OpenMP apps probably will have scaling issues at 50+ cores, I think the bigger issue that the article hardly mentions is the new wide vector unit in the MIC cores. That's where all the FLOPS happen, and it's a completely new instruction set. x86 is a sideshow. Apps will need to be rewritten for the vector unit to get anywhere near peak performance on MIC.
The whole thing seems like a marketing-driven exercise in fake compatibility. Imagine a company with an old PalmOS app which needs to be ported to iPad if the company wants to stay relevant. They can't just wrap the Palm app inside an emulator and sell that -- it's just not good enough for a port because it doesn't leverage any of the new platform's strengths. That's pretty much what Intel is proposing.
"If you currently have access to MIC chips and have been testing
real applications, I would love to hear from you"
The information that was used to judge Intel's solution was based only on press releases and not any practical working knowledge. "We can no longer reduce voltage in proportion to transistor size"
The "we" here refers to Nvidia, and not Intel, because Nvidia uses TSMC to fab their GPU chips. "there is no such thing as a “magic” compiler that will
automatically parallelize your code"
That is funny because Nvidia's newest GPU design offloads scheduling decisions on to the compiler instead of handling it dynamically on-core.So perhaps Nvidia is warming up the FUD machines because of their newest range of GPUs which focus on gaming at the expense of computing.
eg. the Double Precision FLOP/s is now less than 10% of the Single Precision FLOP/s performance
Sometimes (I don't know about this case) the low end CPU's and GPU's are chips that have manufacturing errors in them and don't pass full quality control checks. The defective parts of the chip are then disabled with microcode (or equivalent) and sold as low-end chips. That's a good practice, a lot better than simply throwing the slightly defective chips to waste.
All of you who want inexpensive DP performance should look at AMD Radeon cards. AMD does not cripple any of its GPUs.
I hate that, but I guess it makes sense from a business perspective.
As for double-precision performance, it is also well known that the GTX 680 is only for consumers. The HPC version of Kepler is due later this year and I will not be surprised to see it offer much improve double precision performance compared to current Teslas. Intel MIC is not a consumer product, so it is unfair to compare it to consumer products like GTX 680.
[1] http://www.nvidia.com/object/openacc-gpu-directives.html
Wouldn't it possible to have a virtual machine that presents itself as a single CPU and distributes the job efficiently across the many cores that it is running on?
http://en.wikipedia.org/wiki/Hypervisor
In my very layman's opinion, this is a much different problem. Specifically the cost-benefit of discrete vs. MIC for calculations. HPC guys have tooled with it for a while trying to come up with a fast, yet cost effective method of crunching more data. Floating point acceleration specifically offers huge increases in speed for scientific modeling, but comes at a significant cost in terms of design, implementation and execution.
This NVIDIA fluff PR piece was trying to prove something to the effect of, "Don't use on-processor floating point, use our discrete solution!"
Imagine trying to make a virtual machine that would rewrite any bubble sort as a quicksort. Parallelizing algorithms is much more difficult than even that.
Imagine some layer in between an abstract syntax tree and micro-ops. Again, just thinking out loud.
Alot of decisions are made by non-technical or are technical but lack the depth of understanding.
e.g. propose intel MIC vs Telira (similar thing but with MIPS cores)