Intel's 72-Core “Knight's Landing” Xeon Phi Chip Cleared for Takeoff
hothardware.com
hothardware.com
[0] - https://software.intel.com/sites/default/files/article/33016...
[0] - http://www.anandtech.com/show/9794/a-few-notes-on-intels-kni...
The big difference besides memory and organisation is that the KNL cores are beefed-up Silvermont cores (the new Atom), which is an out of order core and pretty much what you would expect of a modern x86 architecture. The Knight's Corner cores were custom cores, somewhere between barrel and vector processors. The Knight's Landing should be even more general-purpose than the old Xeon Phi's and easier to program for.
EDIT: Looks like this is in the works already [0].
[0] - http://www.nextplatform.com/2015/03/25/more-knights-landing-...
The two main reasons Knights Landing is so competitive compared to the previous-generation Knights Corner is that (1) it doubled the raw compute performance per core thanks to its two 512-bit VPUs (vector processing units) per core compared to only 1 VPU per core for Knights Corner, and (2) it upped the number of cores from 61 to 72 while maintaining and even slightly increasing the clock frequency from 1.25 GHz to 1.3 GHz. All this was possible because Knights Landing is manufactured in 14nm while Knights Corner was 22nm, so its logic gates are 2-2.5x denser. Meanwhile all AMD and Nvidia discrete GPUs are still stuck at 28nm. The only reason GPUs still perform comparably to Knights Landing is that their execution units are simpler and smaller than Intel MIC cores.
What are people using the Xeon Phi for?
nVidia seems to be making a killing off machine learning applications. The entire GTC 2015 was about deep learning.
I assume that Xeon Phi is not nearly as effective for machine learning (?).
The Xeon Phi is a modern reincarnation of a barrel processor. It understands x86-64 opcodes but your data structures and algorithms need to be designed quite differently than vanilla CPU code if you want to optimize throughput. No one is doing that though.
In principle, if your code is correctly designed for the architecture, the Xeon Phi should significantly outperform both CPUs and GPUs for a wide range of use cases. The strength of these architectures is that their throughput is relatively insensitive to both latency and lack of trivial parallelism, which are the major bottlenecks to a lot of modern software performance. It is why Intel resurrected this style of computing architecture.
Basically, very few people know how to design software for these architectures even though it is pretty easy (much easier than GPU code). I have experience designing software for exotic architectures like this, I know what they are capable of in terms of throughput, but have never worked with a Xeon Phi. Nonetheless, the silicon specs suggest that someone that actually knows what they are doing should be able to significantly outperform e.g. GPUs for things like machine learning. Right now, it is basically being wasted because most developers treat them like weird CPUs.
I'd love to play with one of the new Xeon Phi processors to characterize its true performance but I am unlikely to see one. But I would not dismiss their performance; it is an extremely efficient kind of architecture for a surprisingly wide range of workloads if used well.
Would you be willing to expand on that? It sounds fascinating.
I'd also gladly take links or key words to google for.
I'm sure you can do better by hand than the code generated by the CUDA compiler. But people need something to get started with, and it's probably good enough for a lot of applications. In order to get adoption, the vendor has to meet developers closer to the application.
There are very few people writing assembly code from scratch anymore... and those that do are probably the kind of people who are designing their own hardware anyway!
It makes more sense for the vendor to be writing libraries in assembly, since they know the architecture best. Then apps can build on top of that in higher level languages.
http://shop.oreilly.com/product/9780124104143.do is an introduction and weights more than 400 pages.
Yeah. But I think it's good for Intel to follow through here instead of abandoning it like they did Larrabee. Intel has some real design talent, so it would be great if they can get their software team and hardware team to work on a solution that competes well with GPUs.
It doesn't help Phi that Intel's OCL support isn't quite there yet, IMO.
If they could just put their weight behind OCL and add their own extensions to support the cases where their accelerator design is different from GPUs, it would be attractive to HPC types to port their software to support it, IMO.
> What are people using the Xeon Phi for?
I don't know but I suspect it will shine for CPU-bound code that folks don't want to waste time porting x86 executables to a new language/framework. Memory-bound stuff will probably not see an advantage over GPUs. I don't think they'll attract enough attention to TBB/Cilk or whatever they're promoting alongside Phi.
(I'm assuming you were referring to the possibility of bending pins by trying to mount a processor incorrectly? Might be way off.)
It looks really beefy.
http://wccftech.com/amd-exascale-heterogeneous-processor-ehp...
http://www.pcworld.com/article/3003113/components-processors...
The contention about core count on the Bulldozer architecture seems spurious to me. At the time Bulldozer was being developed, multi-core systems were just being introduced to the consumer market. It was unclear what sort of architecture would be most performant. AMD made a (bad only in hindsight) bet that Bulldozer would be a viable architecture for general purpose compute loads. It turns out that combining 1 FPU with two integer pipelines is not as effective as an SMT architecture.
At the time Bulldozer was developed and released, what exactly constituted a core was still not precisely defined. It turns out that Intel's SMT architecture is much more effective, and thanks to market- and mind-share, people associate the definition of a core with Intel's specific implementation.
AMD's new development is on an SMT architecture known as Zen. Like all AMD news and marketing, it sounds exciting. Hopefully they execute well and it actually turns out to be exciting.
http://www.fudzilla.com/news/processors/38402-amd-s-coherent...
Intel Xeon Phi next year: 400 GB/s, 6 teraflops
NVIDIA Pascal next year: 1 TB/s, >10? teraflops
NVIDIA still wins for my applications. Intel is aiming too low.
I'm going to guess that your applications do some kind of streaming numerical processing or maybe some kind of large linear algebra. And, GPU cores are great that.
There's other applications that have a more complex work profile in terms of interleaving branchy logic with numerical processing. A good example of this might be a OLAP Database. They need do some query parsing, building a plan, optimizing that plan, work with indexes, do data decompression (traditional schemes and data specific schemes) then process that data, and do operations on it (join, filter, aggregate).
There's so many steps in that process and some of them require branching logic some of then numerical processing and some of them a combination of both. If you breakup a large query (partition) the Phi being good at both kind of computation and having fast RAM makes it an ideal platform to develop this.
So this means you can run the same code on a desktop Xeon as this one.
It doesn't make up for being a generation+ behind, but I expect they'll catch up over time.
It's a shame NVIDIA's Project Denver custom CPU didn't seem to work out very well. I was really hoping they'd be able to produce a CPU worthy of pairing with their top-end GPUs, to make an x86-free workstation and lessen their dependence on Intel. Wouldn't that be something?
Just a wild guess, but perhaps different people have different applications where GPUs don't fit in as well.
Are they offering something closer to general purpose CPUs, something that could schedule OS threads to run on it?
This is much more powerful than a GPU, which is fundamentally SIMD -- it can hit TFLOPS speeds running hundreds of threads doing different things. While not everyone will need this power, I suspect it will open up entire new classes of applications.
But this Xeon Phi is 8+ teraflops (single-precision) and 3 teraflops (double-precision).
So Knights Landing is a bit faster for SP, and ~15x faster for DP(!)