Why didn't Larrabee fail?
tomforsyth1000.github.io
tomforsyth1000.github.io
Oh god the horror that was OpenCL on Xeon Phi.
1) It was extremely buggy. Any OpenCL program of decent complexity was bound to encounter bugs in their driver.
2) FLOPS are not the only measure for computations / simulations. The bandwidth is is extremely important for many applications. Whatever Xeon Phis OpenCL was doing was not achieving anything close to its peak. Any OpenCL kernels we had tuned would end up improving the performance of the host CPU (16 x 2 Xeon) as well. It would usually result in the Xeon CPU performing better than the Xeon Phis.
3) NVIDIA released multiple generations of their GPUs since Xeon Phi was released. The performance per watt on NVIDIA and AMD GPUs has been improving substantially over the last 3 years while Xeon Phi's stagnated in 2012.
4) Even their claim for best FLOPS / WATT is wrong. The Tesla K20 released around the same time had the same power usage (225W) but about 50% higher FLOPS (2TFLOPS vs 3TFLOPS single precision).
The only selling point was that they were x86 cores and hence did not require rewriting your code. But you had to write additional directives indicating what part of your code you wanted to offload to your accelerator and when you wanted to get it back. The best performance ironically came after a lot of tuning for your application.
4 years down the line, GPU based accelerators still dominate the market and growing. and I wonder how many of the researchers who went with Xeon Phis regret their decision.
Traditional multi core CPUs simply aren't an option to advance in the HPC space any more. Per generation there's not enough of a performance improvement to make it effective. And buying more of them isn't feasible from a power perspective. So everyone is going to have to re-write for massive parallelism. I think Intel's KNL and future Knights will be the most attractive path for a lot of people. You have the tooling and familiarity of x86 and a simpler memory model than GPUs.
The true land of milk and honey is prophesied in Volta when paired with Power 9 and CAPI, first appearing in the Summit and Sierra systems.
The entire reason why GPU has been so successful, is that its performance is relatively predictable compared to x86 cores because of its much simpler architecture - and putting hundreds of them on a single chip doesn't change that. Granted, there are applications that will perform better on KNL than GPU (I expect), because it has a somewhat greater degree of freedom (although the bandwidth you get when actually using that freedom will be a deciding factor on whether it's actually worth it over the latest CPUs in those cases).
Do you still need to use vector instructions on KNL to get top performance?
Yes, but this is also the case for regular Xeons. If so then forgot about "familiarity" when rewriting.
When I say "familiarity" I'm referring to tooling in things like compilers and debuggers as well as just the instruction set.Not quite; the selling point was that they were x86 cores and hence did not require hiring graphics programmers to write your GPGPU code. You certainly did have to rewrite and refactor to get code shoved onto the Phi, but a non-HPC-focused team could do such a rewrite+refactor with just the people they already had.
x86 was and is still a weak selling point. Developers from all backgrounds are deploying their apps on varying architectures like ARM or the JVM without much stress. The hard part about writing code that is fast for an architecture is only made more complex by having your compute units be x86 rather than simple SP vector units of a GPU.
I know many folks without a graphics background writing cuda apps but no one outside of hpc research environments dabbling in the complexities of xeon phi.
The reason for this is simple; if you get your code to be in a cuda friendly structure you have created a data parallel rewrite that leverages the highly parallel memory bus of GPUs to get a pretty easy speedup. By being constrained to a semi opinionated programming interface people can see real speedups and not get bogged down in multithreading, multiprocessing, and buggy device drivers.
It's all well and good for them to aim for 'prestige buys' like Tianhe 2, but when the marketing smoke clears, the product is not achieving the things we in HPC buy products to achieve. Utilization rates are in the toilet across the ten or so Phi contracts I touch. I'm sure KNL's integrated architecture will change this, since everyone will just get this by default from vendors, but I don't know a single site who decided to buy Phi twice.
As far as I can tell, Phi is a play at a numbers game: you can ram up the Rmax of your system under LINPACK, regardless of whether it creates any actual value for your users. In other words, it's a cheap way to claw some false cred out of your top500 position, and not so much a useful HPC tool yet.
Time will tell, I guess.
You really can't not use the intel compilers for them. Gcc won't optimize well for it at all. This means that in addition to the higher entry price for the hardware, you have a sometimes painfully incompatible compiler toolchain to buy as well, and then code for. Which means you have to adapt your code to the vagaries of this compiler relative to gcc. These adaptations are sometimes non-trivial.
I am not saying gcc is the be-all/end-all, but it is a useful starting point. One that the competitive technology works very well with.
From the system administration side, honestly, the way they've built it is a royal pain. Trying to turn this into something workable for my customer has taken many hours. This customer saw many cores as the shiny they wanted. And now my job is (after trying to convince them that there were other easier ways of accomplishing their real goals) to make this work in an acceptable manner.
The tool chain is different than the substrate system. You can't simply copy binaries over (and I know they have not yet internalized this). The debugging and running processes are different. The access to storage is different. The connection to networks is different.
What I am getting to is that there too many deltas over what a user would normally want.
It is not a bad technology. It is just not a product.
That is, it isn't finished. It's not unlike an Ikea furniture SKU. Lots of assembly required. And you may be surprised by the effort to get something meaningful out of it.
As someone else mentioned elsewhere in the responses, the price (and all the other hidden costs) are quite high relative to the competition ... and their toolchain stack is far simpler/more complete.
The hardware isn't a failure. The software toolchain is IMO.
Hum not at all... Larrabee gets destroyed by the competition in terms of flops-per-watt. Knights Corner is rated 8 GFLOPS/W at the device level (2440 GFLOPS at 300 W). For comparison Nvidia Tesla P100 rates 4× better: 32 GFLOPS/W (9519 GFLOPS at 300 W) and AMD Pro Duo rates ~6× better: 47 GFLOPS/W (16384 GFLOPS at 350 W although in the real-world it probably often hits thermal limits and throttles itself, so real-world perf is probably closer to 30-40 GFLOPS/W).
Also, if Intel wants to make Larrabee gain mindshare and marketshare, they need to sell sub-$400 Larrabee PCIe cards. Right now, everybody and their mother can buy a totally decent $200 AMD or Nvidia GPU and try their hands at GPGPU development. And if they need more performance, they can upgrade to a multiple high-end cards, and their code almost runs as is. But because Larrabee's cheapest model (3120A/P) starts at $1695 (http://www.intc.com/priceList.cfm), it completely prices it out of many potential customers (think students learning GPGPU, etc).
While I expect the graphics cards to have the edge in performance per watt for brute-force math, I'm surprised that the current GPU's would have that large of an advantage. Are you sure you are comparing the same size "flops" for each?
And while it's in keeping with the article, I don't think lumping all the generations together as "Larrabee" makes sense when comparing system performance. While availability is still minimal, Knights Landing is definitely the current generation (developer machines are shipping), and as you'd expect from a long-awaited update at a smaller process size, efficiency is a lot better than older generations.
Here's performance data for different generations of Phi: http://www.nextplatform.com/wp-content/uploads/2016/06/intel...
And here's for different generations of Tesla: http://www.nextplatform.com/wp-content/uploads/2016/03/nvidi...
From these, the double-precision figures appear to be:
2016 KNL 7290: 3.46 DP TFLOPS using 245W == 14 DP GFLOPS/W.
2016 Tesla P100: 5.30 DP TFLOP using 300W == 18 DP GFLOPS/W.
I don't know if these numbers are particularly accurate (I think actual KNL numbers for the developer machines are still under NDA), but I think the real world performance gap will be something more like this than one approach "destroying" the other. If you are making full use of every cycle, the GPU's will win by a bit. If your problem requires significant branching, Phi won't be hurt quite as badly.
Also, if Intel wants to make Larrabee gain mindshare and marketshare, they need to sell sub-$400 Larrabee PCIe cards.
The interesting move Intel is making is to concentrate initially on KNL as heart of the machine rather than as an add-in card. This gets you direct access to 384GB of DDR4 RAM, which opens up some problems that current graphics cards are not well suited for. I think this plays better to their strengths: if your problem is embarrassingly parallel with moderate memory requirements, modern graphics cards are a better fit. But if you need more RAM, or your parallelism requires medium-strength independent cores, Phi might be a better choice.
But because Larrabee's cheapest model (3120A/P) starts at $1695
While it's true that Intel's list pricing shows they aren't targeting home use at this point, List Price may not be the best comparison. For a while last year, 31S1P's were available at blow out prices that beat all the graphic cards. In packs of 10, they were available at $125 each: http://www.colfax-intl.com/nd/xeonphi/31s1p-promo.aspx.
It's true that even at that price they were not a great fit for most home users: they require a server motherboard with 64-bit Base Address Register support (https://www.pugetsystems.com/labs/hpc/Will-your-motherboard-...) and they require external cooling. But if you were willing to be creative (https://www.makexyz.com/3d-models/order/b3015267f647354f96ff...) and your goal was putting together an inexpensive many-threaded screaming number cruncher, the price was hard to beat.
KC: 8 GFLOPS/W at the device level (Xeon Phi 7120: 2440 GFLOPS at 300 W).
KL: 28 GFLOPS/W at the device level (Xeon Phi 7290: 6912 GFLOPS at 245 W).
Compare this to 32-47 GFLOPS/W at the device level for current-generation AMD and Nvidia GPUs.
(All my numbers are SP GFLOPS.)
It'd be awesome if Intel made such a machine available the same way they make NUCs. Everyone vaguely involved with HPC should have one to play with.
It might be possible to pull off similar tricks as on normal Intel Xeons. I wonder what kind of memory controller it got. Does it for example have hardware swizzling and texture samplers accessible?
But we use exclusively Nvidia for our number crunching product. And since the relevant code is implemented in CUDA, we have an Nvidia platform lock-in. Other options are not even on the table, and will not be.
Intel would be very smart to offer a migration path from CUDA code.
For me, "Larrabee" died during it's first public demo in IDF 2009 [1]. At the time I was working on CUDA libraries at NVIDIA. I remember everyone watching the stream to see how much of a threat it would be. When we saw the actual graphics, everyone started laughing. "Welcome to the 1990s!" At that point it was obvious that Larrabee would not be a graphics threat to NVIDIA, it just had too far to go. It was not a discrete GPU killer.
You could probably install FreeBSD on it.
I think you should do some reading on what an operating system is.
However I don't know if that is the case here.
From a KNC card:
Linux <hostname>-mic0 2.6.38.8+mpss3.4.3 #1 SMP Mon Feb 16 16:08:55 PST 2015 k1om GNU/Linux
From KNL:
Linux <hostname> 3.10.0-229.20.1.el6.x86_64.knl2 #2 SMP Tue Dec 8 22:27:38 MST 2015 x86_64 x86_64 x86_64 GNU/Linux
What I hear from my friend who works in the GPU industry, one of the main outcomes from the Larrabee research project was: It is possible to do most of graphical operations on a SIMDed general purpose CPU, with good performance. Everything except texture decoding, which really needs dedicated texture decoding silicon. And with texture units taking up 10% of the silicon, Intel really needs to decide if they want to sell it as a GPU or as a general purpose compute unit with 10% more cores. Intel chose the latter, and you can't really blame them, as you can sell dedicated compute units for more money.
Sony ran into the same problem with the PS3 and the cell. They originally designed it so the game developers could implement whatever rendering method they wanted, in software, on the SPUs. But the performance wasn't high enough, partly due to texture decoding taking up too much time. By the time it was discovered this was a problem, the cell was more or less finalises. Sony were considering adding a second cell to the console, to brute force through the problem, but eventually they asked Nvidia to hack together a traditional GPU.
The cell now has 7 vector units, with comparatively more memory, but there was no default job for them, all the vertex transformation now ran on the GPU's vertex shaders. And Sony initially stuck to their guns of "SPU programs should be written in assembly, in a spreadsheet"
Because the single PPU really sucked and was nowhere near fast enough for anything, Sony eventually relented and releases a version of GCC which would compile c++ code to the SPUs. Fast to develop, but nowhere near the performance of an excel spreadsheet designed SPU program.
This resulted in a whole bunch of games running code on the SPUs that was really badly optimised. But at least it reduced load on the PPU.
What was supposed to be the advantage of this over a traditional plain text assembly language? Using Excel macros to generate code?
Both Architectures had exposed pipelines, meaning the result of an operation would take a few cycles to show up in the destination register and some operations would take longer than others. You might have to insert a bunch of NOPs to make sure the data would be ready for the next instruction that needed it. Both Architectures were also dual issue, meaning two completely independent operations, operating on completely independent registers would be manually packed into a single instruction by the programmer. There also would be restrictions on which types of instructions could go in each half of the instruction, if you didn't have an instruction, you have to put a NOP there.
I'm pretty sure Sony liked the spreadsheets because it forced the programmer to see where all the NOPs were. The programmer would be expected to refactor things and manually unroll loops until all the NOPs were filled with useful instructions and peak performance was reached.
https://www.youtube.com/watch?v=4ZFtP8LbUYc (An Uncharted Tech Retrospective)
From the article:
> Remember - KNC is literally the same chip as LRB2. It has texture samplers and a video out port sitting on the die. They don't test them or turn them on or expose them to software, but they're still there - it's still a graphics-capable part.
So the core space is still used right? They didn't choose 10% more cores, they just chose to turn it off an not test it, but it still uses the die space.
It was a long term advantage, they removed the texture samplers for the next version, knights landing, which allowed them to fit more cores on that chip.
I remember by the time Gran Turismo was launched, there was an event showing the new tooling to developers.
For the typical case of 2D uncompressed textures with some kind of trilinear/anisotropic filtering enabled, you are basically just calculating 8 addresses based on texture coordinates and clamping and mipmap levels, doing a gathered load (and CPUs hate doing gathered loads). Remember to optimise for the case that each group of 4 addresses are typically, but not always right next to each other in memory.
Then you use trilinear interpolation to filter your 8 raw texels into a single filtered texel. With the exception of the gathered load, none of these operations is actually that expensive on CPUs, but shaders do a lot of these texture reads and it's really cheap to and faster to implement it all in dedicated hardware.
You can also put these texture decoders closer to the memory, so the full 8 texels don't need to travel all the way to the shader core, just the final filtered texel. And since each texture decoder serves many threads of execution, you have chances to share resources between similar/identical texture sample operations.
And while the texture sampler is doing it's thing, the shader core is free to do other operations (typical another thread of execution). It's not that the CPU can't decode textures, its just that CPU cores without dedicated texture decoding hardware can't compete with the hardware which has the dedicated texture decoding hardware.
leave a openCL layer there for the unwashed masses, but provide a decent API for anyone wanting exactly what the product is.
If they had went full on OpenCL, including a unified programming for vector and multicore parallelism (the way CUDA works), I think Phi might have taken off much more already.
From my knowledge that claim only makes sense for the ever-shrinking domain of applications that haven't been ported to GPGPUs with _simpler_ cores. I also don't understand how the best "flops-per-watt" can be extrapolated from the fact that the Phi is used in top ranked TOP500 machines.
The Phi really just seems like an intermediate between CPUs and GPGPUs in the ease-of-programming vs. perf/watt pareto curve.
I see NVIDIA near the top at 4.7 GFLOPS/W and the old Xeon Phi at #44 at 2.4 GFLOPS/W which supports your argument.
Now, though, it sounds as if Intel may be making the opposite mistake. The Larrabee/KnightsThingummy family was built around x86 cores to make it easier to port existing code to it, but now it looks like this time people really are willing to rewrite their code to work on GPGPUs after all.
- The perf/watt has not been competitive
- For the most part it has not been loved by its users
- The market share/penetration is small
- There is no evidence sales have been material to Intel; In fact it's not even clear it's been profitable
On the upside, though the product has failed to date, it's not dead. Maybe the 2016 release can turn things around.
Also it's nice that AVX512 came out of Larrabee, but the last line claiming credit to a single person seems dubious.
I think we can state it did fail as a graphics processor.
http://www.anandtech.com/show/2580
http://www.anandtech.com/show/3738/intel-kills-larrabee-gpu-... http://www.cnet.com/news/intels-larrabee-more-and-less-than-...
http://www.pcper.com/reviews/Graphics-Cards/Larrabee-New-Ins...
Read the post carefully.
GDCE 2009, "Rasterization on Larrabee: A First Look at the Larrabee New Instructions (LRBni) in Action"
http://www.gdcvault.com/play/1402/Rasterization-on-Larrabee-...
There were another Larrabee talks, but only this one is online.
The marketing message was not only about graphics, but how the GPGPU features of Larrabee would revolutionize the way of writing games.
in terms of engineering and focus, Larrabee was never primarily a graphics card
Note the author says primarily and is specifically talking about it in terms of engineering and focus.
At GDCE it was being sold as a way of doing graphics, AI and vector optimizations of code in more developer friendly than GPGPU.
Of course, framing the graphics feature is a nice way of sidelining the issue that it also didn't delivered those other features to the games development community.
> Of course, framing the graphics feature is a nice way of sidelining the issue that it also didn't delivered those other features to the games development community.
There is clear evidence that the high-end cards that NVidia delivered at these time could outcompete (or keep pace) with Larrabee at that time. Thus when released Larrabee would not have been a strong contender to NVidia or AMD at that time. In this sense I stand by my position that Larrabee failed as GPU.
On the other hand I see no evidence that the rival products by AMD and NVidia could keep pace with Larrabee for AI and vector optimizations. Thus Larrabee was not a failure here. So Intel probably just concluded that in HPC there is much more money to be earned than for consumer devices and thus Larrabee was not released to consumers. And I can see good reasons: If game developers want to exploit the capabilities Larrabee has to offer, they have to depend on consumers having a Larrabee card in their computer. If Larrabee were an outstanding GPU the probability that some enthusiasts will get it was much higher than if Larrabee is just an add-on that some exotic applications/games additionally require.
Hint: For a page with nothing but text content you're doing it wrong. Even more so since this is a static site and you have to go to extra effort to break normal degradation of HTML content.
Personally I also browse with javascript disabled and thought the mandatory JS was silly. Hence I did view source, and lo and behold it is a tiddly wiki, with 2MB of html containing every article.
While github now has pages.github.com and the ability git pull/push your wiki (git clone git@github.com:username/project.wiki.git), I'm not sure how well that works for using on another web server or locally.
Why is this the case? It seems counter-intuitive given the recent Nervana acquisition (presumably to compete with NVIDIA)
When I say "Larrabee" I mean all of Knights, all of MIC, all of Xeon Phi, all of
the "Isle" cards - they're all exactly the same chip and the same people and the
same software effort. Marketing seemed to dream up a new codeword every week,
but there was only ever three chips:
Knights Ferry / Aubrey Isle / LRB1 - mostly a prototype, had some performance gotchas, but did work, and shipped to partners.
Knights Corner / Xeon Phi / LRB2 - the thing we actually shipped in bulk.
Knights Landing - the new version that is shipping any day now (mid 2016).
I can kind of see what the author is getting at here but it's worth reminding
general readers KNL is a wildly different beast from KNF or KNC.1. It's self hosting, it runs the OS itself/ there is no "host CPU".
- Yes KNC ran it's own OS but it was in the form of a PCI-E card which still had to be put into a server of some kind.
- Yes KNL is slated to have PCI=E card style variants but the self hosting variants are what will be appearing first and I honestly don't think the PCI-E variants will gain much if any market traction.
2. It features MCDRAM which provides a massive increase in total memory bandwidth within the chip. Arguably this is necessary with such large core counts.
3. Some KNL variants will have Intel's Omni-Path interconnect directly integrated which should further drive down latency in HPC clusters.
Behind all that marketing, the design of Larrabee was of a CPU with a very wide
SIMD unit, designed above all to be a real grown-up CPU - coherent caches,
well-ordered memory rules, good memory protection, true multitasking, real
threads, runs Linux/FreeBSD, etc.
Here's an interesting snippet, drop "Larrabee" from the above and replace it
with "Xeon". I'm just wanting the point here that the Xeon and Xeon Phi product
lines are only going to look more similar over time. For a while now the regular
Xeons have been getting wider. Which is to say more cores and fatter vector
units. The Xeon Phi line started out with more cores but is having to drive the
performance of those cores up. Over time we're going to see one line get
features before the other e.g. AVX512 is in KNL and will be in a "future Xeon"
(Skylake). And it's widely rumored that "future Xeons" will support some form of
non-volatile main memory, something not on any Xeon Phi product roadmap today.
The main differentiator will be that Xeon products need to continue to cater to
the mass market whereas Xeon Phi can do more radical things to drive compute
intensive workloads (HPC). Larrabee, in the form of KNC, went on to become the fastest supercomputer in the
world for a couple of years, and it's still making a ton of money for Intel in
the HPC market that it was designed for, fighting very nicely against the GPUs
and other custom architectures.
Yes at the time Tianhe 2 made it's debut it was largely powered by KNC and it
placed at number 1 on the top 500.Did KNC as a product make a lot of money for Intel? Probably not. Yes they shipped a lot of parts for Tianhe 2 but for other HPC customers, not so much. At the time putting a KNC card against a Sandy Bridge CPU it wasn't leaps and bounds faster. Yes you're always going to have to re-factor when making such a change in architecture but the gains just weren't worth it in most cases. Not to say there haven't been useful deployments of KNC that have contributed to "real science", but it was not a wild success by any measure.
As for "fighting very nicely against the GPUs and other custom architectures", I don't think so. When KNC got to the market Nvidia GPUs already had a pretty well established base in the HPC communities. Many scientific domains already had GPU optimized libraries users could pick off the self and use. Whilst KNC was cheaper than the Kepler GPUs on the market at the time it wasn't much cheaper. Again back to regular Xeons, for a cost/ performance benefit KNC wasn't worth it for a lot of people.
Its successor, KNL, is just being released right now (mid 2016) and should do
very nicely in that space too.
I do agree with this, KNL is going to do very well. There are whole systems
being built with KNL rather than large clusters with some KNC or some GPUs. 1. Make the most powerful flops-per-watt machine.
Not sure what is mean here, most efficient machine in the top 10? Tianhe 2 is
not an efficient machine by any stretch of the imagination. For reference here's
the top 10 of the Green 500 for when Tianhe 2 placed as number 1 on the Top
500.http://www.green500.org/lists/green201306
SUCCESS! Fastest supercomputer in the world, and powers a whole bunch of the
others in the top 10. Big win, covered a very vulnerable market for Intel, made
a lot of money and good press.
Yes KNC provides the bulk of the computing power for Tianhe 2 but at the time of
writing of the above is number 2, not number 1. Secondly that is the ONLY
machine in the top 10 that uses KNC AT ALL.I'm sure emulating OpenGL on the CPU has it's uses outside of being a reference for drivers and such - but it's not really that exciting, GPUs are already excellent at doing that kind of rendering.
When it comes to ray-tracing in particular, as far as I know there really doesn't exist (except perhaps as research prototypes) any hardware platform that's great at ray tracing. CPUs are okay but they need more parallelism. GPUs are okay but (as far as I understand it) they aren't good with lots of branches and erratic memory accesses. It seems like KNC/KNL ought to be ideal for this kind of thing (and in fact, it appears a fair bit of effort has gone into optimizing Embree for Xeon Phi).
Now with Metal, DX12 and Vulkan GPU languages, maybe that is ever more approachable, specially in cards like Pascal.
I had a number of long conversations with D about soft texturing; I didn't understand Tom Forsyth's arguments about why soft texturing wouldn't work. (My back-of-the-envelope numbers were encouraging.) So, in winter of 2011 (12?) I decided to write my own soft texturing unit. What I found was that soft texturing is completely do-able---just not in any sort of reasonable power budget.
At the same time a friend, S, decided he wanted to implement a rasterizer (I'd gotten my fill). Between the two of us, we had a rockin' still-frame bilinearly sampled Stanford bunny. A third friend, W, decided to implement a threading model. Threading models for very-high count processors are enormously difficult. I ended up writing much of an OpenGL driver, a shader compiler (through LLVM), and part of a pipeline JIT, i.e., a vertex loader, fixed-function-to-programmable compiler, etc. That project was called SWR; the sanitized projection is OpenSWR.
(Looking at the current sources, I'd say that I probably no longer have any code in there.)
I see what you did there. :)
Searched for 'Xeon Phi' on Newegg just now and only desktop CPUs show up. 'GTX 1080' brings up 59 results.