The New Intel: How Nvidia Went from Powering Video Games to Revolutionizing AI
forbes.com
forbes.com
1. Add "good enough" functionality to its high-end processors. Intel is already working on that: https://news.ycombinator.com/item?id=12709220
2. Contribute open-source code to the main branches of the most popular DL/AI frameworks (Tensorflow, Theano, Torch) so these frameworks support the new chip functionality "out of the box," without requiring any additional tweaking. This is not yet happening, but I'm hoping it will soon.
Many DL/AI developers would be content with "good enough" performance out-of-the-box from CPUs if it means not having to pay extra for Nvidia cards or deal with Nvidia's proprietary drivers.
It would be great for Nvidia to get real competition in this space.
It seems obvious AMD will be a player, since they have similar hardware (GPUs) and with a proper port (using HCC, openCL, etc. ) for cuDNN should also have excellent performance for Tensorflow, Theano, etc. 'out of the box'.
NVIDIA knows how to do 2 things right; understanding that software matters, even more than hardware and deliver on that notion; and how manage developer relations.
While AMD makes it easy to access information without NDA and sign ups NVIDIA blows it it if the water the moment you show even the slightest interest at signing up with them.
NVIDIA at this point can probably spin off a pure consulting division for AI/ML and bring in more revenue than AMD in its entirety.
> NVIDIA at this point can probably spin off a pure consulting division for AI/ML and bring in more revenue than AMD in its entirety
AMD had $3.99 Billion in total revenue in 2015. Nvidia had 5 Billion.
Perhaps you mean profit? Amd was operating at a loss and nvidia was very profitable.
The compute capabilities are pretty much the same for the latest NVidia/AMD cards in the same price range. It's a software thing - lack of optimised DL ops, particularly fast convolution kernels that is hurting the perf of AMD in deep learning
The main reason people don't use AMD for deep learning is lack of fast libraries, and AMD's inability to provide a CuDNN equivalent optimised for their hardware
>>We don't happen to have the resources to pay someone else to do that for us.
It makes zero sense for them to place increasingly large bets on proprietary drivers when they're clearly not the market incumbent.
yes and they are priced accordingly ;-)
"This time is different" might apply here though, because training is "just" a lot of low precision matrix math. This requires a lot less new ecosystem / software for Intel (Intel has a long history in matrix math), so much like SSE, AVX, and now AVX-512 (or as I still call it LRBni) Intel can easily make some tweaks to at least get 70% of the gains.
The threat to Nvidia from Intel is in the Nervana chips they recently acquired. Those are presumably using HBM2 and could potentially beat GPUs for neural net training performance.
If future Xeons integrated RAM like Knights Landing does, this would make GPUs less interesting. But KNL can already be used as the main CPU, if I remember correctly.
Intel's offering costs as high as €4000 per processor, which come with a meager 8 cores.
NVidia's offering sells, right now, for less than €1000 a pop.
Furthermore, the performance of each of NVidia's GPUs falls somewhere between 5 and 20 teraFLOPs. A Haswell Xeon gets you about half a teraFLOP for around 1/3 of the power consumption of a NVidia Tesla P100 GPU.
Comparing GPUs with Xeons makes no sense.
On the Nervana front, outrunning a GPU for neural nets is not that hard with an ASIC. I checked your profile, I bet your employer knows something about that :)
I'd recommend reading this paper about matrix math on GPUs from friends of mine way back in the day: https://graphics.stanford.edu/papers/gpumatrixmult/gpumatrix... . While NVIDIA has built much larger register files and L2 caches since then, a modern Xeon still is unbeatable when something fits in L2 or even L3 cache.
I believe streaming the input data in to update the weights is usually done with non Cache polluting instructions (mm_stream equivalents with the NT hint), so it's not hard to keep it fed.
Trying to peg this onto one of their CPU's will have a result much like they have had in pegging the graphics processing stuff onto their CPU's. That is, very meagre results that is fitting for only the lowest resource intensive things that a consumer may want to do.
30MB to hold a model in is almost nothing. Even 12GB is insufficient to train something like imagenet.
But in the same vein as their iGPU implementation, it may be just enough to satisfy the requirements of the average consumer, and still be able to make a large dent in Nvidia's market share.
But while FP16 is useful for audio and imaging, FP16 nearly killed NVIDIA over a decade ago* when it lacked the dynamic range for DirectX 9 HDR effects in contemporary games without banding. FP32 was more than enough for the task, but power hungry, and thus NV30 could be used to figuratively fry eggs while AMD GPUs had FP24, which was just enough for these effects.
These days, it's all going in the opposite direction w/r to deep learning, but I find it ironic that INT8/INT16 is missing from the Tesla flagship P100, but present on GP102 and GP104, the consumer GPUs (and yes, I know about Tesla P40, but that lacks fast FP16).
I agree with the top poster that Intel could make quite a comeback here given how far it's currently behind. I also agree that it would be hard to dethrone NVIDIA without higher bandwidth memory, but I don't think that's necessary, a bloody nose is more than enough to turn heads IMO.
Also backpropagation makes you iterate through the whole model at each step
Intel made massive gains promoting OpenCL, when they helped to redo OpenCV's umat infrastructure.
Intel is trying to compete with NVidia by offering inferior performance at a premium price.
In massively parallel applications, cost is closely tied to performance. NVidia has both aspects secured very well.
In fact, NVidia is already a prominent part in supplying components for the 3rd place in the Top500, Oak Ridge National Laboratory's Titan system, which mixes up AMD Opterons with NVIdia k20x GPUs.
...except that Intel has been trying to do that for over a decade, and Intel has been failing repeatedly for over a decade.
If that was as simple and easy as you lead to believe, Intel would've already done it by now.
Except it didn't, because it can't.
Xeon Phi is their contender, but we'll have to see enough
A high-end GPU is around 10x faster than your avg CPU https://www.nvidia.com/object/gpu-accelerated-applications-t...
http://files.shareholder.com/downloads/AMDA-1XAJD4/341168899...
Also: The level of their gross profit margin indicates that they are in fact a monopolist:
http://marketrealist.com/2016/11/driving-nvidias-profit-marg...
(These are not the financials of company operating in a healthy, 20-year old industry - they are the financials of a company that after operates without meaningful opposition in a 20 year old business.)
Many of Nvidia's "gaming" cards (e.g., higher-end GTX cards) are in fact used NOT for gaming but for deep learning by numerous AI researchers, developers, and startups with small budgets who buy them through retail channels.
a) AI/startups
b) gamers
what would you guess?
My guess is somewhere between 1:50 and and 1:500.
Imagine a world where every car has semi-autonomous technology, the average security camera is doing object detection, localization, more sophisticated language translation, smart drones, etc. The potential for deep learning is undeniable and Nvidia (currently) is leading this race.
Nvidia is pushing PX2 but I have a feeling they are too expensive for most cars - don't know the exact price but "a few thousand dollars" was quoted for the Teslas that use them. Some specs - http://wccftech.com/nvidia-drive-px2-pascal-gtc-2016/
At the same time Qualcomm is pushing their much less performant chips for self-driving cars, but at a much lower price point (hundreds of dollars)
The key here isn't that this isn't a full power desktop GPU, but it will run inference on a CUDA-compatible kernel.
That give NVidia a complete end-to-end modelling, training and driving platform. No one else has that.
It seems if they are used in development the sales don't scale anywhere near directly with the number of units sold.
Maybe I'm missing something?
Nvidia, meanwhile, has built a massively parallel beast that's still decent at gaming. But make no mistake; their engineering priorities are set on general parallel computation. Gaming performance is in many ways a similar problem, and thus benefits, but it would be a mistake to think games still come first for their engineers.
Intel ship about 2/3 of the graphics chips for laptops and desktops. Nvidia have more than half of the remainder. That isn't a monopoly.
In the discrete GPU space AMD are also there and ship almost half as many unit as Nvidia. That's no monopoly.
The graphics systems for the current consoles are shipped by AMD aren't they?
Nvidia is a profitable company. Calling them a monopolist because they are profitable is unwarranted.
* Edited to fix grammar.
That said AMD hasn't managed to ship a 350-400$+ card that's worth spending money on compared to the competition since probably the 7950...
Nobody is claiming Tesla has a monopoly on electric cars because they make the best. And nobody should claim Nvidia has a monopoly on GPUs since they are on top.
And AMD will have again a very expensive card to manufacture GDDR5/5x vs HBM2 with high power consumption based on the fact that the RX480 draws nearly as much power as a 1080 atm, and HBM2 isn't exactly power efficient (unlike HBM1 vs GDDR5), HBM2 is nice but the power requirement currently increase with the density, and it leaks voltage pretty badly with 4 stacks...
I wouldn't be so sure that the comparison to the 40-year behemoth of silicon, which extracted quasi-monopoly profits for 35 of them, is yet valid.
But for mixed, semi-sequential/semi-parallel workloads, I/O becomes the bottleneck. Having to retrieve intermediate results, combine and then redistribute can eat up your gains from parallelisation. It will be interesting to see if Nvidia start pushing for a new standard to replace PCIe, and maybe invest in low-latency, high throughput networking R&D. Who knows? Maybe they'll just build some kind of networking functionality in to their GPUs directly, so data can be transferred NVRAM -> NVRAM without having to travel up and down the stack (until they need to interact with userspace; once on the way in and once on the way out).
Given there are RDMA adapters that can manage ~20ns latency already, it would be interesting to see what could be achieved before we start pushing up against the laws of physics. Even at the physical layer, my understanding is that most networking fibre optics manage about 60% the speed of light.
So, just to make a wild prediction: within the next year, Nvidia will acquire a HPC network adapter company (e.g. Mellanox).
x86 systems like the DGX-1 can only run NVLink between the GPUs, and still rely on PCI-E for the CPU connection.
I assume it will be infiniband that will be the commodity hard adopted to distribute Deep Learning. However, Nvidia are currently trying to segment the market by providing GPU->GPU direct access via infiniband on only the Tesla cards - not the commodity ones (titanx, 1080 gtx). That strategy will probably be the death of them in Deep Learning.
Full disclosure: I am (and have been since I found about about cudnn) long nvidia.
HPC -> Mainstream servers (and niches like ML) -> 'Professional workstations' -> Consumer desktops
Intel's response (at least in part) seems fairly sensible: throw money and developers at open-source and open-standardisation efforts to try and rally the rest of the market (e.g around stuff like OpenCL). Otherwise network effects from stuff like widespread CUDA usage will become too strong to overcome, at which point Nvidia will have a monopoly.
However, I think they're missing an opportunity when it comes to high-speed networking. Specifically, I think they should be looking to release cheap, simple and open-standards based networking gear without trying to cross-sell Xeons or whatever at the same time. I realise that this means cannibalising revenue from their CPU market, but at this point someone is going to do it so it may as well be them. Hell, they could probably even get away with tying the functionality in to some fairly widespread Intel CPU instruction set (say, from Ivybridge onward). However, anything past that and it's a non-starter for consumer segments (which, in my mind, is also the developer segment).
10gbit RDMA is actually sort of affordable nowadays (at least if you pick up gear from eBay). But it's currently a complicated mess of competing and incompatible standards. Some SFP+ modules are even vendor-locked; if they sense there's a different SFP+ branded module on the other end they stop functioning due to an 'unsupported configuration'. These kinds of things make it impossible for a decent sized consumer market to form. If Intel could just accept that the world is moving on with or without them and get a proper consumer market going, at the very least they'd keep the market open (and at the very best the might even tilt the odds in their favour).
AMD, on the other hand, I have no idea wtf they are doing (and neither do they, by the looks of it). It could just be my biases, but I really think the best strategy for them is to basically bet the house on open-source and really up their engagement with Linux core dev. While I get it's attractive to target the gaming market because it's profitable in the present, it probably means death in the future. Actually their interests are pretty aligned with Intel's, and they have highly complementary work forces. It sounds crazy, but I wouldn't be too surprised if Intel were to acquire AMD at some point, or at least buy out their GPU division...
(Disclaimer: I'm the system architect of InfiniPath, which is one of the things that evolved into Omni-Path.)
As I mentioned in another post, I really think Intel are missing an opportunity when it comes to high speed networking. Why don't they drop all the crazy server market lock-in vendor shenanigans (e.g. brand to brand SFP incompatibility, 'call us' pricing etc.) and just jump straight to the consumer market? I know this means cannibalising CPU revenue, but that's going to happen anyway.
Developers are consumers. In many cases the direction that their skills develop in is largely determined by the consumer hardware they're running at home. The same seems to be true for Ops folks, who all seems to be running vSphere homelabs atm. Based on what others have said here, it sounds like NVIDIA have chosen to cross-sell, and work their way down from HPC -> ... -> Consumer, presumably to extract maximum rents, despite not yet having strong enough network effects o lock out competition.
This seems like the perfect opportunity for Intel to cut the legs out from under that strategy, by aggressively pushing widespread consumer (or at least 'power-user'/developer) take-up of high speed, low-latency networking. And hasn't iWARP been endorsed as an IETF standard (or is on its way to be)? It seems like there's a window of opportunity here, albeit one that gets smaller and smaller as Nvidia slowly work their way down market.
But if you want to buy the PCI Express cards, Newegg has 'em. Intel has a great channel organization.
- ~$8000 USD for a 24 port switch is way beyond my price range. I doubt any consumer market could support this (http://www.newegg.com/Product/Product.aspx?Item=9SIA6ZP53G35...)
- ~$550 USD for a single port adapter is also way outside of my budget, especially since I'd need to buy at least half a dozen of these for my home lab (http://www.newegg.com/Product/Product.aspx?Item=N82E16833106...)
- It's very difficult to find solid information on the product. The Ark page doesn't tell me anything (http://ark.intel.com/products/92007/Intel-Omni-Path-Host-Fab...). The link to the whitepaper (which I'm assuming is actually a marketing brochure) 404s.
- So even if I could afford this as a consumer/dev, I probably wouldn't buy it due to the uncertainty around total cost of ownership. Will it accept non-Intel branded QSFP transceivers? Same question with optical cabling.
- Given the product brief mentions 'fabric performance [will] scale automatically with ongoing advances in Intel Xeon processors...', it leaves me wondering if this is somehow dependent on some specialised Xeon CPU instruction sets, making it useless for most home labs (which, at best, will be running older generation Xeons from second-hand servers, but more likely CPUs from Intel's consumer line-up).
On costs: If these are just reflective of the cost of production, then fair enough, I guess I'll just have to wait (for either Intel, or another manufacturer, to devise a cheaper production process). If it's not, then I think you're missing the opportunity outlined above, and you'll lose the HPC fight given Nvidia occupy the high ground here.
On compatibility: If this is just some idiot 'multi-channel cross sales' strategy that some genius with an MBA has cooked up, you're going to have a bad time. It falls apart the second someone releases a commodity adapter (which may already exist, for all I know).
It's also why I figure Infiniband would be a likely acquisition if Nvidia wanted to head in that direction (i.e. GPU to GPU networking a.k.a NVRAM RDMA).
Where is Intel heading? I dont see Intel gaining a foot in the AI / ML market. Not Knight XXX. And all these supposedly weapon such as next gen AVX, Nervana aren't coming any time soon. And Nvidia knows full well, the same thing about GPU and AI/ML isn't in the silicon, it is in the drivers and library.
So not only this is a huge first mover and mature software advantage, the AI / ML is difference in gaming, where the human cost of Nvidia AI/ML knowledge is significantly more then what the Hardware is worth, which is next to nothing when compared to those salary.
So in the next 2 - 3 years, I dont see Intel getting much from Nvidia. They failed Mobile. Windows 10 is now slowing working towards ARM.
It leaves them with 1), a shrinking, or downtrend PC market. 2) Their Server CPU market with heavy margin which will finally be getting a lots of competition from AMD Zen x86 and Qualcomm ARM Server Chip.
Zen Server chip is highly unlikely to be competitive against intel in performance. But there are lots of margin for AMD to attack. It means a lot of pressure from Intel to lower price.
A lot of ARM Server will be coming in the next few years. Not low end blade, but powerful Xeon like chips. I have always been skeptical of this, but it seems ARM, which is owned by Softbank, whom is also the largest shareholder of Alibaba, whom also operate a gigantic Cloud infrastructure, are All in on this.
This is not saying Intel will die in 5 years time. They are likely to be around for decades after, but i dont see where they could grow and head.
0. Finally paying attention to software part of equation : hardware for compute has always been very good
1. drivers are now excellent
2. whole line of GCN devices (since 2012) are supported and still optimized, unlike notorious Nvidia nerfing of older cards
3. Has x86 license - if Zen is a success, can build high-bandwidth fabric between CPU and GPU, unmatchable by nVidia except on exotic Power8 arch.
4. new project focusing on HPC on Linux
5. they keep pushing OpenCL, which will win over Nvidia-only CUDA
6. New tools to semi-automatically port CUDA code to run on their hardware.
Well, here's hoping AMD gives them a run for their money in 2016 </fanboy> 6.
By now supporting porting CUDA "natively", still no real answers to the datacenter CUDA features like DMA, proper virtualization, and networking.
1. drivers are now excellent
The Crimson driver suite is a step in the right direction, if only every 6 months they didn't had a release that actually physically damages cards.
2. whole line of GCN devices (since 2012) are supported and still optimized, unlike notorious Nvidia nerfing of older cards
Yes that's an admirable feat, but then again not one really cares about 5 year old hardware for either gaming or ML
3. Has x86 license - if Zen is a success, can build high-bandwidth fabric between CPU and GPU, unmatchable by nVidia except on exotic Power8 arch.
NVIDIA also has the x86 license it's more limited and AMD (w/ Intel) did bash them when they tried to add x86 interoperability to their HPC parts but NVIDIA has a very very large IP portfolio.
4. new project focusing on HPC on Linux
AMD has a new project every 2 years, they tend to not die, CUDA on Linux is excellent it's also the recommended platform.
5. they keep pushing OpenCL, which will win over Nvidia-only CUDA
OpenCL performance on NVIDIA GPU's is still better, and OpenCL will beat CUDA has been touted for nearly 10 years now...
6. New tools to semi-automatically port CUDA code to run on their hardware.
With still pretty poor results in many cases, and it only works if you use the most basic use cases under CUDA and this was the bare minimum they had to do so people would even look at AMD hardware at large scales these days since virtually every high performing library is written in CUDA simply because it offered a better solution and NVIDIA actually provides whatcha call it - ah right support...
Well, here's hoping AMD gives them a run for their money in 2016 </fanboy> 6.
http://www.forbes.com/sites/moorinsights/2016/11/21/intel-co...
I dunno where this company is going next or even if it will be around (independently) in 10 years, but if that's really the attitude of its CEO, then at least today it's definitely in good hands.
The same can happen to NVIDIA, especially when Intel decides to launch something inside their CPUs that is "good enough" for many people.
Even the best current GPUs can't handle the inevitable future high-res VR. This is going to be a stable business for a while yet.
Intel hasn't had real pressure to innovate for years. They watched the battle and fall of NVIDIA/AMD, and in the meanwhile enjoyed good sales of CPUs for all the cloud computing.
Now, they have intense pressure to innovate - as in the server space ARM is going to eat them up, due to the superior power efficiency, laptop/desktop sales are down (again, because a four-year old computer works just fine if you upgrade its RAM to accomodate Chrome).
Intel needs something mind-blowing and that quick. And NVIDIA too, casual gamers are shifting to console (where AMD has quite a customer base) and mobile (where Samsung/Apple with ARM are in the lead). I haven't seen any performant Intel or NVIDIA mobile SoC solution yet.
Nvidia does have the Tegra, for the Android tablet market.
Interesting that they've ceded Windows on ARM (aka the rumoured Surface Phone) to Qualcomm - I would have thought Denver's code-morphing architecture would be better suited to emulating x86 instructions than a Snapdragon 835.
Which isn't used anywhere except a couple of tablets according to Wikipedia. Samsung, inarguably leader of the high-class Android manufacturers, uses either its own SoCs or Qualcomm. Many cheaper phones use Mediatek or Allwinner chipsets.
It's not used because most likely NVIDIA never really wanted to push it, it checked the water a bit with a few 1st and 3rd party tablets/android gaming consoles but this always looked more like to recoup some investment on the side while developing the SOC than an end goal.
The lessons learned from Tegra allowed NVIDIA to build the best (or at least the most powerful) automotive integrated solution currently on the market and at least as far as software compatibility and performance go they have no real competition.
The result is called a Tensor Processing Unit (TPU), a custom ASIC we built specifically for machine learning — and tailored for TensorFlow. We’ve been running TPUs inside our data centers for more than a year, and have found them to deliver an order of magnitude better-optimized performance per watt for machine learning. This is roughly equivalent to fast-forwarding technology about seven years into the future (three generations of Moore’s Law). (https://cloudplatform.googleblog.com/2016/05/Google-supercha...)
That sounds like a custom ASIC for this purpose. And if Google can do it, so can Intel.
Nvidia (a) owns the high end deep learning space right now, (b) has shown with the Nintendo Switch arguably the future of gaming, (c) by abandoning PS4 etc they are moving to higher margin businesses.
For me they are looking pretty damn good.
--
This is totally normal for companies to do. In the 1920s Nvidia made car carborateurs. In the 1820s they made horsebuggy whips. In the 1700s they made ships of the line, which they had made going back to the 15th century. It's little known fact that the Niña, the Pinta and Santa María (Columbus's three ships) were powered by NVidia sails. yawn.
nintendo really was a playing-card company foudned in the 19th century though.