Jensen Huang’s vision for data center dominance may destroy the Arm ecosystem
semianalysis.com
semianalysis.com
Sounds like the bullshit that IBM salespeople were peddling about POWER circa 2008. The real strategy was licensing capacity with physical segmentation. Nvidia would probably pursue a similar strategy -- sell chips with lots and lots of cores and lease capacity on demand.
You could pave a road with the gravestones of technology vendors who pursued industry dominance with faster, more expensive chips. If IBM, Sun, Unisys, HP, Digital, SGI and a dozen others failed, why would a company with almost zero datacenter presence succeed?
Even with the pitiful state of 2020 Intel, I'd put my money on commodity ethernet and x86 in the datacenter. Most datacenter architecture doesn't require GPU. Nvidia could market with a kick-ass Oracle-like engineered solution for high-performance compute, game hosting or AI. But I don't think Oracle moved the needle in how we buy database compute, and I doubt Nvidia would do so here.
I have very much the same opinion, and I think that Huang is at least aware of that.
> Nvidia would probably pursue a similar strategy -- sell chips with lots and lots of cores and lease capacity on demand.
But this will be pretty much the same IBM thing. Some banks, and Oracle/SAP buyer types may buy in, but that will be it. It does not change anything to how the companies in the above list of "big iron" vendors attempted to do it before.
I want to remind that Huang is also, a very big "investment relations" player, and he likes to make an impression of "grandiose" plans to impressionable investor guys.
Impressionable investor guys tend not to be investor guys for very long. Big-vision bluster can can work with media and some end users but it turns off real market analysts.
I personally feel like distributed (coherent) shared memory is a bit overkill. I would much rather the kind of grain that a solid 10gbps link works well with. I.e. anything that is sensitive to latencies measured in microseconds or lower just lives on the same node where it needs to be sampled.
I can't think of too many applications that couldn't be designed to work just as well using 10GbE vs Infiniband/RDMA/shared memory. Even most real-time simulations could be designed to run in a practically syncrhonous fashion across a number of distributed hosts using stuff you can buy at Newegg or Amazon.
There is a lot of research about doing databases on GPUs, and Apache Spark runs on GPUs today, with much higher performance than on CPUs.
But even more than that, for the price of a single high end Tesla (approx. 10k USD), you can build a high-end COTS x86 machine with a shitload of RAM, NVMe, and then install ClickHouse on it. That machine will scale to trillions of rows with ease and millisecond response times, whether or not everything fits in memory. It will cost less money and also cost less energy and it will scale out easier, and have better utilization of the hardware.
I'd wager that unless you have infinite money to dump on Nvidia or exceedingly specific requirements, any GPU database will get soaked by a comparable columnar OLAP store in every dimension.
https://en.wikipedia.org/wiki/Netezza#Technology
Furthermore, that line of reasoning is fallacious. The non-existence of a thing doesn't prove that the idea is bad.
I also ran datacenter services on Sun and IBM hardware that smoked anything offered by Intel platforms. But at the end of the day, commodity beats premium for 80%+ of the market. We run Linux on POWER exclusively as a cost-savings mechanism to reduce the Oracle bill.
Intel/AMD is allowed to build a moat in data center but Nvidia can't(?)
>Nvidia’s endgame isn’t more revenue from licensing costs. Their endgame is a fully vertically integrated data center provider. They will want to make and control every part of the three legged stool. This means they slowly destroy the idea of Neoverse. Whether through making that IP extremely costly, or having their own in house designs be a generation ahead, Nvidia will build a moat around Arm server CPUs. Over time, Jensen Huang will muscle out other Arm vendors
Intel doesn't even licensed out x86. AMD is grandfathered in. So Intel is allowed to control x86 but Nvidia can't. this is like Oracle bought MySQL along with Sun and Michael Widenius cry about it. You know what you are doing when you sold MySql to Sun. another company can come along to buy it. this is what closed ecosystem/software does. MySql is saved by dual licenses so maybe its time to ditch this idea of building an ecosystem on closed hardware/software.
So why would it pay $30bn+ to own Arm? Because it would have the ability to directly hinder the businesses (and customers) of its current and potential competitors who have invested in the development of Arm based products. Not just in the data center but in mobile and in products that are used by billions of people around the world. That can't be right.
They have extensive contracts in place. It's not like ARM instruction set is provided to nVidia and then they can just revoke it because they don't like nVidia.
If so Nvidia will surely buy ARM.
I think that's the big irony here. It's taken ARM decades to break out of embedded/mobile and get to the point where they're seriously considered for workstation or server use. If Nvidia were to acquire ARM it would likely antagonize everyone to the point where they'd rush back to x86 and Nvidia would gain nothing.
It's enough for a short term gains, which is enough for a nifty bonus to a CEO.
I think they would be expending the money to grow their CPU team. The ARM engineers would be working on making Nvidia CPUs better, neglecting “public” ARM cores.
This would hurt small fish, but Apple, Amazon, Google, Tesla, Facebook, etc already have their own in house ARM teams.
The thing is, there are 3 companies that control the future of datacenter growth. Amazon, Microsoft and Google. They can move the industry to RISC-V if NVidia doesn't play ball, or convince Intel to manufacture on TSMC's process.
Nvidia won't be in a particularly strong negotiating position. My understanding is that lots of ARM licenses are perpetual, or very broad. The existing players can just keep making ARM chips, what can NVidia do to extract more money from them, new instructions?
I don't know what those perpetual licenses look like, but, if there are any strings attached at all, perhaps Nvidia can tug on them in order to apply a little good old-fashioned embrace-extend-extinguish to ARM? Extend ARM with new instructions that offer performance improvements when you've bought 100% into Nvidia hardware, maybe, and then stifle competition by making them implement these features in order to claim ARM compatibility, thereby making it more difficult/expensive for anyone else to pair ARM CPUs with other companies' GPUs?
Similarly, AWS is in competition with nobody in this context. AWS just influences the competitive landscape and chooses who wins by awarding contracts.
IMHO, it simply looked like he was was quietly moving out cash from the company.
The are big players in Internet business, but they are not big enough consumers to benefit economically from owning chips. Even if they will be all combined, they will not be even a double digit of top tier Xeon buyers.
And this is partly the reason why AMD was so keen on playing ball with them. AMD among other things was making a proprietary CPU for Amazon, which the later discarded. It was later rebranded Seattle, and thrown on the open market.
AWS Graviton 2 (Arm) would like a word about that with you. https://aws.amazon.com/ec2/graviton/
Pretty much the only way they could've done that economically, in relative terms, was at the time of obscenely high prices on high core count xeons.
Once the ball is rolling, they can get 3rd party ARM chips or full servers in and stop with their own development on that side if they want to.
Also, the x86_64 patents don't last forever.
No, but Intel and AMD have a decent reason to work together to create new key feature additions to x86 that can be newly patented. FMA is under patent for 6 more years. AVX2 is 12 more years. AVX512 another 15.
If they were buying directly, Intel would have even more leverage to milk them, than if they stood behind the back of some beefy OEM.
No, they build their own servers. They definitely aren't racking Dell or HP servers.
Even Amazon's x86 based servers are extremely custom. That whole "Nitro" system isn't exactly a USB stick, and AWS is often using customized variants of Intel processors in their servers.
Let me repeat that: Intel literally makes custom processors specifically for AWS.
> If they were buying directly, Intel would have even more leverage to milk them, than if they stood behind the back of some beefy OEM.
That's not how bulk buying works... at all.
The scale of AWS is so much larger than you seem to think it is.
Sort of. It's the same mask for everyone with the per customer special features either fused off, or just requiring a special MSR knock sequence.
AWS operates at a large enough scale that they do exactly that, and AWS already purchases directly from Intel, both of which are my main points. Baybal2 doesn’t seem to be aware of either of these facts based on their comment.
An option to fuse an individual performance profile, or few more ME features doesn't amount to be something really custom.
Why can't you just ask questions instead of misleading people?
With my knowledge of their sales tactic, I would imagine Intel could've simply said "You are free to buy AMD," with full knowledge that just few years ago their clients had no other option than coming back to them eventually.
Yes, they buy from Quanta, Supermicro, Wistron etc, the ones who make servers for Dell, and HP.
> Let me repeat that: Intel literally makes custom processors specifically for AWS.
Some unlockable features, and perf profiles. Those weird SKUs popup on liquidation sale websites from time to time.
Unless this is some kind of performance art, you really should stop commenting and switch to reading.
Same difference, compared to building their own from parts. I don't think parent's point lies in the OEM vs ODM semantics detail.
Both of these events would be very positive for the industry.
In addition to selling hardware, it seems like ARM-owning Nvidia would probably be happy to sell whatever amazing new ARM cores they cook up to whoever is willing to pay to use the design. It's not obvious RISC-V will benefit from any of this. Somebody, somewhere would need to create equally amazing RISC-V cores and find a way to sell them.
My impression is that the Cortex-Ms are popular because of the price, the rich sets of peripherals they tend to come with, and not least because they come with a proven, performant C compiler. I don't think it's impossible that a RISC-V core could make some dents in this markey fairly quickly, especially if Nvidia make the future of the platform look uncertain.
I'm still not sure, after reading all these comments, why anyone thinks an Nvidia that owned ARM would do anything other than continue to sell ARM cores and ARM IP rights to other companies, as well as continuing to sell actual hardware products using ARM cores. Suspecting the worst in everything is sort of standard nowadays, but buying ARM Ltd. so Nvidia can stop promoting the platform and selling the IP does not make much sense.
(unless they literally have a master plan to drive everyone back to AMD64, which doesn't make any business sense since they can't make that hardware, or drive everyone to RISC-V and profit off of that, which doesn't make any business sense since they could attempt that without buying ARM Ltd. and spending tens of billions of dollars)
First get C, C++, Go, Rust, Java, Erlang, .NET,.... generating competitive native code, port all their main libraries and IDE tooling, and then it might stand a chance.
You're right that a number of mainstream software packages are still work in progress, for example V8 doesn't target it yet.
But RISC-V is thriving in other areas like industrial controllers, and Alibaba Cloud just announced a few days ago they are using it in a new server processor, positioned as an eventual competitor to AWS Graviton.
Topically: Nvidia uses RISC-V in some GPUs as a controller (not the GPU core). Not sure if those are released yet, but they gave a presentation on it at a recent RISC-V meet-up.
Anybody know what SCYL is? Based on the low number of google hits I'm thinking it's a typo
Not niche traction and announcements here and there, actual heavy investments...
Or is it more like an "open arch" fantasy thing, kind of the cpu analogous to Ogg Vorbis and Open Moko?
Also I don't know why you count Ogg Vorbis as a fantasy kind of thing, it's used for pretty much any conference app and it has support on most consumer hardware and softwware for playback. Programs like EAC and dbpoweramp support it. I use it for my music collection on my phone, at half the size of mp3 it sounds completely transparent to me with my earbuds.
The idea in the day was it would replace mp3/aac/etc which it obviously didn't. That players support it is also not a mark of great success, since most people don't and won't care to use it. It just means that for the rare person caring enough, it will work.
From that aspect, the fact that conference apps use it as the internal format is hardly relevant, they could license and use whatever and we wouldn't know any difference.
It did for anyone that cares about it and about file size. My portable library is vorbis. Mp3 is just very popular and there are a lot of mp3 files out there that can't be converted without further degradation of the sound quality so mp3 will remain relevant for a long time.
>From that aspect, the fact that conference apps use it as the internal format is hardly relevant, they could license and use whatever and we wouldn't know any difference.
It's relevant in terms of market support. Vorbis isn't going anywhere. The random codecs that have popped over the years, open or proprietary never got a lot of support from other companies so they went away.
And there are only some Linux special variants for it.
And in teaching. If you're going to get a sophomore to build their first 3 stage RISC pipeline or a senior their first OoO machine with register renaming then RISC-V's consistent design removes a lot of unnecessary distractions.
I assume people are using it in research too. If I were to redo my thesis on improving adder efficiency today I'd probably look at using an open source RISC-V design as a base.
I'm skeptical of RISC-V's ability to succeed as an application core like ARM's A series or modern x86. But none the less I'm bullish on it having an important future.
But RISC-V is thriving in other areas like industrial controllers, and Alibaba Cloud just announced a few days ago they are using it in a new server processor, positioned as an eventual competitor to AWS Graviton.
Topically: Nvidia uses RISC-V in some GPUs as a controller (not the GPU core). Not sure if those are released yet, but they gave a presentation on it at a recent RISC-V meet-up.
What if Intel asked NVIDIA to buy ARM?
(https://en.wikipedia.org/wiki/Betteridge%27s_law_of_headline...)
Companies like Lightedge will take on the HVAC and physical security. The future is chiplets and computational memory: https://www.researchgate.net/profile/Michael_Stumm/publicati...
But ISA still matters. Switching from Intel to AMD likely isn't much of a switch, but switching to a different ISA in the datacenter will continue to be an uphill battle.
As the article notes, the uphill battle will take longer for RISC-V and if Nvidia follows the strategy laid out they hope to profit in the intervening gap.
Running on local is essential - things are so much easier to debug, especially things deep in the dependency tree.
It really depends on what kind of stuff the datacenter runs.
I know of datacenters where almost all the code running is Java. For those kind of datacenters, so long as JDK supports ARM (which it does), I can't see why it would be such an uphill battle.
The uphill battle is really for sites who run lots of closed source COTS software, especially that which is written in C/C++/etc, where moving to ARM needs support from the vendors and the vendors might hesitate due to the amount of work involved. There are sites where close to everything is either open-source or else developed in-house in managed languages (Java, .Net, Python, Ruby, PHP, JavaScript, etc), and those sites are likely to find it a lot easier.
I'm wondering how this could happen.
I'm guessing the script is using uname to detect the platform, and gets confused by Linux on ARM.
I used to see a lot of shell scripts which detected Linux vs Solaris vs AIX vs HP-UX and did different things on each, especially due to differences in what commands and options are available. Given those commercial Unices are now shadows of their former selves, you don't see that so much any more.
But I still see it in scripts that have to run on both macOS and Linux. I've even written a few of those scripts recently.
This shouldn't be an issue, though, if you just do `uname -s` – you should get e.g. `Linux` on both ARM and x86. Maybe some people, for whatever reason, are checking `uname -p` or `uname -m` instead or as well, or even trying to parse the output of `uname -a` – not a very good practice
I'd expect datacenter to be far easier than in a laptop or such; datacenters run a lot more software that's either open source (and already working on multiple architectures, most of the time) or in-house (in which case the happy path is "add an extra build job that builds for the new systems"). It's not like consumer space where most users are tied to proprietary software that they couldn't port if they wanted.
The team size I see on the photo don't give off much confidence about them pulling that out.
Saving a few pennies is fine, but it hardly seems like a game changer.