Intel CEO: 'The entire industry is motivated to eliminate the CUDA market'
tomshardware.com
tomshardware.com
Google is barely limping along with it's XLA interface to pytorch providing researchers a decent compatibility path. Same with Intel.
Any company in this space should basically setup a giant test suite of IDK, every model on hugging face and just start brute force fixing the issues. Then maybe they can sell some chips!
Intel is basically doing the same shit they always do here, announcing some open initiative and then doing literally the bare minimum to support it. 99% chance openvino goes nowhere. OpenAIs Triton already seems more popular, at least I've heard it referenced a lot more than openvino.
If PyTorch worked fine on Intel GPUs, a lot of people would be happy to switch.
And they don't. My work recently had me working with rocRAND (Rocm's answer to Curand). It was frankly pretty bad- the design, performance (50% slower in places that don't make any sense because generating random numbers is not exact that complicated), and documentation (God it was awful).
Now, that's a small slice of the larger pie. But imagine if this trend continues for other libraries.
But, to be honest, it's not that hard either. I'm surprised their API is 2x slower, Philox is 10 years old now and I don't think there's a licensing fee?
I know! I just wrote a whole paper and published a library on this!
But really, perhaps not as much as many from outside might think. The core of a Philox implementation can be around 50 lines of C++ [1], with all the bells and whistles maybe around 300-400. That implementation's performance equals CuRAND's , sometimes even surpasses it! (the API is designed to avoid maintaining any rng states on device memory, something curand forces you to do).
> running the same PRNG with the same seed on all your cores will produce the same result
You're right. Solution here is to utilize multiple generator objects, one per thread, ensuring each produces statistically independent random streams. Some good algorithms (Philox for example), allow you to use any set of unique values as seeds for your threads (e.g. thread id).
[1] https://github.com/msu-sparta/OpenRAND/blob/main/include/ope...
It's not the generation that matters so much, it's the gathering of entropy, which comes from peripherals and not possible to generate on-die.
If you don't need cryptographically secure randomness, you still want the entropy for generating the seeds per thread/die/chip.
https://github.com/DEShawResearch/random123
if you accept the principles of encryption, then the bits of the output of crypt(key, message) should be totally uncorrelated to the output of crypt(key, message+1). and this requires no state other than knowing the key and the position in the sequence.
the direct-port analogy is that you have an array of CuRand generators, generator index G is equivalent to key G, and you have a fixed start offset for the particular simulation.
moreover, you can then define the key in relation to your actual data. the mental shift from what you're talking about is that in this model, a PRNG isn't something that belongs to the executing thread. every element can get its own PRNG and keystream. And if you use a contextually-meaningful value for the element key, then you already "know" the key from your existing data. And this significantly improves determinism of the simulation etc because PRNG output is tied to the simulation state, not which thread it happens to be scheduled on.
(note that the property of cryptographic non-correlation is NOT guaranteed across keystreams - (key, counter) is NOT guaranteed to be uncorrelated to (key+1, counter), because that's not how encryption usually is used. with a decent crypto, it should still be very good, but, it's not guaranteed to be attack-resistant/etc. so notionally if you use a different key index for every element, element N isn't guaranteed to be uncorrelated to element N+1 at the same place in the keystream. If this is really important then maybe you want to pass your array indexes through a key-spreading function etc.)
there are several benefits to doing it like this. first off obviously you get a keystream for each element of interest. but also there is no real state per-thread either - the key can be determined by looking at the element, but generating a new value doesn't change the key/keystream. so there is nothing to store and update, and you can have arbitrary numbers of generators used at any given time. Also, since this computation is purely mathematical/"pure function", it doesn't really consume any memory-bandwidth to speak of, and since computation time is usually not the limiting element in GPGPU simulations this effectively makes RNG usage "free". my experience is that this increases performance vs CuRand, even while using less VRAM, even just directly porting the "1 execution thread = 1 generator" idiom.
Also, by storing "epoch numbers" (each iteration of the sim, etc), or calculating this based on predictions of PRNG consumption ("each iteration uses at most 16 random numbers"), you can fast-forward or rewind the PRNG to arbitrary times, and you can use this to lookahead or lookback on previous events from the keystream, meaning it serves as a massively potent form of compression as well. Why store data in memory and use up your precious VRAM, when you could simply recompute it on-demand from the original part of the original keystream used to generate it in the first place? (assuming proper "object ownership" of events ofc!) And this actually is pretty much free in performance terms, since it's a "pure function" based on the function parameters, and the GPGPU almost certainly has an excess of computation available.
--
In the extreme case, you should be able to theoretically "walk" huge parts of the keystream and find specific events you need, even if there is no other reference to what happened at that particular time in the past. Like why not just walk through parts of the keystream until you find the event that matches your target criteria? Remember since this is basically pure math, it's generated on-demand by mathing it out, it's pretty much free, and computation is cheap compared to cache/memory or notarizing.
(ie this is a weird form of "inverted-index searching", analogous to Elastic/Solr's transformers and how this allows a large number of individual transformers (which do their own searching/indexing for each query, which will be generally unindexable operations like fulltext etc) to listen to a single IO stream as blocks are broadcast from the disk in big sequential streaming batches. Instead of SSD batch reads you'd be aiming for computation batch reads from a long range within a keystream. (And this is supposition but I think you can also trade back and forth between generator space and index hitrate by pinning certain bits in the output right?)
--
Anyway I don't know how much that maps to your particular use-case but that's the best advice I can give. Procedural generation using a rewindable, element-specific keystream is a very potent form of compression, and very cheap. But, even if all you are doing is just avoiding having to store a bunch of CuRand instances in VRAM... that's still an enormous win even if you directly port your existing application to simply use the globalThreadIdx like it was a CuRand stateful instance being loaded/saved back to VRAM. Like I said, my experience is that because you're changing mutation to computation, this runs faster and also uses less VRAM, it is both smaller and better and probably also statistically better randomness (especially if you choose the "hard" algorithms instead of the "optimized" versions like threefish instead of threefry etc). The bit distribution patterns of cryptographic algorithms is something that a lot of people pay very very close attention to, you are turning a science toy implementation into a gatling gun there simply by modeling your task and the RNG slightly differently.
That is the reason why you shouldn't do the "just download random numbers", as a sibling comment mentions (probably a joke) - that consumes VRAM, or at least system memory (and pcie bandwidth). and you know what's usually way more available as a resource in most GPGPU applications than VRAM or PCIe bandwidth? pure ALU/FPU computation time.
buddy, everyone has random numbers, they come with the fucking xbox. ;)
in practice, assuming a gradient-descent event needs a lot of random numbers, having one keystream for a single GD event might be too much and that's where key-spreading comes in. if you take the "weightIdx W at GradientDescentIdx G" as the key, you can have a whole global keystream-space for that descent stage. And the key-spreading-function lets you go between your composite key and a practical one.
https://en.wikipedia.org/wiki/Key_derivation_function
(again, like threefry, there is notionally no need for this to be cryptographically secure in most cases, as long as it spreads in ways that your CBRNG crypto algorithm can tolerate without bit-correlation. there is no need to do 2 million rounds here either etc. You should actually pick reasonable parameters here for fast performance, but good enough keyspreading for your needs.)
I've been out of this for a long time, I've been told I'm out of date before and GPGPUs might not behave exactly this way anymore, so please just take it in the spirit it's offered, can't guarantee this is right but I've specifically gazed into the abyss the CuRand situation a decade ago and this was what I managed to come up with. I do feel your pain on the stateful RNG situation, managing state per-execution-thread is awful and destroys simulation reproducibility, and managing a PRNG context for each possible element is often infeasible. What a waste of VRAM and bandwidth and mutation/cache etc.
And I think that cryptographic/pseudo-cryptographic PRNG models are frankly just a much better horse to hook your wagon to than scientific/academic ones, even apart from all the other advantages. Like there's just not any way mersenne twister or w/e is better than threefish, sorry academia
--
edit: Real-world sim programs are usually very low-intensity and have effectively unlimited amounts of compute to spare, they just ride on bandwidth (sort/search or sort/prefix-scan/search algorithms with global scope building blocks often work well).
And tbh that's why tensor is so amazing, it's super effective at math intensity and computational focus, and that's what GPUs do well, augmented by things like sparse models etc. Make your random not-math task into dense or sparse (but optimized) GPGPU math, plus you get a solution (reasonable optimum) to an intractible problem in realtime. The experienced salesman usually finds a reasonable optimum, but we pay him in GEMM/BLAS/Tensor compute time instead of dollars.
Sort/search or sort/prefix-sum/search often works really well in deterministic programs too. Do you ever have a "myGroup[groupIdx].addObj(objIdx) stage? that's a sort and prefix-sum operation right there, and both of those ops run super well on GPGPU.
You can't bench implementations of random numbers against each other purely on execution speed.
A better algorithm (better statistical properties) will be slower.
But covering all the important kernels acros all the crazy architecture out there and with relatively good performance and numerical accuracy ... Much harder
I've asked before if they'll merge it back into PyTorch main and include it in the CI, not sure if they've done that yet.
In this case I think the biggest bottleneck is just that they don't have a fast enough card that can compete with having a 3090 or an A100. And Gaudi is stuck on a different software platform which doesn't seem as flexible as an A100.
I tried the a770, but returned it. Half the stuff does not work. They have the CPU side and GPU development on different branches (GPU seems to be ~6 months behind CPU) and often you have to compile it yourself, (if you want torchvision or torchaudio) it also currently on 2.0.1 of pytorch so somewhat lagging, and does not have most of the performance analysis software available. You also, do need to modify your pytorch code, often more than just replacing cuda for xpu as the device. They are also doing all development internally, then pushing intermittently to public. A lot of this would not be as bad if there was a better idea of feature timeline, or if they made their CI public. (Trying to build it myself involved a extremely hacky bash script, that inevitably failed halfway through.)
You have a massive organization full of gazillions of engineers many of whom are really excellent. Before you open your mouth in public and say something is a priority, deploy a lot of them against this and manifest that priority by actually doing the thing that is necessary so people can use your stuff.
It’s really hard to take them seriously when they haven’t (yet) done that.
There was a post on HN a few months ago about how Nvidia's CEO still has meetings with engineers in the trenches. Contrast that with what we know of Intel, which is not much good, and a lot of bad. (That they are notoriously not-well-paying, because they were riding on their name recognition.)
(They don't advertise this as "support" because they have higher standards for what that term means. PyTorch includes a zillion different "operators" and some of them might be unimplemented still. Besides performance is still lacking compared to CUDA, Rocm or Metal on leading hardware - so only useful for toy models.)
TCP/IP completely displaced IPX to the point that most people don't even remember what it was. Nobody uses WINS anymore, even Microsoft uses DNS. It's rare to find an operating system that doesn't implement the POSIX API.
The past is littered with the corpses of proprietary technologies displaced by open standards. Because customers don't actually want vendor-locked technology. They tolerate it when it's the only viable alternative, but make the open option good and the proprietary one will be on its way out.
Except that it has barelly improved beyond CLI and daemons, still thinks terminals are the only hardware, everything else that matters isn't part of it, not even more modern networking protocols that aren't exposed in socket configurations.
While POSIX was state of the art when it was invented, it shouldn't be today.
Lots of research was thrown in recycle bin because "Hey, we have POSIX, why reinvent the wheel?", up to the point that nobody wants to do operating systems research today, because they don't want their hard work to get thrown into the same recycle bin.
I think that people who invented POSIX were innovators and have they live today, they would come with a totally new paradigm, more fit to today's needs and knowledge.
And this strategy regularly works out for companies. It's Commoditize Your Complement.
If you're Intel you support Linux and other open source software so you can sell hardware that competes with vertically integrated vendors like DEC. This has gone very well for Intel -- proprietary RISC server architectures are basically dead, and Linux dominates much of the server market in which case they don't have to share their margins with Microsoft. The main survivor is IBM, which is another company that has embraced open standards. It might have also worked out for Sun but they failed to make competitive hardware, which is not optional.
We see this all over the place. Google's most successful "messaging service" is Gmail, using standard SMTP. It's rare to the point of notability for a modern internet service to use all proprietary networking protocols instead of standard HTTP and TCP and DNS, but many of them are extremely successful.
And some others are barely scraping by, but they exist, which they wouldn't if there wasn't a standard they could use instead of a proprietary system they were locked out of.
Routing is what killed IPX/SPX and that was going to happen anyway. I am a former Netware admin and it was the post-LAN environment that killed IPX as well as pre-IP NetBEUI and all the rest.
But that's absolutely false about the Nvidia moat being only training. Llama.cpp makes it far more practical run inference on a variety of devices. Including ones with or without Nvidia hardware.
The larger issue is that they need to fix the access to their high end GPUs. You can't rent a MI250... or even a MI300x (yet, I'm working on that myself!). But that said, you can't rent an H100 either... there are none available.
They're having to announce it so much because people are rightly sceptical. Talk is cheap, and their software has sucked for years. Have they given concrete proof of their commitment, e.g. they've spent X dollars or hired Y people to work on it (or big names Z and W)?
MI300x and ROCm 6 and their support of projects like Pytorch, are all good steps in the right direction. HuggingFace now supports ROCm.
I have a bridge to sell.
"AMD is not serious, and neither is Intel for that matter. Their software are piles of proprietary garbage fires. They may say they are serious but literally nothing indicates they are."
"Yes, and ROCm also doesn't work on anything non-AMD. In fact it doesn't even work on all recent AMD gpus. T"
If they want to grab a piece of NVIDIA's pie, they do NOT need to build something better than an H100 right away. There are a million consumers who are happy with a 4090 or 4080 or even 3080 and would love for something that's equally capable at half price, and moreover, actually available for purchase, from Amazon/NewEgg/wherever, and without a "call for pricing" button. AMD and Intel are much better at making their chips available for purchase than NVIDIA. But that's not enough.
What they DO need to do to take a piece of NVIDIA's pie is to build "intelcc", "amdcc", and "qualcommcc" that accept the EXACT SAME code that people feed to "nvcc" so that it compiles as-is, with not a single function prototype being different, no questions asked, and works on the target hardware. It needs to just be a drop-in replacement for CUDA.
When that is done, recompiling PyTorch and everything else to use other chips will be trivial.
That's fine but the job of software is to abstract that out. The code should at least compile, even if it is 10x less efficient and if multiple awkward sets of instructions need to be used instead of one.
If Pytorch can be recompiled for AMD overnight with zero effort (only `ln -s amdcc nvcc`, `ln -s /usr/local/cuda /usr/local/amda`) they will gain some footing against Nvidia.
That's simple, just do what Nvidia did. Be better than the competition.
There is more than just a hardware layer to adoption.
CUDA is a platform is an ecosystem is also some sort of attitude. It won’t go away. Companies invested a lot into it.
Really? Because I’m confident Intel knows exactly what it’s about. Have you looked at their contributions to Linux and open source in general? They employ thousands of software developers.
Oh, wait…
Them understanding what was coming doesn't mean they have a magic wand to instantly have a fully competitive product. You can't write a CUDA competitor until you've gotten the framework laid. The fact they invested so heavily in their GPUs makes it pretty obvious they weren't caught off-guard, but catching up takes time. Sometimes you can't just throw more bodies at the problem...
If they didn’t get it’s the software, they wouldn’t catch up.
You didn’t say they were late to the party, you said they don’t understand what needs to happen. My point is they understand exactly what needs to happen they just didn’t have the technology to even start to tackle the problem until recently.
Leverage LLMs to port the SW
Until Intel finds a CEO who is as technical and strategic, as opposed to the bean-counters, I doubt that they will manage to organize a successful counterattack on CUDA.
Did you just call Gelsinger a "non-technical"? wow, how out of touch with reality
>Gelsinger first joined Intel at 18 years old in 1979 just after earning an associate degree from Lincoln Tech.[9] He spent much of his career with the company in Oregon,[12] where he maintains a home.[13] In 1987, he co-authored his first book about programming the 80386 microprocessor.[14][1] Gelsinger was the lead architect of the 4th generation 80486 processor[1] introduced in 1989.[9] At age 32, he was named the youngest vice president in Intel's history.[7] Mentored by Intel CEO Andrew Grove, Gelsinger became the company's CTO in 2001, leading key technology developments, including Wi-Fi, USB, Intel Core and Intel Xeon processors, and 14 chip projects.[2][15] He launched the Intel Developer Forum conference as a counterpart to Microsoft's WinHEC.
I used to work at Nokia Research. The problem was on full display during the period Apple made it's entry into mobile. We had plenty of great software people throughout the company. But the leadership had grown up in a world where Nokia was basically making and selling hardware. Radio engineers and hardware engineers basically. They did not get software all that well. And of course what Apple did was executing really well on software for what was initially a nice but not particularly impressive bit of hardware. It's the software that made the difference. The hardware excellence came later. And the software only got better over time. Nokia never recovered from that. And they tried really hard to fix the software. It failed. They couldn't do it. Symbian was a train wreck and flopped hard in the market.
Intel is facing the same issue here. Their hardware is only useful if there's great software to do something with it. The whole point of hardware is running software. And Intel is not in the software business so they need others to do that for them. Similar to Nokia, Apple came along and showed the world that you don't need Intel hardware to deliver a great software experience. Now their competitor NVidia is basically stealing their thunder in the AI and 3D graphics market. Intel wants in but just like they failed to get into the mobile market (they tried, with Nokia even), their efforts to enter this market are also crippled by their software ineptness.
This is a lesson that many IOT companies struggle with as well. Great hardware but they typically struggle with their software ecosystems and unlocking the value of the hardware. So much so that one Finnish software company in this space (Wirepas), has been running an absolute genius marketing campaign with the beautiful slogan "Most IOT is shit". Check out their website. Some very nice Finnish humor on display there. Their blunt message is that most hardware focused IOT companies are hopelessly clumsy on the software front and they of course provide a solution.
He spent almost decade on VMware which is... software company that significantly grew during his time
>just like they failed to get into the mobile market (they tried, with Nokia even), their efforts to enter this market are also crippled by their software ineptness.
Microsoft, which is software company failed at it too.
Nokia kept doing crazy hardware to show off on hardware side. But these old companies can't stop nickle and dimeing - so you'd get stuff w crazy drm etc. And software wasn't there or invested in fully.
The downfall began with Krzanich who had no goal besides raising the stock price and no strategy other than cutting long-term projects and other costs that got in the way. What a shame.
This is interesting - because what I heard (within Intel at the time, circa 2015) was Xeon Phi was a disaster. The programming model was bad and they couldn't sell them.
It still doesn't change mistake in your original message.
>Gelsinger was appointed out of desperation, but the problem runs much deeper.
How much "much deeper"? VPs? middle level managers? engineers?
The example goes from the top, so if he can change the culture at the top, it will eventually get deeper.
It seems that this is an AND not an OR.
Precisely. The problem I found on HN is that It is hard to have any meaningful discussion on anything Hardware. Especially when it is mixed with business or economics models.
I was greedy and was hoping Intel could fall closer to $20 in early 2023 before I load up more of their stock. Otherwise I would put more money where my mouth is.
You will need to observe Intel more closely. It was not out of desperation. And Gelsinger is more technical and strategic than you implies.
That’s very different from “non-technical.”
He was clearly capable of leading a team to develop a new processor but that’s not the issue here.
I believe his vision/strategy is really sound
I’ll have to take your word for it on his vision/strategy though. I worked for him at VMware and never saw him articulate one.
Even if Intel falls over its own feet, the incentives to bring in more chip manufacturers are huge. It'll happen, the only question is whether the timeframe is months, years or a decade. My guess is shorter timeframes, this seems to mostly be matrix multiplication and there is suddenly a lot of money and attention on the matter. And AMD's APU play [0] is starting to reach the high end of the market with the MI300A which is an interesting development.
[0] EDIT: For anyone not following that story, they've been unifying system and GPU memory; so if I've understood this correctly there isn't any need to "copy data to the GPU" any more on those chips. Basically the CPU will now have big extensions for doing matrix math. Seems likely to catch on. Historically they've been adding that tech to low-end CPU so it isn't useful for AI work, now they're adding it to the big ones.
How many programmers one can employ is determined by profits, and Nvidia has monopoly profits thanks to CUDA, while "the entire industry" can at best hope to create some commiditized alternative to CUDA. Companies with real market power can beat entire industries of commodity manufacturers, Apple is the prime example.
https://www.cnbc.com/2023/12/06/meta-and-microsoft-to-buy-am...
> That is a lot more programmers than Nvidia can afford to employ
How do you account for the increased complexity those developers have to deal with in an environment where there are multiple companies with conflicting incentives working on the standard?
My gut reaction is to worry if this is one of those problems like "9 people working together can't have a baby in one month".
Of course, based on what we see right now that standard would be Nvidia's CUDA; but while CUDA is impressive I don't think running neural nets requires that level of complexity. We're not talking about GUIs which are one of the stickiest and most complicated blocks of software we know about, or complex platform-specific operations. I'd expect that the need for specialist libraries to do inference to go away in time and CUDA to be mainly useful for researching GPU applications to new problems. Training will likely just come down to raw ops/second in hardware rather than software.
It isn't like this stuff can't already run on other cards. AMD cards can run stable diffusion or LLMs. The issue is just that AMD drivers tend to crash. That is simultaneously a huge and a tiny problem - if they focus on it it won't be around for long. CUDA is an advantage, but not a moat.
I mean, this statement is technically true, but it's true for any proprietary technology. If things work like this then we won't have any industry where proprietary techs/formats are prevalent.
I was interested in being part of this AI thing, what stopped me wasn't lack of CUDA, it was that my AMD card reliably crashes under load doing compute workloads. Then when I see George Hotz having a go, the problem isn't lack of CUDA; it was that his AMD card crashed under compute workloads (technically I think it was running the demo suite). That is only anecdata, but 2 for 2 is almost a significant number of people with the small number of players and lack of big money in AI historically.
Lacking CUDA specifically might be a problem here, but I've never seen AMD fall down at that point. I've only ever see them fall down at basic driver bugs. And I don't see how CUDA would matter all that much because I can implement most of what I need math-wise in code. If I see a specific list of common complaints maybe I'll change my mind, but I'm just not detecting where the huge complexity is. I can see CUDA maintaining an edge for years because it is convenient, but I really don't see how it can stay essential. The card can already do the workload in theory and in practice assuming the code path doesn't bug out. I really don't need CUDA, all I want rocBLAS to not crash. I suspect that'd go a long way in practice.
If nothing else, I would be curious to know more about the issue. Personally, I want to know how well ROCm functions on every AMD GPU.
Granted, I may have just gotten off on the wrong foot. The first thing he said to Pivotal during the acquisition announcement was, “You were our cousins, but you’re now more like children.” So the whole tone was just weird.
Hardware without software is just expensive sand. Every semiconductor company knows this. Intel was the one to perfect the whole package with x86 in the first place…
In the GPU compute space CUDA is x86. It’s ubiquitous, de facto standard and will be disrupted. Question is if it takes a year or a decade.
Cuda is enormous, very complicated and fits together relatively well. All the semiconductor startups have a business plan about being transparent drop in replacements for cuda systems, built by some tens of software engineers in a year or two. That only really makes sense if you've totally misjudged the problem difficulty.
He should be sweating.
Hardware development cycles are closer to 5 years. So while he might have gotten some adjustments done on the designs so far, if he turned the ship around it'll take a while longer to materialize.
The software side is more agile, so any tea leave reading to discern what Gelsinger's strategy looks like is best done over there.
It’s only when you put them up against Nvidia and AMDs comes-with-decades-of-experience offerings that Intel’s GPUs seem less than stellar.
They announced the first chips based on Intel 4 today, which is more or less equivalent to TSMC’s 5nm.
They may fail, but the goal is clear and ambitious.
[0] - https://www.xda-developers.com/intel-roadmap-2025-explainer/
And you dont blame that to Raja Koduri but to Gelsinger?
> ...as opposed to the bean-counters
Intel whining about the CHIPS Act, while doing huge layoffs, while continuing massive stock buybacks, while continuing to pay dividends...
I'm not impressed.
I'd expect a corporation facing an existential crisis would make some tough decisions and accept the consequences. I know that Wall St (investors) is the tail wagging the dog. But I expect a leader to hold off the vultures (ghouls) long enough to do the necessary pivots.
At least Intel's doom spiral isn't as grimm as the condundrum facing the automobile manufacturers. They need to prop up their declining business units, pay off their loan sharks (investors), somehow resolve their death embrace with the unions, transform culture (bust up their organizational siloes), transform relationships with their suppliers, AND somehow get capital (cheap enough) to fund their pivots.
I'm sure Intel has the technical chops to do great things. But financially their half-measures strategy doesn't instill confidence.
Source: Just a noob reading the headlines. Am probably totally off base.
"Yes, you know all those r/amd and r/nvidia posts about people ditching their RX 5700 XT and switching over to an Nvidia RTX 2070 Super… AMD is reading them and has obviously been jamming them down the throats of its software engineers until they could release a driver patch which addresses the issues Navi users have been experiencing."
https://www.pcgamesn.com/amd/radeon-rx-5700-drivers-black-sc...
So, its very likely Intel has more software engineers than NVIDIA. Intel has far more products than NVIDIA though, so NVIDIA almost certainly has more software engineers working on GPU.
I think it was 500 people working on Windows XP? A hundred for Windows 95. Etc.
/S
Until then, it's a bit funny claim, especially considering what a failure OpenCL was (programmer's experience and fading support). Or trying to do GPGPU with compute shaders in DX/GL/Vulkan. Are they really "motivated"? Because they had so many years and the results are miserable... And I don't think they invested even a fraction of what got invested into CUDA. Put your money where your mouth is.
But what you mention is important, and also a reason for the ultimate demise of e.g. Xeon Phi. Intel surely realized they could just scale their existing Xeon designs up-and-out further than expected. Like from a product/SKU standpoint, what is the point of having a 300 core Phi where every core is slow as shit, when you have a 100 core 4-socket Xeon design on the horizon, using an existing battle-tested design that you ship billions of dollars worth every year? Especially when the 300 core Xeon fails completely against the competition. By the time Phi died, they were already doing 100-cores-per-socket systems. They essentially realized any market they could have had would be served better by the existing Xeon line and by playing to their existing strengths.
This made me wonder a couple of things-
What kind of workloads and problems is that best suited for? It’s a lot of cores for a CPU, but for pure math/compute, like with AI training and inference and with graphics, 288 cores is like ~1.5% of the number of threads of a modern GPU, right? Doesn’t it take particular kinds of problems to make a 288 core CPU attractive?
I also wondered if the ratio of the highest core count CPU to GPU has been relatively flat for a while? Which way is it trending- which of CPUs or GPUs are getting more cores faster?
And what you suggest could be done, but it would likely flop commercially if you made it today, which is why they aren’t doing it. SIMD machines are faster on homogenous workloads, by a lot. It would be a bummer to develop a CPU with thousands of cores that is still tens or hundreds of times slower than a comparably priced GPU.
SIMD isn’t going away anytime soon, or maybe ever. When the workload is embarrassingly parallel, it’s cheaper and more efficient to use SIMD over general purpose cores. Specialized chiplets and co-processors are on the rise too, co-inciding with the wane of Moore’s law; specialization is often the lowest hanging fruit for improving efficiency now.
There’s going to be plenty of demand for general programmers but maybe worth keeping in mind the kinds of opportunities that are opening up for people who can learn and develop special purpose hardware and software.
If you don't want that, program the GPU directly in assembly or C++ or whatever. A kernel is a thread - program counter, register file, independent execution from the other threads.
There isn't a Linux kernel equivalent sitting between you and the hardware so it's very like bare metal x64 programming, but you could put a kernel abstraction on it if you wanted.
Core isn't very well defined, but if we go with "number of independent program counters live at the same time" it's a few thousand.
X64 cores are vaguely equivalent to GCN compute units, 100 or so if either in a 300W envelope. X64 has two threads and a load of branch prediction / speculation hardware. GCN has 80 threads and swaps between them each cycle. Same sort of idea, different allocation of silicon.
Traditional compute? The cores were too weak.
Number crunching? Okay-ish but gpus were better.
Useless stuff.
I'd love to hear about what didn't work. OpenMP support seemed ok maybe but OpenMP is just a platform, figuring out software architectures that's mechanistically sympathetic to the system is hard. It would be so interesting to see what Xeon Phi might have been if we had Calcite or Velox or OpenXLA or other execution engine/optimizers that can orchestrate usage. The possibility of something like Phi seems so much higher now.
There's such a consensus around Phi tanking, and yes, some people came and tried and failed. But most of those lessons, of why it wasn't working (or was!) never survived the era, never were turned into stories & research that illuminates what Phi really was. My feeling is that most people were staying the course on GPU stuff, and that there weren't that many people trying Phi. I'd like more than the heresay heaped at Phi's feed to judge by.
Math guys came up with a list of algorithms to try for a search engine backend.
What we needed was matrix multiplication and maybe some decision tree walking (that was some time ago, trees were still big back then, NNs were seen as too compute-intensive for no clear benefits). So we thought that it might be cool to have a tool that would support both. Phi sounded just right for both.
And things written to AVX-512 did work. Software surpisingly easy to port.
But then comes the usual SIMD/CPU trouble: every SIMD generation wants a little software rewrite. So for both Phi generations we had to update our code. For things not compatible with the SIMD approach (think tree-walking) it is just a weak x86.
In theory Phi's were universal, in practice what we got was: okay number crunching, bad generic compute.
GPU was somewhat similar: the software stack was unstable, CUDA just did not materialize as a standard yet. But every generation introduced a massive increase in compute available. And boy did NVIDIA move fast...
So GPU situation was: amazing number crunching, no generic compute.
And then there were a few ML breakthroughs results which rendered everything that did not look like a matrix multiplication obsolete.
PS I wouldn't take this story too seriously, details may vary.
- Very bad performance at existing x86 workloads, so a major selling point was basically not there in practice, because extracting any meaningful performance required a software rewrite anyway. This was an important adoption criteria; if they outright said "All your existing workloads are compatible, but will perform like complete dogshit", why would anyone bother? Compatibility was a big selling point that ended up meaning little in practice, unfortunately.
- Not actually what x86 users wanted. This was at the height of "Intel stagnation" and while I think they were experimenting with lots of stuff, well, in this case, they were serving a market that didn't really want what they had (or at least wasn't convinced they wanted it).
- GPU creators weren't sitting idle and twiddling their thumbs. Nvidia was continuously improving performance and programmability of their GPUs across all segments (gaming, HPC, datacenters, scientific workloads) while this was all happening. They improved their compilers, programming models, and microarchitecture. They did not sit by on any of these fronts.
Ironically the main living legacy of Phi is AVX-512, which people did and still do want. But that kind of gives it all away, doesn't it? People didn't want a new massively multicore microarchitecture. They wanted new vector instructions that were flexible and easier to program than what they had -- and AVX-512 is really much better. They wanted the things they were already doing to get better, not things that were like, effectively a different market.
Anyway, the most important point is probably the last one, honestly. Like we could talk a lot about compiler optimizations or autovectorization. But really, the market that Phi was trying to occupy just wasn't actually that big, and in the end, GPUs got better at things they were bad at, quicker than Phi got better at things it was bad at. It's not dissimilar to Optane. Technically interesting, and I mourn its death, but the competition simply improved faster than the adoption rate of the new thing, and so flash is what we have.
Once you factor in that you have to rewrite software to get meaningful performance uplift, the rest sort of falls into place. Keep in mind that if you have a $10,000 chip and you can only extract 50% of the performance, you more or less have just $5,000 on fire for nothing in return. You might as well go all the way and use a GPU because at least then you're getting more ops/mm^2 of silicon.
In practice, people weren’t afraid to roll up their sleeves and write CUDA code. If you wanted good performance you had to think about data parallelism anyways, and at that point you’re not benefiting from x86 backwards compatibility. It was a fascinating dream while it lasted though.
There's no intrinsic reason why multiplying matrices requires massive parallelism, in principle it could be done on few cores plus good management of ALUs/memory bandwidth/caches.
[1]: https://github.com/bevyengine/bevy/blob/main/examples/shader...
> setting up the graphics pipe
I’ve picked D3D11, as opposed to D3D12 or Vulkan. The 11 is significantly higher level, and much easier to use.
> compiling, binding
The compiler is design-time, I ship them compiled, and integrated into the IDE. I solved the bindings with a simple code generation tool, which parses HLSL and generates C#.
> No proper debugging
I partially agree but still, we have renderdoc.
Indeed, in D3D they are called “wave intrinsics” and require D3D12. But that’s IMO a reasonable price to pay for hardware compatibility.
> no cooperative matrix multiplication
Matrix multiplication compute shader which uses group shared memory for cooperative loads: https://github.com/Const-me/Cgml/blob/master/Mistral/Mistral...
> tensor cores
When running inference on end-user computers, for many practical applications users don’t care about throughput. They only have a single audio stream / chat / picture being generated, their batch size is a small number often just 1, and they mostly care about latency, not throughput. Under these conditions inference is guaranteed to bottleneck on memory bandwidth, as opposed to compute. For use cases like that, tensor cores are useless.
> there's no way D3D11 can compete with either CUDA
My D3D11 port of Whisper outperformed original CUDA-based implementation running on the same GPU: https://github.com/Const-me/Whisper/
One of the features that would get compute shaders far ahead compared to now would be pointers and pointer casting - Just let me have a byte buffer and easily cast the bytes to whatever I want. Another would be function pointers. These two are pretty much the main reason I had to stop doing a project in OpenGL/Vulkan, and start using CUDA. There are so many more, however, that make life easier like cooperative groups with device-wide sync, being able to allocate a single buffer with all the GPU memory, recursion, etc.
Khronos should start supporting C++20 for shaders (basically what CUDA is) and stop the glsl or spirv nonsense.
So yet another thing where Khronos APIs are dependent on DirectX evolution.
It used to be that AMD and NVidia would first implement new stuff on DirectX in collaboration with Microsoft, have them as extensions in OpenGL, and eventually as standard features.
Now even the shading language is part of it.
Modern Vulkan is looking pretty good now. Cooperative matrix multiplication has also landed (as a widely supported extension), and I think it's fair to say it's gone past OpenCL.
Whether we get significant adoption of all this I think is too early to say, but I think it's a plausible foundation for real stuff. It's no longer just a toy.
[1] https://community.arm.com/arm-community-blogs/b/graphics-gam...
It's been awesome seeing folks like Keras 3.0 kicking out broad Intercompatibility across JAX, TF, Pytorch, powered by flexible executuon engines. Looking forward to seeing more Vulkan based runs getting socialized benchmarked & compared. https://news.ycombinator.com/item?id=38446353
> [1] https://community.arm.com/arm-community-blogs/b/graphics-gam...
"Using a pointer in a shader - In Vulkan GLSL, there is the GL_EXT_buffer_reference extension "
That extension is utter garbage. I tried it. It was the last thing I tried before giving up on GLSL/Vulkan and switching to CUDA. It was the nail in the coffin that made me go "okay, if that's the best Vulkan can do, then I need to switch to CUDA". It's incredibly cumbersome, confusing and verbose.
What's needed are regular, simple, C-like pointers.
Maybe they should look into their own failures first.
https://www.intel.com/content/www/us/en/developer/tools/onea...
My work involves writing software that runs on many GPU platforms at once. So far we have been going through the Kokkos route, but SYCL is looking pretty good to me recent days. There is some consolidation happening in this space (Codeplay gave up working on their own implementation and merged with Intel). It was pretty easy to setup on my Linux machine for Nvidia card. Documentation is very good and professional, unlike AMD's, which can be frankly horrible at times. And Intel has a good track record with software.
I genuinely believe if someone is going to dethrone CUDA, at this point SYCL (oneAPI) is a far more likely candidate than Rocm/HIP.
This would be the first step. Then, if we want to move away from Cuda into hardware that's as ubiquitous and performant as Nvidia's (or better), someone would need to write an abstraction layer that's more convenient to use than Cuda. I did play a little bit with Cuda and OpenCL, but not enough to hate either.
Any Cuda-compatible software layer only has two options, be a second class CUDA implementation by being compatible with a subset like AMD ROCm and HIP effforts, or be compatible with everything always playing catchup.
The only way is to use middleware that just like in 3D APIs, abstract the actual compute API being used, as man language bindings are doing nowadays.
Implementing programming languages and runtimes is pretty difficult in general. Note that cuda doesn't have the same semantics as c++ despite looking kind of similar. Wherever you differ from expected behaviour people consider it a bug, and implementing based on cuda's docs wouldn't get you the behaviour people expect.
Pretty horrendous task overall. It would be much better for people to stop developing programs that only run on a gnarly proprietary language.
Yet another example on how Intel and AMD failed to take up on CUDA.
SYCL (SYCL-2020 spec) supports multiple backends, including Nvidia's CUDA, AMD's HIP, OpenCL, Intel's Level-zero, and also running on the host CPU. This can either be done with Intel's DPC++ w/ Codeplay's plugins, or using AdaptiveCpp (aka. hipSYCL, aka openSYCL). OpenCL is just another backend.
It is also a very long way from OpenCL C++. The code is a single C++ file, and you don't need to write any special kernel language. The vast majority of SYCL is just C++, so -if you avoid a couple of features- you can use SYCL in library-only form without even any special compiler! This is possible for instance with AdaptiveCpp.
Also some of that work, we have to thank Codeplay for, before their acquistion from Intel.
I'm not an expert here, am I missing something? Saying the x86 industry is motivated to move away from what nvidia provides, intel needs to tick some of these 'better somehow' boxes.
I attribute this to opencl being the common subset a bunch of companies could agree could be implemented. I wrote some code that compiles as cuda, opencl, C++ and openmp, and the entire exercise was repeatedly "what, opencl can't do that either? damn it".
Apple also doesn't care about HPC, and has their own stuff on top of Metal Compute Shaders, just like Microsoft has DirectML.
Unless a random gamer with a random AMD GPU can go to amd.com and download pre-packaged, officially supported tools that work out of the box on their machine and after a few clicks have working GPU-accelerated pytorch (which IMHO isn't the case, but admittedly I haven't tried this year) then their "doubling down" isn't even meeting table stakes.
I predict that access to the newer cards is a more likely scenario. Right now, you can't rent a MI250 or even MI300x, but that is going to change quickly. Azure is going to have them, as well as others (I know this, cause that's what I'm building now).
I'm considering adding ROCm support for some ML-enabled tool - no matter if it's a commercial product or an open source library - the thing I need from AMD is to ensure that the ROCm support I make will work without hassle for these end-users with random old AMD gaming cards (because these are the only users who need the tool to have ROCm support), and if ROCm upstream explicitly drops support for some cards because AMD no longer regularly test it, well, the ML tool developers aren't going to do that testing for them either; that's simply AMD intentionally refusing to do even the bare minimum (do a lot of testing for a wide variety of hardware to fix compatibility issues) that I'd expect to be table stakes for saying that "AMD is doubling down on ROCm".
What they really need is to support the less expensive cards, of which the older cards are a large subset. There are a lot of people who will make contributions and fix bugs if they can actually use the thing. Some CS student at the university has to pay tuition and therefore only has an old RX570, and that isn't going to change in the next couple years, but that kind of student could fix some of the software bugs currently preventing the company from selling more expensive GPUs to large institutions. If the stack supported their hardware.
I get your point about AMD not wanting to spend money on supporting old hardware, but how do they expect to build a market without a fan base?
https://www.phoronix.com/news/Radeon-RX-7900-XT-ROCm-PyTorch
It's clear to everybody that it's not the hardware but the software - which is the CUDA ecosystem.
I've played a bit in the past with ML, but at the level of understanding I had - training some models, tweaking things, I was using higher level libraries and as far as I know, it's pretty much an if statement in those libraries to decide which backend use.
So let's suppose Intel and others does manage to implement a viable competitor - am I wrong in thinking that the transitions for many users would be seamless ? That's probably not the case for researchers and people pushing the boundaries, but for most companies, my understanding is there would be not a lot of migration costs involved ?
The actual PC clones, pure drop-in compatibles, were made in Taiwan and they took over the market. Which is to say that large companies don't want a commodified market where prices are low and everyone competes on a level playing field - which is what "seamless transition" gets you. So that's why none of these companies are working to create that.
device = "cuda" if torch.cuda.is_available() else "cpu"
and I've yet to see a single one implement the AMD NN middleware: https://www.amd.com/en/developer/zendnn.html $ python3 -m venv .
$ ./bin/pip3 install torch --index-url https://download.pytorch.org/whl/rocm5.6
$ ./bin/python3 -c 'import torch; print([torch.cuda.is_available(), torch.cuda.get_device_name(0)])'
[True, 'AMD Radeon RX 6600 XT']The only reasonable way to break the cycle would be proactive intervention by AMD to contribute code and testing and easy "works out of the box" installation to all the major popular projects to add AMD support so they can sell more their hardware later, but I'm not seeing AMD doing that.
I wish AMD&Intel would extend compute shaders with pointers&pointer casting, arbitrary large buffers instead of just 4GB, device-wide sync, and function pointers. Those are kinda my must-have functionality. Even better, just use C++ for compute shaders.
At least Nvidia built something to facilitate advances in AI. They made a bold bet that paid off.
If you see the presentations by Nvidia's Bill Daly [1] it shows they've been evolving hardware and software together. The CUDA programming model was a massive bet a long time ago. And they deservedly won ML/AI.
Unless Intel and AMD do a radical change it's game over for them. They will lose to ARM and Nvidia.
Nvidia is really good at building/buying moats:
PhysX where lack of Nvidia GPU used to default you into unoptimized FPU slow path on SSE capable CPUs https://arstechnica.com/gaming/2010/07/did-nvidia-cripple-it... https://www.realworldtech.com/physx87/3/ "For Nvidia, decreasing the baseline CPU performance by using x87 instructions and a single thread makes GPUs look better."
"The Way It’s Meant To Be Played" program paying of studios to directly cripple AMD. Ubisoft retracting DX10.1 patch is one example https://techreport.com/news/14707/ubisoft-comments-on-assass...
"GameWorks" program where Nvidia went one step further paying game studios to directly embed Nvidia crippleware libraries in games. https://techreport.com/review/21404/crysis-2-tessellation-to... https://arstechnica.com/gaming/2015/05/amd-says-nvidias-game... https://wccftech.com/fight-nvidias-gameworks-continues-amd-c...
However it's worth noting that in the end, it's the consumer that loses out when these technical/capability moats maintain a de facto sort of monopoly on certain use cases.
I think in most cases you are right, but for CUDA, the consumer is winning. CUDA isn't some special secret complex algorithm just for nvidia GPUs. nvidia just spent a decade giving a fuck about the developer experience across a wide range of industries. CUDA isn't just AI, it's also physics, numerical modeling, cryptography, biology and more. They found a thousand use cases and spent time listening to customers and built a platform around it - it just turned out AI ended up being a huge money maker.
The problem is, and will continue to be, that Intel and AMD will only see the "AI" money bags and ignore every other part of platform, from debugging, compilers, language integration, GUIs, and even bug fixing. If Intel is saying here "we want to eliminate CUDA by investing billions into OpenCL and ensure OpenCL has a top of line developer experience and platform" than I'd be excited. But what I'm reading is "we will replace a few function calls for CUDA in Pytorch", which might be fun for a while up until you have to debug some performance issue and you realize you can get in touch with a CUDA engineer on github instead of emailing some dead Intel mailing list.
Intel pushed 14nm to 14nm+++ and used backhand tactics to push OEMs not to use AMD. how long did Intel stay on some trend until it fizzled for Intel?
while Nvidia has been working on CUDA forever?
NVIDIA seems to be able to do software markedly better than all the other hardware folks. And software is the real product that end users actually interact with.
In many ways it feels like the same situation we had in the 70’s and 80’s personal computers, with many different software bases divided by different hardware implementations. In that regard, Unix was a boon in that it allowed one to write software that’d be compiled out of a single source into multiple architectures (rather than the PC-clone way, where non-identical implementations were driven to extinction).
(Do we know each other?)
I dabbled a bit with it in college for a course
replacing it ain't gonna be easy (or even a desired thing for some)
this is where it is going all south imo. some initiatives push their "open-source" nature, but they are still distributed efforts in the end.
the vast majority of developers are using a framework and are completely abstracted out of the inner workings. they are not the main audience when marketting your ecosystem. but those are the ones calling the procurement shots unfortunately.
at this rate, some futuristic "Xeon Phi" like architecture, where the new surge of multithreaded processing on CPU side can just be generalized, might be more likely than "we got CUDA at home" nonsense.
The problem is most people are using NVIDIA libraries which ship with proprietary binary blobs, so good luck running that on your AMD GPU. Until they make all these extra libraries as well, they're SOL.. NVIDIA has zero reason to port all their IP to AMD or Intel.
One would imagine that they already have internal ports of their software to both these platforms to do benchmarking. Again, the technical problem is simply one of software.
Going back to the topic at hand. CUDA language and runtime are easy. Both amd and Intel have competitors here. They don't have the libs. The libraries contain thousands of variants of hand optimized code. Good luck.
Nobody really seems to be taking CUDA compatibility seriously.
No? Then is just talk.
Intel should have an advantage here as MKL and OneAPI are quite good, but the AMD’s roc libraries are not very good.
He has wanted GPGPU or at least massively parallel compute for a long time and says it would be 10 years of work to make that happen.
And secondly, do you really think he's wrong about this part: "As inferencing occurs, hey, once you've trained the model… There is no CUDA dependency," Gelsinger continued. "It's all about, can you run that model well?"
Why run inferencing on an expensive H100? Gaudi seems like a pretty good idea.
Any new business problem in a Microsoft shop is going to go with a Power Platform/Teams/SharePoint Online solution as first prize.
Personally, I like it, even if it has some annoying challenges. I've not seen anything else come close in terms of rapid application development, integration and deployment; there's very little to license and what there is left falls under existing procurement relationships for most big enterprises that rely on Microsoft already.