All the companies that are now screaming hellfire because of Nvidia's market maker position are also the companies that gave Nvidia a warchest filled with billions of dollars. How is AMD supposed to compete when the whole market is funding their rival?
It has only been the last 4 or 5 years that AMD has had any real money to put into their GPU/AI accelerator sector and that seems to be developing quite well now (though they seem to be mostly interested in super computer/massive data center deployments for now)
it's turtles (broken AMD technical leadership) all the way down.
this is simply not a field AMD cared about, whether you think they were right or wrong (based on their financials or otherwise). and now that it has turned into a cash fountain, everyone wishes it had gone differently. "should have bought bitcoin" but instead of pollution funbux it's investing into your own product.
I agree that correctness and robustness of the vendor impolementations went poorly, but it's fixable if there was commitment between vendors, the infighting leading to Apples departure must not have helped that side either.
In the end I think including the high level compilers and APIs in the proprietary driver stack responsibility is unsustainable. Microsoft seems to have a better model, or there could be even a standard JIT code consumed by drivers that emitted by open source stacks shared across platforms. The latter way would also better support development of nicer GPU languages.
OK, let's talk MS.
What NVIDIA has done with PTX is basically the same thing as MSIL/CIL. There is a meta-ISA/meta-language that has been commonly agreed, and if you invoke no tokens that the receiving driver doesn't understand, the legacy compiler understands and emits executable (perhaps not optimal) code.
The legacy hardware stays on a specific driver and a specific CUDA support level. The CUDA support level is everything, that is your entire "feature profile". It's not GFX1030/GFX1031, it's RDNA3.0/3.1/whatever. NVIDIA has been able to maintain their feature support (not implementation details) in a monotonically increasing sequence.
Additionally, they also maintain a guarantee that spec 3.1 guarantees that you can also compile and execute 3.0 code. "Upwards family correctness" I guess. Again, doesn't have to be optimal but the contract is there's no tokens that you don't understand. A Tegra 7.0 compiler can compile desktop 7.0 CUDA (byte/)code etc.
https://en.wikipedia.org/wiki/CUDA#Version_features_and_spec...
Legacy driver support doesn't change, so, you can always read PTX, you can always read anything that's compiled for your CUDA Capability Level even if it's from a future toolkit that wasn't released at the time. PTX is always the ultimate "emit code from 2020 and run on the 2007 gpu that only understands CUDA 1.3" relief valve though. If you don't write code that does stuff that's illegal in 1.3... it'll compile and emit PTX, and the Tesla driver will parse and run it.
Anyway, let's talk MS. MS could come up with MSIL/CIL. NVIDIA can come up with PTX. Why can't AMD do this? What is unique about GCN/RDNA as a language that you can't come up with an intermediate language representation?
But the actual big opportunity which could make the whole GPU programming field far more dev friendly and could make the gpu vendor competition work in favour of users would be doing the GPU-IL in a cross-vendor way, driven by someone else than a single GPU vendor who is playing in the proprietary lock-in game (or certain to do it after gaining sufficient market position).
oh, right, those are EOL'd, security updates only/separate driver package with a quarterly release cycle, even though they're still on store shelves...
Remember that the "AI core" in the new meteor lake and hawk point stuff is not really all that gangbusters either... it's a low-power sidekick, for the same types of stuff that smartphones have been doing with their AI cores for ages. Enhancing the cameraphone (these cameras would be complete shit without computational/AI enhancement). Recognizing gestures, recognizing keywords for voice assistant activation.
AMD's pitch for the AI core is enhancing game AI. Windows 11/12 assistant. That type of stuff.
Vega can absolutely just brute-force its way through that stuff, and it gets more people onto the platform and developing for AMD. It is crazy that if nothing else they aren't at least making sure the APUs can run shit.
And again, it's pretty damn unethical imo to be pulling support from products that are still actively marketed and sold. That's a cheezy move. AMD dropped support for Terascale before NVIDIA dropped support for fermi, and NVIDIA went back and added Vulkan support to all their older stuff during the pandemic too. Then they dropped the 28nm GCN families, while NVIDIA is still supporting maxwell. And then they dropped Polaris and Vega, and NVIDIA is still supporting Maxwell (albeit I expect them to drop it very soon imo).
The open driver is great under linux because it bypasses AMD's craptacular support. But AMD doesn't support consumer GPUs in ROCm under Linux, and de-facto they don't seem to support ROCm on the open driver in the first place anyway, you have to use AMDGPU-PRO (according to geohot's investigations).
This is such a crazy miss. Yes, it's not a powerhouse, but in the era of Win12 moving to AI everything, and games moving to AI-driven computer opponents, etc - Vega can do that ok, and it at least would give people something to open the door and get them developing.
If there's a contender for breakout against CUDA, sadly it really seems to be Apple. They've got APUs with PS5-level bandwidth and unified memory, and that's just M1 Max, and they have Metal, and it's supported everywhere across their ecosystem. It lets people dip their toes into AI/ML if nothing else, and lets people dip their toes into metal development. That's the kind of environment that NVIDIA spent a decade fostering, and it's also not a coincidence that the second stop for all these data scientists playing with models is not ROCm but their apple laptops. llama.cpp and so on. Everyone likes the hardware they have in their gaming PC or in their laptop, and it's an absolute miss for AMD to not make themselves available in that fashion when they already have the market penetration. Crazy.
https://twitter.com/Locuza_/status/1450271726827413508/photo...
While they don't officially support any consumer GPU aside from 7900XT(X) and VII, I haven't encountered any issues using it on a 6700XT with the open source drivers, pretty much the only tinkering required was to export HSA_OVERRIDE_GFX_VERSION=10.3.0. It was quite a pleasant surprise after never getting my RX480 to work even though it was officially supported back in the day.