Nvidia reveals new A.I. chip, says costs of running LLMs will drop significantly
cnbc.com
cnbc.com
I haven't really gotten too much into it, so I'm not sure how has Nvidia come to absolutely dominate this market; AMD GPUs are quite good in the gaming sector... though I guess... with just two real players in the GPU market it's difficult to really get anywhere.
Meanwhile Intel and AMD, never delivered something at the same level as CUDA for OpenCL, and when they finally decided to react with SPIR and C++, everyone was already too busy to care, and it isn't as if the tooling has improved that much.
Hence why OpenCL 3.0 is basically OpenCL 1.0 rebranded.
Can't say they don't deserve their success in a capitalist sense. Can't say their success doesn't alarm me.
The rest of the market has not been reacting to them for a long time.
They had the foresight to see the coming wave and ensure they have both best hardware and best software stack.
What puzzles me more is why AMD isn't throwing money at their software side to sort that out. Rocm was launched in 2016. 7 years later running pytorch on their flagship GPU still requires a janky laundry list of crowdsourced instructions:
https://github.com/AUTOMATIC1111/stable-diffusion-webui/disc...
https://github.com/RechieKho/IREE.gd -- RechieKho and I collaborate on making this work for Godot Engine, but IREE.gd is at a proof of concept stage.
lood_in_4bit=True will let you run Llama2-7B variants at 6.3GB VRAM.