Testing Intel’s Arc A770 GPU for Deep Learning
christianjmills.com
christianjmills.com
The family started with nothing, multiplied and played the strategic long game to conquer the AI marked to enable Skynet.
But Intel are the Jedi hacker good guys.
I've been waiting on AMD to do that. Embarrassingly, it looks like Intel will get there faster than AMD.
The software/driver woes in their graphics division have been going on for years at this point with no clear path forward, just some statements about how they plan to fix it in the future (e.g. handling the geohot AMD rant back in June).
It's not just ML either - on the gaming side, some people swallow the Nvidia price premium for the better drivers.
I made irrational purchasing decisions in the past to support the underdog, likening it to using Firefox. Unfortunately with AMD you just get a worse product.
This space sorely needs more competition.
We need FAANG to put aside their differences to fund an alternative, or else FAANG will suffer the consequences.
They seem to employ zero regression testing.
Nah, I pretty much gave up on AMD when the 7000 series launch was imminent and still no ML real support for the 6000 series. Not that I was even expecting anything. With the way Intel Arc is coming along, any ML support for Radeon is just a bonus. I refuse to pay the Nvidia tax until it becomes clear there is no other option again. Nvidia just really rubbed me the wrong way with the latest generation.
Interesting to put AMD and Nvidia at the same level. From experience developing GPGPU applications they are not even close..
It'll be really hard to compete for AMD Radeon cards and nVidia itself against Intel if they aggresively price dump their GPUs and use typical megacorporation tactics to push out and destroy competition.
Why would you want that? Why not root for an independent GPU company? Where does this love for megaconglomerates consolidating whole markets come from?
AMD's market cap has been higher than Intel's for the last year, and NVIDIA is more than seven times higher. Apple is twice again as big. Surely this logic goes the other way. It's Intel and AMD and ARM Ltd. that are at the mercy of the big players dumping their way to dominance.
But obviously that's not happening. The competitive landscape in this industry is just fine. More products means more innovation and lower prices and bigger markets. That's just capitalism 101, really.
> Where does this love for megaconglomerates consolidating whole markets come from?
Nowhere, because that's just a strawman. I'm sure people here would love to see a GPU startup launch a product. But we're talking about Arc in this thread, and that seems like a good product too.
To me is a mystery that CPUs/motherboards still have extensible RAM, integrated RAM (if you don't skimp and go for cheap RAM) can be very fast - for example the Steam Deck has quad-channel 5500 MT/s LPDDR5, very console-like.
You're forced to make everything slower to make things consistent.
That being said, a more optimized connector, board layout and form factor could make large-capacity, lower-latency modules possible. Not sure for example if there is anything out there besides the Dell proposed CAMM form factor, and I don't remember that one being specifically promoted as "faster", mostly "smaller".
That performance is shockingly good, am I missing something? Comparable to a Titan RTX? Yes, that card is 4 years older but it cost an order of magnitude more, has almost twice the die area, and has twice the GDDR6 bus width. The A770 theoretical FP16 is apparently a bit higher (39.2T vs 32.6T) but in gaming workloads the Arc cards far under-perform their theoretical performance compared to competitors. I guess ML is less demanding on the drivers, or perhaps some micro-architectural thing. In gaming performance the Titan RTX is about 70% faster than the A770.
I'm really looking forward to the next generation now.
For Linux gamers, DXVK is how you'd be playing the game regardless. It's a first-gen card that's not a terrible prospect for the right sort of tinkerer.
I just checked ebay sold listings and Titan RTX cards are selling consistently for between $800 - $1600.
A 3090.
Ampere will be well supported until its very obsolete because of the A100.
These already exist, they're just disproportionately expensive compared to consumer cards. You can get a Nvidia A100 with 80GB VRAM on eBay for about $14,000.
3.33x the memory for 9x the price.
In Llama V1, 7B was clever but kinda dumb, 13B was still short sighted, but 33B was a big jump to "disturbingly good" territory. 7B/13B finetunes required very specific prompting, but 33B would still give good responses even outside the finetuning format.
16GB is probably perfect for a ~13B model with a very long context and better (5 bit?) quantization, or just partially offloading a ~33B model.
Llama in general is not great for code completion and fact retrieval, from what I have tried
I can't run llama2 70B - but I remember llama1-Guanaco30B was a major leap over the base llama30B model for me.
And you can run this locally with enough combined RAM + VRAM.
First some background: llama is divided into prompt ingestion code, and "layers" for actually generating the tokens.
There are different offloading schemes, but what llama.cpp specifically does is map the layers to different devices. For example, ~7GB of the model could reside on the 8GB GPU, and the other ~9GB would live on the CPU.
During runtime, the prompt is ingested by the GPU all at once (which doesn't take much VRAM), and then for each word, the layers are run sequentially. So the first ~half of a word would run on your GPU, and the last half would run on your CPU, and the alternation repeats till all the words are generated.
The beauty of llama.cpp is that its cpu token generation is (compared to other llama runtimes) extremely fast. So offloading even half or two thirds of a model to a decent CPU is not so bad.
The M1 and M2 Apple processors might change that equation.
The 32gb and 64gb MacBooks can run inference on a lot of these big models.
They might be quite a bit slower in TFLOPS than a 4070, but if they can run the model while the 4070 can't run it due to limited RAM, then they come out the winner
The long and short of it is, right now, CUDA beats TFLOPS.
Now for training, $10K for an M2 with 192GB memory could break even somewhere around a dozen models with current costs, iff the compute is not saturated. The limit of memory bandwidth is only a function of the compute; it doesn't scale infinitely.
The card is 256 bit... Intel could have done this, just like the 16GB RTX 4060 with a 128 bit bus.
Why are gpu vendors deciding how much ram the gpu gets? This feels so 1980s. GenAI workloads have radically variable memory footprints, some folks want 256GB per card so they have a spitting shot at trying something out - and other folks want 8GB per card to maximize throughput.
To add complication, I really don’t want to deal with copying 256GB models between cpu/card. We should have unified memory.
Intel/AMD are reportedly coming out with wide M1-Pro like chips.
Just saw one yesterday, 128GB PCIE 5.0 x8 which is a maximum of 32 GB/s in and out, where the 4090 has 1 TB/s memory bandwidth.
IMO for inference, dual socket EPYC (24 channel DDR5) is the way to go and CXL can theoretically allow you to bump up the bandwidth in that situation (assuming that you can optimize the software properly). Already in llama.cpp there seems to be some issues using 2P servers regarding numa.
...because they make them?
You can go buy 256gb systems if you want, at extreme cost. You can also use an 8gb card for whatever you want too. What are they doing wrong?
> To add complication, I really don’t want to deal with copying 256GB models between cpu/card.
For deployment, it doesn't really matter. You just memory-map it to the CPU or keep it hot-loaded in GPU memory. With a unified memory model you're still limited by your disk speed, so the performance difference probably wouldn't come out great either way.
To my knowledge, this is pretty much how InfiniBand works on datacenter-scale systems like the GH200. The incentive and upsell for bringing this to consumer systems seems weak to me.
The problem is latency. Apple can do what they do in terms of performance because they place the RAM extremely close to the SoC and don't have sockets that present their own RF interference challenges.
If you want to run modular, unified RAM, it gets really messy really fast.
[1] https://www.ifixit.com/Guide/Macbook+Air+(M2+2022)+Logic+Boa...
Movidius AI accelerators are going to be there as part of every single Meteor Lake chip.
Yeah this is sad.
Intel is supposedly turning it into a tensor core-ish accelerator on Falcon Shores... And, hopefully, lower end Arc cards.
On LLM inference we’ve got a baseball game.