Similar to rewind.ai, I want a 100% offline AI with Vision to run on every file and image I touch. Every file I use can be categorized and available to prompt against.
Similar to rewind.ai, I want a 100% offline AI with Vision to run on every file and image I touch. Every file I use can be categorized and available to prompt against.
There is nothing stopping AMD from allowing the sale of an affordable 48GB 7900 (or 32GB 7800) right now... In fact they already do in the form of the W7900, for $4500. They just don't sell those at competitive consumer prices, as they are playing the anticompetitive artificial segmentation game right alongside Nvidia, because they somehow think their tiny workstation market is more valuable than seeding AI development on their cards.
Technically speaking, AMD could sell "cheap" 96GB+ GPUs by simply swapping the 7000 series memory controller die. This would take some time and money, but not an eternity + fortune like Nvidia would need to tape out a whole new GPU die.
Intel has no workstation market to lose. I suspect the Arc A770 was kinda late and skimpy for a 32GB "AI" SKU, but the next generation could be very interesting.
AMD only supports small fractions of the deep custom eco-systems built around models such as Stable Diffusion. Good luck getting half of the shit built for ComfyUI or Automatic1111 to run at all, and frequently those features are the killer ones which motivate folks to want to do AI in the first place. There's a reason why Nvidia is a first class citizen among all of the AI clouds. AMD is even behind Intel in this regard, because at least Intel had an existing robust ecosystem of accelerators (i.e. intel optimized sklearn/pytorch, OpenVino, etc)
AMD can be 4x better in every way compared to Nvidia hardware wise but if they can't get developer communities to build around their equivalent to CUDA and its ecosystem than it's not looking rosy for AMDs future.
Dec 6th was an announcement from Lisa Su about the MI300x. It also included a full on commitment to ROCm as well as other projects like pyTorch. They are clearly aware of what you're talking about, and working on it. It isn't something that can magically be fixed over night, but it is something that people can complain about on HN over and over and over again...
In my eyes, the incentives weren't there before. But, they sure are now!
You can't build a software platform by just willing it into existence. It means sitting down with developers, writing documentation and support users. It means tackling use cases other than llama so that users feel confident in building the next greatest thing on your platform. At it is today, AMD will always be playing catch up and people will continue to push nvidia because the latest and greatest will be on Cuda.
I guess for those of spending money on real gpus for the last 10 years, nvidia was incentivized to take our money but amd wasn't.
Games, rendering and crypto.
If I was a hardware company, that's all that I would have focused on. Certainly AMD could have done better on the software/developer front, nobody is arguing that.
I tend to live in the present and what matters is what they are saying and doing today. You can choose to be pissed off about the past, or you can work towards the future.
Right now, they are cranking out ROCm releases on schedule, contributing to a bunch of AI open source projects, listening to crazy people like George Hotz and have released a GPU product that leapfrogs the H100 and is built on OAM-UBB, at 1/3 the FLOPS/$, while also promoting fast ethernet standards.
I'd call that a pretty good start.
Better maintained inference projects are picking up considerable AMD/Intel support. Sometimes even Apple Silicon support.
ComfyUI has found a place in the image pipeline design niche, but A1111 still seems like the best choice for new users who just want to spin up Stable Diffusion and generate some art for a D&D campaign or whatever.
If there was a 40GB-48GB 7900 SKU, I would have bought it over a 3090 in a heartbeat, even at a big premium over the 24GB card. And I'd be debugging ROCm support in ML projects left and right.
But I am not. The 24GB SKU is just not worth the trouble over a 3090, which is more RAM efficient out of the box. I can't pay $5K for a W7900.
Intel is supporting SR-IOV in Linux with the latest Xe GPU driver they are developing, but it isn't clear to me that they'll allow it in the consumer cards. If they do, and if I can run a Linux host and Windows guest simultaneously on one card, I'll be on it like stink on a monkey.
Product segmentation might be anti-consumer (in the sense that they can charge higher prices overall), but it's hardly "anticompetitive".
>than seeding AI development on their cards
*ceding
W7900 margins are enormous. In a competitive market, AMD would massively undercut the 48GB RTX A6000 and the 24GB 4090, which Nvidia is price gouging even more. They absolutely can... But they don't.
Separately, its also (IMO) a bone headed business decision.
There are significant compromises in Apple's offerings.
> DDR5 octuples the maximum DIMM capacity from 64 GB to 512 GB.[8][3] DDR5 also has higher frequencies than DDR4, up to 8000 MT/s which translates into 64 GB/s (8000 MT/s * 64-bit width / 8 bits/byte = 64 GB/s) of bandwidth per DIMM.
Then there are issues of prefetching algorithms, cache eviction and concurrency strategies as well as speed and associativity, TLB size, branch prediction algorithm especially as it relates to pipeline length and flushes, decoder width and complexity and execution units available for dispatch.
There's just a lot more that goes into fast than memory bandwidth go brrrrr. Also x86 (and increasingly ARM and Risc-V) have many SKUs above those Apple offers in sleek packaging. Particularly silicon fabbed for high end desktop workstation or datacenter duty is a cut above what's available via other channels. For the cost of a Mac Pro you can easily afford an HP Z series or a whitebox from SuperMicro with dual sockets, 16 channels of DRAM, and enough PCIe lanes to make starving children cry.
If you can afford a max specced M3 you can also afford 2 RTX 4090 which should theoretically be faster.
If all you want to do is AI/ML then yes, a M3 is not your best option but if you want an extremely powerful laptop that can run AI models locally fairly easily (and it's well supported) then the M3 is a great choice. I love being able to download and test out AI models on my M3.
Still, trying to do significantly fun or interesting things in any major AI ecosystem, such as the Stable Diffusion one through Automatic1111 or with the LLMs in Oobabooga (which supports nearly all LLM backends i.e. llama.cpp), will be mostly crippled and stuff will break in ways that simply don't happen to Nvidia hardware. I'll take the 2 4090s all day, because I know that quantization techniques which shouldn't even be possible (who knows maybethe fabled 1 bit quantization seems inevitable at this point) will make it possible for me to stuff even the most bloated LLM into my measly 48 GBs will be available on Nvidia first and Apple (maybe) second.
That or devs will start building Linux hosts for their gpus at home.
Might be a good hold-over for X years until the consumer hardware catches up and/or the model optimizations make the same hardware perform up to today's DC hardware.
Unfortunately I suspect VRAM capacities in consumer cards will continue to be limited in order to differentiate enterprise products.
[0] https://www.dramexchange.com/
[1] https://www.tomshardware.com/news/gddr6-vram-prices-plummet
I doubt 48GB cards are using 8 Gbit chips. How are you going to physically route 48 chips on a PCB?
What a lot of people dont understand is that RSP dont change with commodities ups and downs. If you price your product $1500 cheaper now due to DRAM cost reduciton, you will have a hard time getting it back up $1500 if things ever resume to normal. So generally speaking RSP comes down very slowly.
But the upshot for a manufacturer for higher density memory, if the 'failure rate' of fresh parts is the same, well it's better to have less parts of a fixed failure rate ofc.
Not that it isn't appealing... If the price is right.
Also, both AMD and Intel are both reportedly coming out with M Pro-like SoCs. AMD's "Strix Halo" is rumored for 2025. Less is known about Intel's Arrow Lake SKU at the moment.
In some cases one can replace rocm/openvino incompatible libraries with pure PyTorch, but the performance hit is often severe, and not many inference projects are using torch.compile or triton kernels to compensate.
Performant Pytorch compatibility can be achieved without CUDA compatibility.
> In some cases one can replace rocm/openvino incompatible libraries with pure PyTorch
This is the crux of my point: AMD ought to focus on fixing these Pytorch incompatibilities instead of splitting focus with chasing CUDA translations.
The point really is that while deep learning libraries are amazing, at the end of the day they are DSL and really pull towards one specific way of computing and parallelization. It turns out that way of parallelizing is good for deep learning, but not for all things you may want to accelerate. Sometimes (i.e. cases that aren't dominated by large linear algebra) building problem-specific kernels is a major win, and it's over-extrapolating to see ML frameworks do well with GPUs and think that's the only thing that's required. There are many ways to parallelize a code, ML libraries hardcode a very specific way, and it's good for what they are used for but not every problem that can arise.
Unless something fundamentally changes with the capabilities of produced tech I think demand will be saturated before 5 years.
As a result of this split I think consumer hardware will still continue to grow slower. Not because of lack of capability but lack of market interest. People don't have a use case for hundreds of cores or the like, they'd rather smaller, quieter, and more efficient at ever decreasing price points. I'm much more worried investment in consumer GPUs and CPUs for traditional desktops will actually die out and start to stagnate just because the market is getting squeezed ever smaller from both the low end and high end sides.
Take almost any given model that fits in X memory pool for some task. Double that, and everything gets better. Training quality goes up, inference quality potentially goes up, you gain the capacity for more batching which dramatically increases performance. Maybe more caching, depending on the app. You can add more models to the pipeline without much fuss.
But the big models that are in fashion now... GPUs are basically always memory capacity limited to some extent.
Intel canceled, and AMD sidelined, big APUs because they found the market didn't want them.