The state of silicon and the GPU poors
latent.space
latent.space
Debian runs like a champ on my 8 year old CPU. When will gpus last that long?
However, I do want to complain about the poor tooling around GPUs, particularly package managers, and closedness of drivers. Getting stuff to work on your GPU is often a shitty and time-consuming experience.
The MacBook Pro I bought in the same year is officially classed as "obsolete" by Apple.
I've also been running the same i5-3470 CPU since it was new.
Computers can last a long, long time if you can bear the thought of not being on the bleeding edge. Otherwise, yeah, you're going to be paying a premium for brand new tech. That's the trade you make chasing the newest and shiniest thing.
Well AMD just ended support for Vega and Polaris architectures so you're shit out of luck there, proving the point of this thread.
And AMD was selling Vega APUs all the way in 2022. So imagine a year later you get the news that the chip in your system is not getting updates because it's architectures is too old. Lol.
I wish I couldn't forget AMD for this.
We give shit to Microsoft for W11 needing quite modern hardware even though the 10 year old W10 will still keep working after support stops, so why don't we hold HW manufacturers to a similar standard for supporting their shit longer than a year?
So forgive me for feeling stabbed in the back by being dumped on only a year after purchase.
My main computer runs a nvidia NVS5400m. It's been out of support for a very long time. I genuinely wasn't even aware that my AMD GPU was still in support at all.
Support affects me exactly none. I'll just use whatever default driver my Linux distro provides and never think of it. I'll get whatever kind of performance it provides. As long as it's above 30fps and textures aren't obviously broken, I don't care.
Add in the increase in energy costs... I don't know, really.
It is currently tailored to gaming, but ML workloads are coming shortly. Everything runs in an invisible ephemeral VM.
The current state of local inference/finetuning is insane, where the hardware that makes any financial sense is ancient Nvidia GPUs, like the 3090 (2020) or the RTX 8000 (2018). Or maybe the rare used AMD Radeon Pro, if you can actually find one.
The only hope seems to be the rumored large APUs Intel/AMD APUs and maybe Intel Battlemage, since AMD is seemingly complicit in preserving the low VRAM status quo with Nvidia.
You run into immense pain the moment you venture outside of llama inference though.
The more VRAM you have, less aggressively you have to quantize models for inference, which in turn has huge speed/quality implications. You can run higher batch sizes, or draft models, or more caching, which increases efficiency. For LLMs specifically, you can load bigger models into VRAM in the first place. You can load more of a multimodal pipeline in VRAM without having to constantly swap everything out.
This is all 10x true for finetuning. Quality and speed is essentially determined by VRAM capacity, as long as you are not on truly ancient GPU like a P40 than't can't even do fp16.
As for architecture... TBH, many operations are heavily bandwidth bound these days. Sometimes a 3090 and a 4090 are essentially the same speed. And waiting a little longer for a finetune is no big deal vs not being able to do it at all, or doing it at low quality.
> They are graphical cards, equipped with consumer grade VRAM (instead of HBM) and IMHO it will take a very big shift before we see them being used as real AI accelerators
I think its important for users to break away from the cloud and APIs, and try to run stuff themself, lest we get locked into an OpenAI monopoly.
But setting that aside, running local is also extremely useful for prototyping and testing. You can see if something works without burning dollars every second you spend debugging on a big cloud instance. Even if that's affordable, just feeling like I am under the clock when debugging/optimizing is stressful to me.
Not sure how well it will perform as it lacks some of the smaller data type acceleration available on newer cores, though.
Of course you'll get a better service if you pay more. Expecting a high performance, easy, turnkey solution for a low price is expecting to have your cake and eat it. It's never happened on any other technology with this level of demand, and expecting it now seems naive at best. There's too much money (and hype) sloshing around in the sector, and limited supply.
Enthusiasts and experimenters have always traded their time for cost - if only because their time investment and learning (and maybe fun) is the whole point, not the end results of their experiments. If you expect financial returns due to your incredible idea, you should be looking for investors not old hardware.
The thing is, they are more geared towards HPC/Scientific computing. The only thing that really makes them appealing for ML is that they have somewhat more price depreciation than Nvidia cards and M1 Maxes, for the moment.
> Expecting a high performance, easy, turnkey solution for a low price is expecting to have your cake and eat it.
But I don't agree with this. This is not a normal "premium" hardware market, its a greedy pseudo monopoly that also dramatically supply constrained at the moment, until maybe next year (as the article points out). Business can pay a premium, but that's different than a lack of competition and supply.
https://old.reddit.com/r/LocalLLaMA/comments/17vcsf9/somethi...
Another thing that jumps out:
> I also have a couple W6800's and they are actually as fast or faster than the MI100s with the same software...
That's insane. The MI100 should be so much faster than the W6800 (a ~6900XT) that its not even funny.
1. It is not a recent GPU, so don't expect blazing inference speeds
2. It is designed for external forced air, so you will need to push air through it somehow
3. The one I got is actually not a normal card, I think it was a firmware testbed, and has no VBIOS and reports its product name as "TBD"
It does work and have 32gb of VRAM though.