The more VRAM you have, less aggressively you have to quantize models for inference, which in turn has huge speed/quality implications. You can run higher batch sizes, or draft models, or more caching, which increases efficiency. For LLMs specifically, you can load bigger models into VRAM in the first place. You can load more of a multimodal pipeline in VRAM without having to constantly swap everything out.
This is all 10x true for finetuning. Quality and speed is essentially determined by VRAM capacity, as long as you are not on truly ancient GPU like a P40 than't can't even do fp16.
As for architecture... TBH, many operations are heavily bandwidth bound these days. Sometimes a 3090 and a 4090 are essentially the same speed. And waiting a little longer for a finetune is no big deal vs not being able to do it at all, or doing it at low quality.
> They are graphical cards, equipped with consumer grade VRAM (instead of HBM) and IMHO it will take a very big shift before we see them being used as real AI accelerators
I think its important for users to break away from the cloud and APIs, and try to run stuff themself, lest we get locked into an OpenAI monopoly.
But setting that aside, running local is also extremely useful for prototyping and testing. You can see if something works without burning dollars every second you spend debugging on a big cloud instance. Even if that's affordable, just feeling like I am under the clock when debugging/optimizing is stressful to me.
You run into immense pain the moment you venture outside of llama inference though.