I'll keep this in mind!
I'll keep this in mind!
I recently bought a 7900 XTX with 24 GB of VRAM, but the model I currently run can easily run in 16 GB (6 bit llama 3 8b). It's fast enough and high enough quality that I can use it for processing information that I don't feel comfortable sharing with hosted services. It's definitely not the best of the best as far as what models are able to do right now, but it's surprisingly useful.
Unless of course you were talking about VRAM, in which case 16GB is still not great for ML (to be fair, the 24GB of an RTX 4090 aren't either, but there's not much more you can do in the space of consumer hardware). I don't think the other commenter was talking about VRAM, because 16GB VRAM are very overkill for everyday computing... and pretty decent for most gaming.
You don’t need a GPU for llm inference. Might not be as fast as it could be but usable.
I don't want to have a laptop over 3 pounds and I'm not spending over 1100$, so a dedicated GPU isn't really an option.