5090 likely won't have more than 32 GB, if even that much.
0. https://manifold.markets/Tenoke/how-much-vram-will-nvidia-50...
I have 3090 Ti and I can run Q4 quant 33b models at 30t/s with 8k context. A 4090 would allow me to do the same but with ~45t/s, both inference speeds are more than fast enough for people so 3090 is the usual choice. In my tests on runpod, H100 with 80GB memory is around the same speed as 3090, so slower than a 4090.
Odd statement. I don't really know what you mean by that. Perhaps 'math _works_, code should too' ?
I would definitely agree that it _should_ work.
I'm of the belief that no one should _have to_ publish (e.g. to graduate, get promotions, etc) in academia, and that publications should only occur if they're believed to be near Novel prize worthy, and fully reproducible by code with packaging that should last and work in 10 years, from data archives that will exist in 10 years.
But it seems I have been outvoted by the administration in academia.
Hence, we get this "ai that doesn't run" phenomenon
Do you want to publicly fund researchers only for the industrial research partner's benefit?
My main point was that there is a lot of noise in scientific journals that are caused from pressures in academia that are requirements if publishing. If these are removed, then the quality of work published increases and quantity decreased.
There are other places to post work that is derivative and non-novel like blogs. The field of biology has an immense amount of work that is mostly observational without strong conclusions or predictivity. A tabulation of observation should definitely be put out by a lab, and it should be much sooner with far less pressures than today, such as the typical dance of putting the data in during publication. The SRA is one example of a place to share data. If the typical way to work was put all data immediately onto a public repo, sometimes comment on it in ways that have been seen before on blogs and other classes below scientific journals, and then if something truly substantial comes out of it (a novel model that is analytical and highly predictive of cell behavior in all situations for example) then publish.
It could alleviate the noise from the signal. LLMs is one case where the noise is very strong in that many papers are simply 'we fine tuned an llm'.
That said, I certainly think that researchers can do more to make their code and data more accessible. We have the tools to do so already but the incentives are often misaligned.
> Odd statement. I don't really know what you mean by that. Perhaps 'math _works_, code should too' ?
It was a typo. I mean "Everything should run in it" as in most LLM should be able to run at least quantized in 4090
https://github.com/vosen/ZLUDA
That's a binary level wrapper. Of course there's also ROCm HIP at the source level, and many other things, such as SYCL
[1] https://www.tomshardware.com/pc-components/gpus/chinese-work...
Stable diffusion, stable video, text models, audio models, I never had issues with anything yet
There's a lot of open weights activity around 7B/13B models which the 4090 will run with ease. But you could can run those OK on much cheaper cards like the 4070Ti (which is of course why they're popular).
And there's a lot of open weights activity around 70B and 8x7B models which are state-of-the-art - but too big to fit on a 4090. There's not much activity around 30B models, which are too big to be mainstream and too small to be cutting edge.
If you're specifically looking to QLoRA fine-tune a 7B/13B model a 4090 can do that - but if you want to go bigger than that you'll end up using a cloud multi-gpu machine anyway.
Traditional CPU-bound physics/simulation models have typically wanted all the RAM they could get; the more RAM the more accurate the model. The same is true for AI models.
I can max out 24GB just using spreadsheets and databases, let alone my 3D work or anything computational.
No matter how much vram you have, there's something that doesn't fit :)
Nvidia have held back the majority of their cards from going over 24GB for years now. It's 2024 and my laptop has 96GB of RAM available to the GPU but desktop GPUs that cost several thousands just by themselves are stuck at 24GB.
This is like Intel and their refusal to support ECC memory; when AMD does on nearly all Ryzens.
—
Note: your laptop is probably using a 64-bit memory bus for system RAM. For GPUs, the 4090 is 384-bit. That takes up a lot more die area for the bus and memory controller.
I guess my point is, rather than give the cards more RAM, the gaming cards should just be priced cheaper.