The results look interesting, however.
Here's hoping that they'll add GTPQ 4bit quantizing so the 65B version of the model can be run on 2x 3090.
The results look interesting, however.
Here's hoping that they'll add GTPQ 4bit quantizing so the 65B version of the model can be run on 2x 3090.
But justai.com would also be apt
There's a subreddit r/LocalLLaMA that seems like the most active community focused on self-hosting LLMs. Here's a recent discussion on hardware: https://www.reddit.com/r/LocalLLaMA/comments/12lynw8/is_anyo...
If you're looking just for local inference, you're best bet is probably to buy a consumer GPU w/ 24GB of RAM (3090 is fine, 4090 more performance potential), which can fit a 30B parameter 4-bit quantized model that can probably be fine-tuned to ChatGPT (3.5) level quality. If not, then you can probably add a second card later on.
Alternatively, if you have an Apple Silicon Mac, llama.cpp performs surprisingly well, it's easy to try for free: https://github.com/ggerganov/llama.cpp
Current AMD consumer cards have terrible software support and IMO isn't really an option. On Windows you might be able to use SHARK or DirectML ports, but nothing will run out of the box. ROCm still has no RDNA3 support (supposedly coming w/ 5.5 but no release date announced) and it's unclear how well it'll work - basically, unless you would rather be fighting w/ hardware than playing around w/ ML, it's probably best to avoid (the older RDNA cards also don't have tensor cores, so perf would be hobbled even if you could get things running. Lots of software has been written w/ CUDA-only in mind).
I haven't tried with LLaMA at all.
> Current AMD consumer cards have terrible software support and IMO isn't really an option. On Windows you might be able to use SHARK or DirectML ports, but nothing will run out of the box.
I was merely sharing that I did not have that same experience that current consumer cards have terrible support.
For RDNA2, you apparently can get LLMs running, but it requires forking/patching both bitsandbytes and GPTQ: https://rentry.org/eq3hg - and this will be true for any library (eg, can you use accelerate? deepspeed? fastgen? who knows, but certainly no one is testing it and AMD doesn't care if you're not on CDNA). It's important to note again, anything that works atm will still only work with last-gen cards, on Linux-only (ROCm does not work through WSL), w/ limited VRAM (no 30Bq4 models), and since RDNA2 tensor support is awful, if the SD benchmarks are anything to go by, performance will still end up worse than an RTX 3050: https://www.tomshardware.com/news/stable-diffusion-gpu-bench...
Absolutely fair and I agree with this part. I started my reply with "FWIW" (For What It's Worth) on purpose.
> For RDNA2, you apparently can get LLMs running, but it requires forking/patching both bitsandbytes and GPTQ: https://rentry.org/eq3hg - and this will be true for any library (eg, can you use accelerate? deepspeed? fastgen? who knows, but certainly no one is testing it and AMD doesn't care if you're not on CDNA).
I haven't tried any of the GPU-based LLMs yet. SD leveraging PyTorch (which seems to have solid ROCm support) worked for me. It will not be faster than NVIDIA for sure but if someone already has a 16GB+ AMD card they may be able to at least play with stuff without needing to purchase an NVIDIA card instead.
Currently the 4090, the rumor is the 4090ti will have 48gb of vram, idk if its worth waiting or not.
The more VRAM the higher paremeter count you can run all in memory (fastest by far).
AMD is almost a joke in ML. The lack of CUDA support (which is nvidia proprietary) is straight lethal, and also even though ROCM does have much better support these days, from what I've seen it's still a fraction of the performance of what it should be. I'm also not sure if you need projects to support it or not, I know pytorch has backend support for it but I'm not sure how easy it is to drop in.
I mean in all honestly there's no reason a gaming card would need 48gb at the moment when so few games even use 24gb.
48GB really only makes sense for workstation cards.
Get a 3090 or 4090. Forget about AMD.
[1] https://www.dell.com/en-us/shop/nvidia-ampere-a100-pcie-300w...
Do I need dualboot? Or is Windows good?
One caveat though, my asus b650e-f is barely supported by the currently used ubuntu kernel (e.g. my microphone doesn't work, before upgrading kernel + bios I didn't have lan connection...) so expect some problems if you want to use a relatively new gaming setup for linux.
4090 24GB is 1800USD, The Ada A6000 48GB is like 8000USD and idk where you buy it? So if you want to run games and models locally the 4090 is honestly the best option.
EDIT: I forgot - there is a rumored 4090ti with 48gb of vram, no idea if thats worth waiting for.
Seems you can get used RTX A6000s for around $3000 on ebay.
I think that's such a silly name for it, but oh well
Thanks for the correction!
Why do they do this? Sometimes consumer products are versioned weirdly to mislead customers (like intel cpus) - but these wouldn't even make sense to do that with as they're enterprise cards?
According to GPT-4 the next generation one will be called Galactic Unicorn RTX 6000 :D
WSL on windows apparently decent, or native PyTorch, dual boot windows/ubuntu still prob best tho.
You can’t switch which GPU Linux is using without restarting the session
Basically, you want nVidia, and you want lots of VRAM. Buy used for much more bang for the buck.
Depending on your budget, get:
- an RTX 3060 with 12GB or
- 1 used RTX 3090 with 24GB (approx twice as expensive as the 3060 but twice the VRAM and much faster) or
- 2 used RTX 3090 cards if you need more than 24GB.
Everything beyond that gets quite a bit more expensive because then you need a platform with more PCIe lanes, you may need more than one PSU and you will have problems fitting and cooling everything.
With two cards and 2x24GB you can run the largest version of the LLaMA model (the 65B variant) and all its descendants with 4-bit quantization inside your GPU's VRAM, i.e. with good performance. Can can also try some low resource fine-tuning variants (LoRa etc).
Oh and while you're at it also get a decent amount of RAM like 64GB or 128GB (it's very cheap right now) and a NVMe SSD. These models are quite large.
Example: A 30B parameter model trained at 16bit FP gets quantized down to 4 bit ints. 4 bits = 0.5 byte. 30 billion * 0.5 byte = 15GB of VRAM (plus a GB or few of other overhead)
For more real world discussion see
M3 is around the corner tho, and there's some announcement to come from intel or arm following their partnership. There's also the new card coming from intel that is supposed to be aimed squarely at machine learning workloads, and they don't have to segment their market by memory sizing like Nvidia do, but they aren't well supported as device targets, but a pair of these will likely be very cost effective if and only if they will get credible compatibility with the libraries and models
A more honest name would be Visual-Vicuna or Son-of-BLIP.
The model as a whole is just BLIP-2 with a larger linear layer, and using Vicuna as the LLM. If you look at their code it's literally using the entire BLIP-2 encoder (Salesforce code).
maybe even add an "a" for extra spice: Son-of-a-BLIP
The number of parameters used for GPT-4 is unknown.
[0] https://twitter.com/SebastienBubeck/status/16441515797238251...
So we're back to guessing ...
A couple of years ago Altman claimed that GPT-4 wouldn't be much bigger than GPT-3 although it would use a lot more compute.
https://news.knowledia.com/US/en/articles/sam-altman-q-and-a...
OTOH, given the massive performance gains scaling from GPT-2 to GPT-3, it's hard to imagine them not wanting to increase the parameter count at least by a factor of 2, even if they were expecting most of the performance gain to come from elsewhere (context size, number of training tokens, data quality).
So in 0.5-1T range, perhaps ?
https://www.reddit.com/r/IAmA/comments/12rvede/im_stephen_go...
Outside of the brand name ChatGPT, lay members of the general public are way more likely to call these chatbots (like Bard and Bing) “AIs” than “GPTs”. And although GPT could technically refer to any model that uses a Generative Pre-trained Transformer approach (although it probably wouldn’t be an open-and-shut case), the mark “GPT-4” definitely is associated with OpenAI and their product, and you can’t just use it without their permission.
Let's not discuss the amount of copyright licenses OpenAI has already infringed, too
At Brewer’s Art in Baltimore, MD they just released a beer called GPT (Green Peppercorn Tripel)[1]. They’re likely allowed to do that because a reasonable consumer would probably not actually think they had collaborated with OpenAI, because OpenAI does not make beer.
OP is releasing a model called “MiniGPT-4”. A reasonable consumer could look at that name and become confused about the origin of the product, thinking it was from OpenAI. This would be understandable, since OpenAI also makes large language models and has a well known one that they’ve been promoting whose brand name is “GPT-4”. If MiniGPT-4 does not meet that consumer’s expectation of quality (which has been built up through using and hearing about GPT-4) it may cause them to think something like “Wow, I guess OpenAI is going downhill”.
Trademark cases are generally decided on a “reasonable consumer” basis. So yeah, they can seem a little arbitrary. But it’s important for consumers to be able to distinguish the origin of the goods they are consuming and for creators to be able to benefit from their investment in advertising and product development.