CPU Version: https://huggingface.co/TheBloke/guanaco-65B-GGML
GPU Version: https://huggingface.co/TheBloke/guanaco-65B-HF
4bit GPU Version: https://huggingface.co/TheBloke/guanaco-65B-GPTQ
CPU Version: https://huggingface.co/TheBloke/guanaco-65B-GGML
GPU Version: https://huggingface.co/TheBloke/guanaco-65B-HF
4bit GPU Version: https://huggingface.co/TheBloke/guanaco-65B-GPTQ
How much ram and vram does one need to run 4,13,33,65B models at a reasonable speed?
edit: instead I'll ask this, what's the best model to run on a system with a 24gb 4090 and 64gb of ram?
You can run 4-bit quantized 65B models on your cpu, but it is slow, 1-2 tokens a second instead of 8-15 people typically get with a gpu, but you need two 24gb or an enterprise card with 48gb of ram to load them there.
https://old.reddit.com/r/LocalLLaMA/wiki/models has the information you are irritated about not being listed.
"Reasonable speed" is subjective, though.