I have 2x 3090 do you know if it's feasible to use that 48GB total for running this?
Ooobabooga: https://github.com/oobabooga/text-generation-webui
Model: TheBloke_Llama-2-70B-chat-GPTQ from https://huggingface.co/TheBloke/Llama-2-70B-chat-GPTQ
ExLlama_HF loader gpu split 20,22, context size 2048
on the Chat Settings tab, choose Instruction template tab and pick Llama-v2 from the instruction template dropdown