So... Generally the quality is worse, but the available set of finetunes is totally different. Some llama v1 33b finetunes are not available in 70B, and extremely good at their niche.
Also 70B should get more than 1 token/sec on a single 3090 offloaded to CPU. I dunno what framework op is using.
- theraputic/friend style chat (Samantha)
- translation (various single language finetines)
- medical advice (can't remember this one)
This is non exhaustive. And Llama V2's extended native context does really help some niches (like storytelling) that a few 33B models are still pretty good at.
You're probably thinking of Clinical Camel: https://huggingface.co/augtoma/qCammel-70-x
Right now, I'm using pure CPU Llama but only the 17B version, based on I believe llama.cpp. How do I mix both CPU and GPU together for more performance?
Then offload as many layers as you can to the gpu with the gpu layers flag. You will have to play with this and observe your gpu's vram.