Which consumer gpu runs llama 70B?
Yi 34B is a much better fit. I can cram 75K context onto 24GB without brutalizing the model with <3bpw quantization, like you have to do with 70B for 4K context.
Llama 70B is a huge compromise at 2.65bpw... This does make the much "dumber." Yi 34B is much better, as you can quantize it at ~4bpw and still have a huge context.
The perplexity graph here is a pretty good illustration: https://github.com/ggerganov/llama.cpp/pull/1684
YMMV, as Mistral and Yi are not necessarily comparable like different sizes of llama, and it depends on the task.
MacBook Pro M3 Max.
Or wait another month or so for https://ChatOnMac.com