So... Generally the quality is worse, but the available set of finetunes is totally different. Some llama v1 33b finetunes are not available in 70B, and extremely good at their niche.
Also 70B should get more than 1 token/sec on a single 3090 offloaded to CPU. I dunno what framework op is using.