Not an answer to your question, but maybe this has more info? I think with these optimizations and quants, it compares negatively. But can these optimizations be applied to models you want? Another question.
https://news.ycombinator.com/item?id=48353348
if you take a model that requires 200GB of VRAM and you run it on the CPU, it requires 200GB of RAM instead. Still unfeasible on consumer hardware. With this approach you can easily do it on 12GB or less of either RAM of VRAM, at several seconds per token instead of tokens per seconds. Very unusable, but certainly interesting!
If token speed doesn't count, you can have 200GB "RAM" in swap space.
But I agree, what the OP does is a lot more efficient than this.