Running the Q4 quant (14GB or so in size) at 46 tok/sec on a 3090 Ti right now if anyone's curious to performance. Want the headroom to try and max out the context.
Given that, I'd expect a single 3060 (if a large enough one existed) to run at about 16 tok/s so 20 tok/s on two isn't bad not being NVLinked.