You can, for instance, connect two RTX 3090 with an NVLink bridge. That gives you 48 GB in total. The 4090 doesn't support NVLink anymore.
Depending on the model the performance is sometimes not all that different. I believe for solely inference on some models the speed difference may barely be noticeable, where for other training activities it may make 10+% difference [1]
[0] https://pytorch.org/tutorials/intermediate/model_parallel_tu...
[1] https://huggingface.co/transformers/v4.9.2/performance.html
Are you sure about that?