I'm super duper curious if there are ways to glob together VRAM between consumer-grade hardware to make this whole market more accessible to the common hacker?
Are you sure about that?
Depending on the model the performance is sometimes not all that different. I believe for solely inference on some models the speed difference may barely be noticeable, where for other training activities it may make 10+% difference [1]
[0] https://pytorch.org/tutorials/intermediate/model_parallel_tu...
[1] https://huggingface.co/transformers/v4.9.2/performance.html