The training methods for current LLMs are parallellizable only by very frequently transferring all the data back and forth, and needing a node to gather and merge all the updates very frequently, and redistribute it to every other node so that they can make any progress. And a GPU that's not connected you with a high-speed link is pretty much useless as you can't make useful progress until you get their part pack, and "their part" is very large (i.e. the update size you need to get back is comparable to all of the model size) and you need to do that very frequently. Training on nVidia many-GPU pods works because of high-speed interconnect (e.g. 600 gigabytes per second for 8 GPUS in A100 pod), and if your internet bidirectional speed is much less than 600gbps, then if you have 100000 free remote RTX 4090s, you simply can do the compute locally faster than you can exchange information with the other GPUs.