You got it wrong. Inference can use crap GPU's. Training needs the 100x more expensive big guns. Our training machine is 100x more expensive than our inference machine.
Then have inference go down to the next layer to use those models as a P2P decentralized network.
Maybe like open router could tap federation networks.