How is the result of training stored? How big is that? It seems reasonable to assume we’ll eventually plateau and all we’ll need is relatively infrequent training.
Then have inference go down to the next layer to use those models as a P2P decentralized network.
Maybe like open router could tap federation networks.