https://arxiv.org/ has a ton of papers on it.
See also: https://github.com/learning-at-home/hivemind
and more to OP's incentive structure: https://docs.bittensor.com/
Latter two intend to beat latency with Mixture-of-Expert models (MoEs). If the results of the former hold, it shows that with a simple algorithmic transformation you can merge two independently trained models in weight-space and have performance functionally equivalent to a model trained monolithically.