That's exciting. So we could build a SETI@home style network of even the largest models.
I wonder if training could be done in this way too.
I wonder if training could be done in this way too.
But I guess the final output needs to be sent back to the first node before it can continue. So if there are 50 nodes with a latency of 40ms each, each token would take 2s to process.
Out of curiosity, would you ever support training or at least fine-tuning?