Yes, so you would have a vector about 8k values long to be transferred on each token generated.
You could do that easily with any modern network.
You could do that easily with any modern network.
I wonder if training could be done in this way too.
But I guess the final output needs to be sent back to the first node before it can continue. So if there are 50 nodes with a latency of 40ms each, each token would take 2s to process.
Out of curiosity, would you ever support training or at least fine-tuning?