What an fascinating concept. I guess this won't be useful for any kind of realtime feedback system, though?
Of course, you could do better if you have enough high-end GPUs to host the entire model yourself (3x A100 or 8x 3090). But if you don't, 1 token/sec is much faster than what you get with other existing methods.
You write that speed can be inferred, but the analogy that was used here is BitTorrent—and my experience with BitTorrent tells me that it certainly cannot be inferred.