A Petals dev here. It is not real-time, but we think the speed of ~1 token/sec may be enough for some interactive apps such as chat bots (especially, if you show tokens to a user once they are generated). You can try one at http://chat.petals.ml (heads-up: it may be laggy right now due to lots of HN users trying out the system).
Of course, you could do better if you have enough high-end GPUs to host the entire model yourself (3x A100 or 8x 3090). But if you don't, 1 token/sec is much faster than what you get with other existing methods.