I got worse than 1 token/sec, and yes, wasn't impressed with bloom results, but I believe it's also very foreign language heavy. I haven't tried it yet but I believe flexGen benchmarked faster as well.
So, I believe 1 sec/token with Petals is the best you can get for the models of this size, unless you have enough GPUs to fit the entire model into the GPU memory (you'd need 3x A100 or 8x 3090 for the 8-bit quantized model).