I've had difficulty obtaining useful results from the smaller (7B-sized) models. The issue lies in the content, not the speed. If you could stream the text-to-speech, the speed alone would be satisfactory.
100wpm: Max typing speed
200wpm: Max speaking speed
300wpm: Max listening speed, max reading speed with subvocalisation
900wpm: Max reading speed without subvocalisation