I wonder how things would change if the text-to-speech were done server-side? I'd imagine one can get significantly better quality if you're not doing text-to-speech on the 400Mhz processor on the Kindle 2.
Also, what if you're using it as a "book on tape" in the car, on a road trip through the middle of nowhere? You lose network connection, you lose your book on tape.
tldr: I think it would be better quality, when you could get it, but at some level reliability is more important for this application.