Did you read the paper? They intentionally steered the quality to ensure they sound fake. Their generated speech is "very easy to detect" according to the reference at the end of the paper.
I found that in my (high quality) studio monitors, the audio sounded fine and hard to distinguish from 24kHz wav. But in headphones, the artifacts were pretty obvious. So probably some reverberation will do a lot to cover up artifacts. In the paper, they only do a subjective comparison between the generated audio and the soundstream-encoded original audio, which seems a bit disingenuous. Listening to soundstream audio in headphones, I can hear those same artifacts.
just to be clear, one could mistake them for some (voice) actor reading a book (maybe) but even to my untrained ear they sound fake and artificial.
Am i missing something?
It's meant to sound artificial. The focus is on speed and consistency