Importantly, the recording should indicate whether it was human or AI generated.
Importantly, the recording should indicate whether it was human or AI generated.
This is all that's necessary. Sometimes I'm fine with mediocre TTS; sometimes I want an actual professional; librivox is somewhere in between, but should clearly specify whether I will be getting an amateur human or a robot.
Historically, being told that a voice recording is AI generated would be enough to tell you to expect basic TTS robotic voice, but with advances in AI voice generation we're approaching the point where AI can sound as good as real humans - it's not yet to the point where it's easy to generate an audiobook as good as a professional reader, but that point will come in the not too distant future.
And equally on the other side, something being recorded by a human doesn't automatically mean it has the quality of a professionally-read audiobook. This is something LibriVox has always had to deal with, by gatekeeping which volunteer recordings to either give feedback requesting improvements to or to not use at all.
In some but not all cases, an amateur human reader can already be as good as a professional, that will soon be true for AI. For both AI and humans it will remain the case that some efforts are not as good, but the line between them (for quality) isn't going to be whether or not they are AI - though I do agree that AI or not should also be labelled.
For an example of a professional audiobook, check out Rob Inglis' version of The Lord of the Rings.
But I disagree with you when you write "it simply doesn't have the information to improve beyond sounding like a human reading words fluently" - it has the same information when reading it as a human does, meaning that the best implementation would have to not only adapt tone to explicit instructions like "... she shouted", but also read between the lines / make subjective choices to suit the different characters.
AI is already capable of doing sentiment analysis on text, and text to speech models are getting better at being able to simulate moods/emotions rather than just speaking flatly, and I don't think we're many years away, if that, from those two sides being paired together in a way that produces the sort of quality output we're talking about for the first time without human involvement. Add to that the fact that AI can train on the many good examples of humans reading things, they may get to the point of emulating not just the core accent but also how each accent should adopt to what meanings in the text and arrive at a great solution without even needing to go through the steps of analysing what the text means to use that to know how to modify the voice being generated.
To take an example, here's an iconic line from the Fellowship of the Ring:
> The wizard swayed on the bridge, stepped back a pace, and then again stood still. ‘You cannot pass!’ he said.
If you think that is a command, you should shout it like Ian McKellen in the movie. If you think it's a statement based on superior knowledge (see https://acoup.blog/2025/04/25/collections-how-gandalf-proved...), you should probably state it with certainty and fatigue. And if you're making a movie with a ton of crazy special effects and swelling music, you should probably make whatever choice goes best in that context.
Even if a model could make some consistent choice there, I wouldn't be all that interested, because the reader conveying their interpretation of the character to the listener is what matters. Sure, it might get enough Spotify plays to make some money, but it's not art.
I just close the tab when I realize it is AI. Not sure how long I can do this.