Do you mean like as a simple text to speech application? There is a huge need for better quality audiobook output.
Do you mean like as a simple text to speech application? There is a huge need for better quality audiobook output.
True human-level audiobook reading would require understanding the whole story, which often assumes general cultural knowledge, which you'll only get from a model trained on LLM-scale data. If you asked GPT-4o's new end-to-end voice mode to read an audiobook you'd probably get a better result than any TTS model. I bet it would even do different voices for the characters if you asked it to.
Voice acting is quite literally done a sentence or at most a paragraph at a time. Often the recording order is completely different from the script.
An actor may very well record his final scene on the first day of a project, after the whole character arc has transpired. But you know, acting. They get fed a line with stage direction and do a bunch of takes and somehow it works.
Heck you might be a full blown Italian who can't say a word in English but with the right kind of jacket it comes out a banger: https://www.youtube.com/watch?v=-VsmF9m_Nt8
You mention Eleven labs being ahead, check out Suno. There is no LLM-scale anything involved there. The voice in this context is a musical instrument and there are lots of viable ways to tackle this problem domain.
Sure, voice acting for games or movies is done piecemeal. But the actor still gets information about the story ahead of time to inform their acting, along with their general cultural knowledge as a human. Most crucially, when acting is done in this way it is done with a human director in the loop with deep knowledge of the story and a strong vision, coaching the actor as they record each line and selecting takes afterward. When the directing is done poorly, it is pretty easy to tell.
Sure, for a movie or game you could direct a TTS system line by line in the same way and select takes manually, but it would be labor intensive and not at all automatic. And to take human direction the model would need more than just the text as input. Either a special annotation language (requiring a bunch of engineering and special annotated training datasets), or preferably a general audio-to-audio model that can understand the same kind of direction a human voice actor gets.
It is, though, I've just done it a few days back. Once you have a clean text extracted from the book (this is actually the difficult part, removing headers, footers, footnote references, etc.), you can feed it into edge-tts (I recommend the Ava voice) and you get something that is, in my opinion, better than most human performers. Not perfect, but humans aren't either (I'm currently listening to a book performed by a human who pronounces GNU like G-N-U).
Inflection, emotion, tone, character-specific affects and all that can really change the audiobook experience for fiction. Your mentioning of footnote references and GNU suggests you're talking about nonfiction, perhaps technical books. For that, a voice that never significantly changes is fine and maybe even a good thing. For fiction it'd be a big step down from a professional human reader who understands the story and the characters' mental states.
> For fiction it'd be a big step down from a professional human reader who understands the story and the characters' mental states.
On the contrary, I don't want the reader to understand anything, I just want the text in audio form and I will do the interpretation of it myself.
I think how you tend to listen might also matter. I mostly use audiobooks when I'm driving or otherwise doing something else that is going to claim a portion of my attention. Following the narrative and dialog is easier when the audio provides cues like vocal tone changes for each speaker / narrator.
You can find it on YouTube easily if you want an example.
To the extent that authors provide supplemental notes or instructions to human actors reading their books, that information would be helpful to provide to an automated audiobook reading model. But there is no reason for it to be in a different form than a human actor would get. Additional structure or detail would be neither necessary nor helpful.
However, in my opinion it would be a huge benefit, if this kind of metadata would be put into the ebook file in some way, so that it would be something extractable and not has to be detected. I think it would be enough to ID the characters and tag a gender and a mood in the book together with citations, so that you could add different speech models for different characters. That would also allow to use different voices for different characters.
I wrote a little tool called voicebuilder (which I will open source next year). It's a "sentence splitter" which is able to extract an LJSpeech training dataset for an audio file, epub file and length matching. Works pretty accurate for now, although it needs manual polish of the extracted model. Still way faster than doing it manually.
This way you can build speech sets of your favorite narrators and although you would never be allowed to publish them, I think for private use they are great!