Also, did you notice that in all major animated films (think Disney or Pixar), while the imagery are all computer-generated, the voices are not?
[0]: https://machinelearning.apple.com/2017/08/06/siri-voices.htm...
https://www.youtube.com/watch?v=DWK_iYBl8cA impresses me and it just feels like a low quality recording.
It's concerning given how computing power and resources advance over time.
and an unnamed start-up a while back: http://web.archive.org/web/20190803012012/https://instaud.io...
What it doesn't do is the acting. In an audiobook, the voice actor will change their voice in various ways for dramatic effect, and in that respect the book becomes something like a radio drama. With TTS you're getting "just the book". I think that's the major difference and the refuge in which voice actors might hope for continued employment.
Having said that, some fairly small scale audiobooks that have the authors narrating them are also very good as you can hear the interest and the passion of the author in their subject:
e.g.
I guess what I'm saying is don't underestimate neuroplasticity. I wager you could even achieve casual fluency with morse code if you listened to it long enough. I'm under the impression that some telegraph operators did.
The best/only way to get most of that computer-generated imagery is by huge amounts of manual labour: designing, animating, simulating, sometimes motion-capturing. It's painstaking detail work involving many people.
The best way to get the voices is with a small amount of manual labour: voice acting.
If you put as much manual effort as the imagery into controlling the nuances of a TTS engine, you might get acceptable results, but it's far easier and cheaper to use voice actors. In fact, the easiest way to tell a TTS engine exactly what you want would probably be to voice act and have it mimic you. This might be worth trying to do if remapping vocal anatomy (e.g. woman voicing man or vice versa, or monster, etc.), but for most purposes it's easier to hire appropriate voice actors and/or manipulate the vocal recording audio than to use it to drive a resynthesis by simulation.
Those are deliberately made to sound unnatural. Not to say it changes anything, and they've already shown up once or twice in anime.
(Though the only example I can name off the bat is Black Rock Shooter, and that doesn't include the voice. It's complicated. Mato is complicated, too.)
Well, big budget theatrical animated films from major studios like Pixar or DreamWorks, sure.
But most animated films aren't those.
It's similar to saying that voice actors (and then actual visual actors) will soon be out of business because we can soon 'automate' that too...
Voice (audiobook) acting is a _performance_, not a routine task to tick the "available on Audible" box.
For anyone who hasn't ever read the Dresden Files, it is a fantasy series set in modern Chicago with a young wizard way in over his head. The series starts off kinda bleh with the first few books, but steadily picks up pace as the main character gets more involved with all the crazy things going on. The author puts basically every single (ok a lot) of diverse mythologies as if it is all real (Odin, Mab, Erlking, Skinwalkers, Trolls, the Fae, Necromancers, 4 different kinds of Vampires, Roman Gods, the White God, Angels, Demons, Lucifer, Dragons, Cthulhu, and about a dozen other bad and good things all with a single coherant plot. The main character starts to see that all of these major supernatural entities are moving their armies like 3D chess. The characters take some time to develop, but are seriously good.
The dresden files subreddit has one of the author's beta readers that sends us some updates although no spoilers of course. She says it will be as big of a whammy as "Changes".
There was also a new Goodman Gray shirt story released recently.
However, I'd guess that it's simply quicker and more efficient for someone to just properly act and speak the lines with a dramatic effect and any explicit annotation of how exactly they should be said takes much more time and effort.
Having the choice between a monotone narration and the voice actors in this case would be a much welcome improvement.
I am not sure if some of the brands advertising on YouTube realise what crap inventory their adverts are being run on.