You can steal the voice of somebody like Lester Holt or Tanaka Rie with about 20 hours of audio where you manually splice segments to compose sound. Sometimes you can lift a sentence, sometimes you compose paired phonemes. It is a lot of work, but it has been possible long before deepfakes.
No matter what you will need to listen to the voice and give feedback, pro narrators have directors just as actors do and it is what makes them sound so good.
As others have said, to an extent you could program this without AI using some current techniques but it would be impractical. An area that might help in this regard is efforts with GST, global style tokens, as this should allow more variation. Clearly more work needs to be done to get it to be more acceptable, but there are some examples here: https://google.github.io/tacotron/publications/global_style_...
No matter what you will need to listen to the voice and give feedback, pro narrators have directors just as actors do and it is what makes them sound so good.
On the other hand, https://www.theverge.com/2019/8/23/20830057/amazon-audible-s...
If a console integrates the system into its OS then maybe that can help reduce file size and also allow indie devs to add a lot of voice to a game without having to record a ton.
Well after we get a best selling book series that's completely written by a bit of software.
So, not very soon.