AudioGen: Textually Guided Audio Generation
felixkreuk.github.io
felixkreuk.github.io
Given that, I don't think AudioGen particularly needs to add full narration. That seems like a very different problem to me, likely requiring a completely different architecture.
[0]https://paperswithcode.com/method/tacotron-2 [1]https://speechresearch.github.io/naturalspeech/
[0] - https://twitter.com/FelixKreuk/status/1575846953333579776
i giggled :)