After that, in theory, you can just take existing template and infrastructure, feed it with needed data, including voice synthesized with a different library (it would be the developer's responsibility to find a fitting one), and generate as many total videos as you would need.