As the tech gets cheaper, this calculous changes.
Though this is very impressive, it seems to take longer and longer to make those tiny improvements that make all the difference wrt believability.
N.B. The most convincing TTS I've ever heard (predating Lyre by quite a bit) generated things like this: http://web.archive.org/web/20190803012012/https://instaud.io...
https://drive.google.com/file/d/1zRvJEGJjTpKvvzel-J0agh3fKBn...
I'm integrating it into a "Snapchat filter" type app with lightweight social features just as a means to bring it to market and hopefully attract Facebook or Snapchat or Tencent into buying it. I'm building it to sell, essentially.
I need capital so I can fund my real ambitious start-up of end to end computational filmmaking. Graph-based story language, light field camera optics, tracking and localization in prerendered environments, content-aware shaders, real time storyboard population and automated editing, posture estimation and mistake correction...
With patent protection, I think it could unseat Disney and make more money than they do with Marvel and Star Wars.
I need a lot of capital to build my lab. Optics (good sensors and glass), a modest studio with rigging and tracking set up for experiments, and a handful of engineers.
https://apps.apple.com/us/app/celebrity-voice-changer-face/i...
I checked the Android app store and all the apps used text to speech before vocoding. I had no idea this existed on iPhone.
Thanks.
However, one of the problems of deep learning is that you need to have a good dataset first, but how are you going to build one? Well, the big game companies wouldn't reveal their models and textures that easily. For indie artists to utilize this technology, there needs to be a centralized community project built around gathering and preprocessing data, rather than waiting for someone like Adobe or Autodesk do it.