Show HN: Python library to add voiceovers to Manim videos programmatically
osolmaz.github.io
osolmaz.github.io
I wasn't, but it's not that hard to integrate a new TTS service. You can copy e.g. AzureService and adapt it to work with AWS:
https://voiceover.manim.community/en/stable/services.html
> How do you approach aligning the words with animations? Is it possible to align it for a specific word?
Indeed, the feature is called "bookmarks". You can see a demo here:
https://voiceover.manim.community/en/stable/quickstart.html#...
> Did you already describe it somewhere?
I did not do a detailed write-up yet, but you can search the repo for "word_boundaries" and "TimeInterpolator". Services like Azure return timestamps for the beginning of each word, and for those that don't return, I integrated Whisper to generate them from the audio. Then, it's a matter of mapping the string indices to audio time via some sort of interpolation (I used linear).
Manim is the math animation library created by the awesome math YouTuber 3blue1brown. This is about Manim Voiceover, a plugin for Manim that provides an API for adding voiceovers to videos. My goal is to create efficient text2video pipelines and make it possible to automatically generate beautiful explainer videos from any educational text on the web.
This is kind of an announcement post. You can check out the documentation itself here: https://voiceover.manim.community/en/stable/
Curious to hear what you think!