Thank you very much!
Yes. You cannot control the video playback on the demo page. I made it so because I wanted a way to showcase how you can switch between languages. You can go from Elon speaking English to German, Russian and Chinese, each with just one click. Activating the player controls would have made the UI more complex and distracting. And it would have also made it harder for me to sync the timing between languages.
Of course, the real output would be a proper player, with all of the controls. Or, for creators, raw files (video and/or audio, plus SRT subtitles).
I also noticed problems in the translation of the Chinese video. I put it up there anyway, because I figured most people coming to my site would be English speakers, and being able to understand a Chinese video might be another interesting aspect, in addition to the idea of being able to turn your own English content into languages you don't speak.
If this had been a pitch deck, I would have cherry picked the samples. But I wanted to share where the project is right now and see if anyone was interested. Premature optimization is the root evil of all programming. I think Knuth said that. And it's a trap I regularly fall into. So I tried to be disciplined this time.
But if any Chinese YouTuber would ask me to dub their work today, I'd make darn sure that the translations were close to perfect. Meaning I'd allow the system to make changes to the way things are phrased if that's necessary for the purpose of timing or cultural context. But I wouldn't allow it to skip a thought from the original video, or say something something different.
I've translated books by hand in the past. So this is something I care about. If the demo isn't perfect in this regard, it's because I didn't know if anyone was going to even look at my project. When I first posted this yesterday, my submissions didn't go beyond one comment for several hours. I already thought I had built another solution looking for a problem. :)
If you're seeing dropped phrases, that's most likely because my arranging function failed. Basically, the translation ran longer than the original. The algorithm tried to speed it up and fit it in. But it failed and dropped it. Better handling of these overruns are on my to-do list. Neither drops nor speedups should be tolerated.
In terms of self-correction, I plan to feed the translated audio back into the transcription engine. Then, an LLM can compare the translation with the original transcript. If anything is missing, the pipeline will be force to run again with slightly different parameters. There shouldn't be a human neccessary in the loop. Translation is what Transformers are best at.