- my actual voice is terrible and i have a lot of noise in vicinity to sit and process it
- what i want is to be able to talk into the mic and an AI voice is generated in real time with a very large degree of accuracy
- bonus if it works with OBS as a plugin or something
I have been playing with Whisper speech-to-text recently (both CLI, and also . Depending on the model & computer performance, it seems like it can stream (listen and transcribe to text) audio. Maybe you can pipe that to a text-to-speech service (remote or local), and configure that as an audio input to OBS.
Are you live-streaming, or just recording video file to edit/publish later? If it's not live, that may relax some of the latency requirements.
Related links:
- https://github.com/openai/whisper (speech-to-text for terminal)
- https://buzzcaptions.com/ (graphical interface to Whisper)
Somewhat related, Descript mentions it does dubbing. I recently used this tool to edit a video, and I was impressed by how easy it was, and accurate too. It may not fit your realtime use-case, but might be worth exploring...
Good luck!
- what exactly is the architecture of speech to speech?
- is it speech to text and text to speech or am I badly understating the actual mechanism