This is a good release if they're not too cherry picked!
I say this every time it comes up, and it's not as sexy to work on, but in my experiments voice AI is really held back by transcription, not TTS. Unless that's changed recently.
This is a good release if they're not too cherry picked!
I say this every time it comes up, and it's not as sexy to work on, but in my experiments voice AI is really held back by transcription, not TTS. Unless that's changed recently.
(I've yet to experiment with giving the LLM alternate transcriptions or confidence levels, but I bet they could make good use of that too)
Even if you can just mark the text as suspicious I think in an interactive application this would give the LLM enough information to confirm what you were saying when a really critical piece of text is low confidence. The LLM doesn't just know what are the most plausible words and phrases for the user to say, but the LLM can also evaluate if the overall gist is high or low confidence, and if the resulting action is high or low risk.
old ASR systems (even models like Wav2vec) were usually combined with a language model. It wasn't a large language model, those didn't exist at the time, it was usually something based on n-grams.
I put together a script a while back which converts any passed audio file (wav, mp3, etc.), normalizes the audio, passes it to ggerganov whisper for transcription, and then forwards to an LLM to clean the text. I've used it with a pretty high rate of success on some of my very old and poorly recorded voice dictation recordings from over a decade ago.
Public gist in case anyone finds it useful:
https://gist.github.com/scpedicini/455409fe7656d3cca8959c123...
git clone <whisper-diarization.git URL>
cd whisper-diarization
python -m venv .
cd scripts
# and then depending on your OS it's activate.sh, activate.ps1, activate.bat, etc. so on linux [0]
your prompt should change to say(whisper-diarization) <your OS prompt>$
now you can type
cd ..
pip install -c constraints.txt -r requirements.txt
python ./diarize.py --no-stem --suppress_numerals --whisper-model large-v3-turbo --device cuda -a <FILE>
next time you want to use it, you can just do like cd ~/whisper-diarization
scripts/activate.sh (or whatever) [0]
python ./diarize.py [...]
[0]
To activate a Python virtual environment created with venv, use the command source venv/bin/activate
on Linux or macOS, or venv\Scripts\activate
on Windows. This will change your terminal prompt to indicate that the virtual environment is active.(the [0] note was 'AI generated' by DDG, but whatever, linux puts it in ./bin/activate and windows puts it in ./Scripts/activate.ps1 (ideally))
https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2
I use it for real-time chat and generating subtitles. It can do a tv show in less than a minute on a 3090.
Whisper always hallucinated too much for me. It's more useful as a classifier.
> Any of you fucking pricks move and I'll execute every motherfucking last one of you.
I'm so tired of the boring old "miss daisy" demos.
People in the indie TTS community often use the Navy Seals copypasta [1, 2]. It's refreshing to see Resemble using swear words themselves.
They know how this will be used.