QuiLLMan: Voice chat with Vicuna-13B
github.com
github.com
https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_...
Compare it yourself.
It does all sorts of things better like upgrading words into proper nouns, and doing it based on context.
Honestly it's the first time speech recognition finally feels "production ready" after years of Siri/Alexa.
I tried a lot of tweaks but this one is definitely the best I've seen.
How did it know how to spell my name correctly when I just spoke into my microphone??? The two L's usually trips up the transcription models. What????
if the variations are pronounced the same? luck, probably.
This is the first time I've ever seen it get the spelling actually right (missing the accent is forgivable) ;)
How is tiny.en so damned accurate?
How much faster does this run natively?
Do the NEON instructions work for arm devices like an RPI or is it just tuned for Apple?
With this and Alpaca 13B you could probably replace an entire window manager.
Edit: it seems like it was less than a year ago that local speech recognition was a slow slog through a swamp of crufty, complicated research projects with most the high quality training data hidden behind walled gardens. Now I can stutter-step over a word and a demo in my browser correctly transcribes what I meant sans stutter. What happened?
Check out textgen, it has voice in/out, graphics in/out, memory plugin, api, plugins, etc, all running locally.
Do you know how to get this working? I looked through the read-me and didn't see any options for it.
You need to enable the extensions.
I only did voice out locally with silero_tts, it also supports voice out with eleven labs api.
Voice input is via whisper tts.
What’s the state of container based ML deployments?
Can I take a container orchestration if these services and just put them on a vps w a GPU and run this?
Is there secret or just special sauce in ML infra?
https://github.com/coqui-ai/TTS
Support for it was recently added to vocode:
Demo at: https://www.youtube.com/watch?v=OmQup3kst5s
Signup at: https://signup.bondsynth.ai
What I really want is a program to waste the time of phone calls making unsolicited sales pitches.
It would do voice to text, run a simple language model to generate responses, then synthesize the voice back. It doesn't need to be a sophisticated model, not much more sophisticated than the classic "Eliza" program. A few years back someone did this with a canned loop of vague responses and it fooled the sales people for surprisingly long:
https://www.youtube.com/watch?v=XSoOrlh5i1k
It seems like it could all run locally for low latency. Probably the most important part to get right would be a TTS system that isn't immediately pegged as a robot.
It was a bit laggy, but for a free demo from an open source project, I should be the one being shamed!
Well done.
It’s a cost-saving measure that costs you mortal wounds in integration expense, and eventually building out a real infrastructure plan.
I presumed the same as the person you're replying to.
As Michael Bolton said in Office Space: "Why should I change my name? He's the one who sucks."
this is distinct from "client side" where no servers are involved, aside from maybe something that serves a static web page.