Here's a demo from a startup in this space. Still very early. https://deck.sindarin.tech/
Also my radio-trained voice is so generic a caller every week-ish assumes I am a bot, so I’m pretty sure the problem isn’t me enunciation or accent.
> [INST] Spell out the following word letter by letter: margorczynski [/INST] m - a - r - g - o - r - c - z - y - n - s - k - i
> So, the word "margorczynski" spelled out letter by letter is: m-a-r-g-o-r-c-z-y-n-s-k-i.
The text between `[INST]` and `[/INST]` is the input. The text after `[/INST]` is the output.
https://twitter.com/karpathy/status/1657949234535211009
I'm not arguing that you can't use single chars just that many of the issues parent discussed are caused by this.
Aye.
I was surprised this morning when it decided I had was talking about a "Mark of chain". 1/3rd of the time it hears "bedroom 100%" as "bedroom off".
When cooking dinner today, I asked for a "ten minute timer", it responded "for how long?" then confirmed my "ten minute minute timer".
Still better than Alexa, which kept telling us it couldn't find «kitchen» on Spotify even though we didn't even have Spotify.
And way better than the voice control on Mac OS Classic; back in the late 90s/early 00s, it interpreted 75% of my attempts to use it as "tell me a joke" (it wasn't even a good joke), and ignored 20%.
This demo is pretty bad compared to what we currently have in development.
We’ve been in code freeze in prod for over two months to get our substantially improved engine finished.
It’ll be out in a few weeks, and it’ll blow this version away in every way that matters.
Thanks for checking us out!
The vast majority of America is within 10ms of a data center. That's nothing.
The current challenge for most interaction is ASR -> prompt processing latency. This will be improved with multimodal models on specialized hardware like Groq.
and another 0.6s or so to get first voice chunks from PlayHT
measuring STT latency is harder, I need to implement a local VAD model first to properly measure it, but I think it's on the order of 0.5s
So this has nothing to do with Groq, really. ChatGPT is just slow (too slow for realtime voice communication).
Why do we add junk words while we think? I think it's probably because we're social animals, we want to hold that person's attention as we think as periods of silence are likely to make them become disengaged. But who knows really.