It doesn't reach the frontier in either latency or accuracy for ai multilingual conversations.
226 karma · joined January 27, 2014
It doesn't reach the frontier in either latency or accuracy for ai multilingual conversations.
1. The response latency / patience can be configured in the quick settings (tap the gear during the call).
2. There already is a bias in transcription. It's something we're actively improving, multilingual transcription isn't a solved problem yet
3. It is supposed to correct you actively if you have corrections enabled (check the settings). But we received feedback that it's not enough, so we will ramp it up soon
There are incorrect reading or Chinese readings occasionally, but you can tell when that happens due to the furigana being different
We are aiming to create a long term companion/tutor that gets to know you more and more, and can create customized curriculums, lessons, etc.
Also, having the AI voice tutor as our main feature allows us to iterate quickly, and be well positioned for future improvements in AI models.
As for marketing and GTM, we're in the super early stages, and there is definitely a lot of competition out there, it won't be easy.
Intelligence is super key here, especially as the context size gets larger (due to memory) and intelligence degrades.
Another major issue is TTS voice quality, but this seems to be improving a lot for small local models.
EDIT: You're right, latency is also a big deal. You need to get each piece under a second, and the LLM part would be especially slow on mobile devices.
We focused on testing and tweaking the most popular ones, we have not tested some of the niche ones. We have removed languages that users have told us have major issues, but there are still some left.
The voices are due to the quality of the TTS services that we use. Openi, 11labs, minimax. Some services don't have many or even 1 good voice. We will add more over time
Sesame also passes in the users voice into the TTS model so that it can vibe well with the users tone and mood, whereas we are just using raw TTS. Their latency is also very low, but this is not quite suitable for language learning.
In the future we hope to move to full voice to voice models, once those become mature and intelligent enough.
Good idea on the export, will add that to our to-do list.
For pitch accent, shadowing is a great way to improve. You can pause and repeat the tutors messages for example, or read out the word when doing flashcard reviews (copying the flashcard audio).
The models are improving though, and they are at a very good place for English at the moment. I expect by next year we will switch over to full voice to voice models.
Plus, you can't do auto translation for languages like Japanese where the grammar is reversed. Auto translation has fundamental limitations.
You do need an app to create a holistic learning experience just for language learning. Customized curriculum, tons of prompting, AI models chosen for transcription accuracy, flashcards/dictionary, etc.
We also support hands free mode, and many other things are customizable like slang, speaking speed, target language usage, etc.
This is available for all proficiencies. It's just much harder to talk for hours in a new language as a beginner. It's usable but requires more effort.
- Don't go over 10k tokens in the prompts as the intelligence and memory degrades
- Summarize sessions and save the summaries, potentially summarize the summaries as well
- use VAPI or realtime api if you want to build fast. Building the full pipeline takes a while
- try out different models and see how personality varies. Our favorite is gpt4.1 with temperature 1.
- goal system. The promot should always contain the current goal, and the next goal. Evaluate goals with another LLM, and dynamically change the prompt
* curriculum, completely customizable, with grammar, roleplay, topics, speaking speech, transcript, dictionary, corrections, etc
* prompting and AI models all chosen to be a better fit for multilingual, easy to understand, etc.
* the tutor actively tries to teach you, it's not an assistant
* integrated flashcards that go hand in hand with the speaking immersion