Show HN: LibreASR – An On-Premises, Streaming Speech Recognition System
github.com
github.com
I've been working on this for a while now. While there are other on-premise solutions using older models such as DeepSpeech [0], I haven't found a deployable project supporting multiple languages using the recent RNN-T Architecture [1].
Please note that this does not achieve SotA performance. Also, I've only trained it on one GPU so there might be room for improvement.
Edit: Don't expect good performance :D this is still in early stage development. I am looking for contributers :)
Another question: could you use for example mozilla voice data to train/test?
Mozilla Common Voice data is already used for training.
I was hoping to use it as a good training base, but the issues I encountered made me wary that the data quality would adversely affect any outcomes.
Relative to what? Paperspace is one of the costlier GPU providers.
For something cheapest I read that post on reddit :
https://amp.reddit.com/r/devops/comments/dqh09n/cheapest_clo...
But you need supervised data too.
'alias downloadmusic='youtube-dl --extract-audio --audio-quality 0 --extract-metadata'
in my .bashrcI find that helps with the annoyance of downloading things off of YT. This is for music obviously, but there's an option to download subtitles as well.
EDIT: Typed this from memory, there may be errors in the alias.
youtube-dl -f bestaudio $URL
Dunno when that went in but it works now.On a side project, I'm looking at the best interface to facilitate further edition (correction) of the recognized text. Target is local councils and regional parliament, where sessions are usually recorded but without transcripts. If xx% accuracy is enough to identify keywords, manual edition is still required to not distort precise meaning.
Nothing special in the interface, but two features seems interesting: 1. Be able to collaborate in real-time. Maybe using Etherpad API to merge multiple editions. 2. Easily validate text and label speakers so as to generate new training data.
Pointers to similar existing solutions would be very appreciated.
I looked into Kaldi and Mozilla Deep Speech but the former seems geared at ASR experts and the latter didn’t seem suited for my particular application (longer recorded audio or real time stream)
I would recommend looking at Vosk too, it converts speech to text much faster than Mozilla DeepSpeech while having slightly better results: https://alphacephei.com/vosk/
premise (noun) a previous statement or proposition from which another is inferred or follows as a conclusion.
premises (noun) a house or building, together with its land and outbuildings, occupied by a business or considered in an official context.
You didn't have any issue understanding the original title.
Humans are very good at live error correction, but that doesn't make it not wrong.
I looked at Twilio but they seem to only offer a means to do it on their VOIP/SIP product.
Yes, Google, Amazon, Microsoft all offer streaming solutions (wouldn't recommend Amazon's however, might recommend Microsoft over Google). wav2letter from FB is the only open-source framework worth looking at, deepspeech is not a seriously usable framework.
What most people don't realize is that it heavily depends on your use case and domain whether any given model/algorithm will work better.
Not sure if relevant though, it's using their SIP product also. If the original service isn't using Telnyx, you could get creative and have a Telnyx shadow user join the group call to receive the stream, etc.
Agora also talk about this, but I haven't used it myself https://www.agora.io/en/
I enjoy a Raspberry Pi, Jetson nano, Arduino, whatever as much as the next person but the seemingly endless stream of projects and resulting blog posts, etc featuring them can get a little old.
Great work!
Thank you for your kind words! :)
[1] https://ai.facebook.com/blog/wav2vec-20-learning-the-structu...
[0] https://github.com/theblackcat102/Online-Speech-Recognition
Do you mind emailing me? Lastname at gmail dot com (see my profile for my name)
Wondering if you might have a better reco of the French President if you have a model per dialect of English?
As a FYI, I was told “the money” was in specific “dictionaries” for medical professionals and so forth. Apparently, doctors liked to dictate straight into text. Might be worth trying that $$$€€€£££?
Thats why there are dozen RNNT projects around, some of them more reasonable, some less, but most of them struggle to demonstrate even a good librispeech WER.
You have a long way to go.
That is to say, what are areas where you think ASR can enable new products or make common and tedious tasks much more efficient?
Also, it's already paying off tremendously. When repercussions can be severe, so can rewards. We have a 2-year old who is incredibly emotionally aware, has a huge vocabulary, enjoys eating anything, explores freely, is learning to play multiple instruments, draws with a pencil grip in both hands, can sit to actively listen to music for 20+ minutes at a time, and learns lyrics incredibly fast.
If you have specific fears, I'm interested in hearing them because that gives us an opportunity to prepare.
I wonder if the drawback will be that he/she will be incredibly bored once released to the "normal" world and the slower development pace of contemporaries and may feel out of place and frustrated. It's the curse of the gifted.
We're working hard to keep from using judgmental/subjective words like good/bad, like/dislike, etc. We're also starting to incorporate Nonviolent Communication patterns and concepts. We use they/them pronouns instead of gendering them. If they want to do something, we strive to help them do it as long as they wont be maimed or killed. We ask them for consent before changing their diaper, touching them, picking them up, taking things from them, and performing medical/dental procedures on them. We've named them Uni Verse All. They wear whatever clothes they choose, no matter what gender they may seem created for. I'm genderfluid and do the same. I also shower once every 1-2 weeks, stopped using shampoo about a year ago and am about to stop using soap on my body, too. We aren't teaching them about property currently and may not ever, choosing to instead describe things as "living with" someone. I'm developing a spirituality with a component I call "radical ignorance," which is essentially a sort of Zen "beginner's mind." It recognizes that ignorance isn't an excuse, but a spiritual reason for doing things, which runs counter to the US's legal reasoning of "ignorance of the law is no excuse for breaking the law." My partner and I are intentionally staying out of the workforce, instead choosing to serve people in our community alongside Uni, which allows both of us to be available so they can have their choice between us. Anytime one of us is choosing to not let them do something without there being a safety issue, the other ideally defaults to helping Uni do what they want. When they get hurt, we bring their attention to the pain and teach them to mindfully experience it while breathing through it.
As for the drawbacks, we're currently designing a community anti-adultist homeschool model that focuses on collaboratively learning our needs over what schools typically teach and allowing the students (of age 0-200+) to choose the contexts for learning. So they'll probably have an interesting intergenerational peer group to blow past the world with.
The primary purpose of the recordings isn't for the sake of analysis, but for documentation. I think the only time we'd probably "go to the tapes" is for reliving what we call "sacred moments" (like last night when they began playing the pump organ in ways similar to me without me doing more than playing and explaining a little bit of the mechanics of the machine) or for when there's a dispute about things. My partner has an automated process of constructing exaggerated narratives and is still learning to notice when it's happening. The videos are more about capturing when we're carrying our own traumas into the relationship and visiting them upon Uni.
If this turns into something where I'm poring over videos, I'm relapsing in my information addiction hard.
Parenting is, ideally, training a new person to know/identify what their needs are and how to meet their needs, including safety, autonomy, exploration, acceptance, and interdependence. Our goal is to only stop them from doing something when they might die or be maimed. And then to get out of the way. Parenting as it's classically been done in many cultures around the world is incredibly adultist, with all kinds of assumptions about what children can and can't do. Even neuroscience and the medical field uses "childhood development" as a reason for ignoring childrens' pains, consent, and autonomy. We don't play that way. Uni is "behind" on their vaccinations due to them not yet saying yes to the ones we're at. We tried respecting their consent for the first few, but the nurses weren't onboard. Even when we had someone setup to receive a flu shot first, the nurse administered it when Uni wasn't looking, keeping them from actually seeing what was happening.
No, if anyone will be analyzing the videos, it won't be us. It'll be other people and algorithms. Any suggestions coming from the analysis will be converted into experiments to conduct with Uni's full informed consent.
Does that make things clearer about what we're trying to do here?
this would give you absolutely perfect sync down to the word, I assume... I don't know about the cost if you paid ratecard though, perhaps you can do some partnership with them since yours is a symetrical product