Voice2json: Offline speech and intent recognition on Linux
voice2json.org
voice2json.org
It's entirely unpackaged: https://repology.org/projects/?search=voice2json https://pkgs.org/search/?q=voice2json
Docker image is broken, how'd that happen?
$ voice2json --debug train-profile
ImportError: numpy.core.multiarray failed to import
Traceback (most recent call last):
File "/usr/lib/voice2json/.venv/lib/python3.7/site-packages/deepspeech/impl.py", line 14, in swig_import_helper
return importlib.import_module(mname)
File "/usr/lib/python3.7/importlib/__init__.py", line 127, in import_module
return _bootstrap._gcd_import(name[level:], package, level)
File "<frozen importlib._bootstrap>", line 1006, in _gcd_import
File "<frozen importlib._bootstrap>", line 983, in _find_and_load
File "<frozen importlib._bootstrap>", line 967, in _find_and_load_unlocked
File "<frozen importlib._bootstrap>", line 670, in _load_unlocked
File "<frozen importlib._bootstrap>", line 583, in module_from_spec
File "<frozen importlib._bootstrap_external>", line 1043, in create_module
File "<frozen importlib._bootstrap>", line 219, in _call_with_frames_removed
ImportError: numpy.core.multiarray failed to import E: The repository 'http://security.ubuntu.com/ubuntu eoan-security Release' does not have a Release file.
And just bumping the image tag to ":groovy" caused subsequent silliness, so this project is obviously only for folks who enjoy fighting with build systems (and that matches my experience of anything in the world that touches Numpy and friends)Maybe you could convince a couple of indie creators or state-run programs to licence their audio? But I'm not sure if negotiating that is more efficient than just recording a bit more audio, or promoting the project to get more volunteers.
Also I don't think there have been any major court cases about this, so there's no clear precedent in either direction.
Easy fix - keep a bloom filter of hashed ngrams ensuring you don't repeat more than N words from the training set.
It would still be an uphill battle to convince them to hand over the training set but the legal department can likely be convinced if the data set they contribute back is heavily chopped up audio of the original content, especially if they have the originals before mixing. I imagine short audio files without any of the music, sound effects, or visual content are pretty much worthless as far as IP goes.
I'm not sure that is a clear copyright violation. Sure, at a glance it seems like a derivative work, but it may be altered enough that it is not. I believe that collages, and reference guides like cliff notes are both legal.
I think a bigger problem would be that the scripts, and even the closed captioning, rarely match the recorded audio 100%
What might work better is using closed captions or subtitles, but I've also seen enough cases where those don't exactly match the actual speech either.
Thanks for taking time to contribute!
(The first clip I got was spoken more or less correctly, but a couple of words are slurred together and the prosody is awkward. Without having a good idea of the standards and goals of the project, I have no idea whether including this clip would make the overall dataset better or worse. My gut feeling is that it's good for training recognition, and bad for training synthesis.)
This seems to me like a major issue, since it should take a relatively small amount of effort to write up a list of guidelines, and it would be hugely beneficial to establish those guidelines before asking a lot of volunteers to donate their time. I don't find it encouraging that this has been an open issue for four years, with apparently no action except a bunch of bikeshedding: https://github.com/common-voice/common-voice/issues/273
[1] https://f-droid.org/packages/org.commonvoice.saverio/
[2] https://discourse.mozilla.org/t/discussion-of-new-guidelines...
One speaker, who sounded like they were from the mid-west United States, was dropping the S off words in a couple clips. I wasn't sure if it was misreads or some accent I'd never heard.
Another speaker, with a thick accent that sounded European, sounded out all the vowels in circuit. Had I not had the line being read, I don't think I'd have understood the word.
I heard a speaker with an Indian accent who added a preposition to the sentence that was inconsequential but incorrect none the less.
I hear these random prepositions added as flourishes frequently with some Indian coworkers, does anyone know the a reason? It's kind of like how American's interject "Umm..." or drop prepositions (e.g. "Are you done your meal?") and I almost didn't pick up on it. For that matter where did the American habit of dropping prepositions come from? It seems like it's people in the North East primarily.
[¹] If that's even fair given it's a dialect in its own right - Americans also say things differently than I would as a 'Britisher'
I'd say it's mostly subtler (I suppose that should be the expected distribution!) things I've noticed though, they're just harder to recall as a result.
(Just want to emphasise I'm not making fun of anybody or saying anything's wrong, in case it's not clear in text. I'm just enjoying learning Hindi, fairly interested in language generally, and interested/amused to notice these things.)
Hindi is much more economical, to put it literally, one says things like 'than/from/compared to orange, lemon is sour', and 'orange is little/less [without comparison] sour'.
Which, I believe, is what gives rise to InE sentences like 'the salt in this is very less' (it needs more salt, there's very little).
Massive mountains of data tends to be incompatible with opensource projects. Even Mozilla collecting user statistics is pretty controversial. Imagine someone like Mozilla trying to collect hundreds of voice clips from each of tens of millions of users!!
Both of those involve entering data about external things. Asking people to share their own data is another thing entirely—I suspect most people, me included, are much more suspicious about that.
Do user collected clips have soemthing so special to the point that it’s critical to collect them?
They do, and it's working! https://commonvoice.mozilla.org/en
Decent designs are in published papers all over the place, so thats a solved issue.
Lots of compute requires lots of $$$, which isn't opensource-friendly.
Lots of data also isn't really opensource friendly.
Sadly this is a niche that the opensource business model doesn't really fit.
* besides the hard part of standing up an EMR is not installing a prepackaged software.
Not really, look up BOINC.
So companies with a lot of data, then.
A FLOSS system would only have my voice to recognise and I would be willing to spend some time training it. Very different usecase from a massive cloud that should recognise everyone's voice and accent.
It's well-documented and works basically out of box. I wish the STT models bundled were closer to the quality of Kaldi but the ease-of-use has no comparisons.
And maybe with time it will surpass Kaldi in quality too.
CMUSphinx and Julius have been around for ~10+ years at this point.
[EDIT] - there's even a useful Quora post[4]
[0]: https://cmusphinx.github.io/
[1]: https://www.kaldi-asr.org/doc/about.html
[2]: https://github.com/alphacep/vosk-server
[3]: https://github.com/julius-speech/julius
[4]: https://www.quora.com/Are-there-any-open-source-APIs-for-spe...
Example: https://www.youtube.com/watch?v=tfcme7maygw
Granted, maybe this is "not good enough", but I feel like I got pretty far with pico2wave, pocketsphinx plus 1980's Zork level "comprehension" technology.
And the open source status of pico2wave is a bit questionable, I'll grant you that.
A bit more detail about the implementation here: https://scaryreasoner.wordpress.com/2016/05/14/speech-recogn...
The TLDR of this project is: a unified command-line interface to different offline speech recognition projects, with the ability to train your own grammar/intent recognizer in one step.
My apologies for the broken packages; I'll get those fixed shortly. My focus lately has been on Rhasspy (https://github.com/rhasspy/rhasspy), which has a lot of the same ideas but a larger scope (full voice assistant).
Questions, comments, and suggestions are welcomed and appreciated!
voice2json is better suited for limited domain speech, where each sentence is a specific voice command (think home automation).
E.g. maybe "dine" maps to d$ and "chine" to c$. So as in keyboard vim you can guess what "dend" and "chend" do.
this guy is already there: Slurp slap scratch buff yank
My aim has been to train "good enough" models for any public/free data I can get my hands on.
Edit - Sorry, I realize that's a tangent. What I'm saying is that when I was evaluating speech to text engines for things like IVR systems using AWS and Google, neither of them supported SRGS. Microsoft does, I think, but they didn't have a telephony component, and IBM was ignored from the get go, so "no one" really means "two very large companies."
I would have preferred to use a standard. Perhaps this is something for a future version.
Might use this with a Raspberry pi to set up some projects around the house. Is it possible to buy higher quality voice data ?
It's from the same author.
Part of the problem is that language support varies dramatically between components. There's usually a pretty obvious "best" set for English, but it gets more difficult with other languages.
The goal of voice2json is to provide a common layer on top of existing open source engines. This common layer lets you train custom speech/intent models with having to know the details of each engine.
The second was a little surprising (and maybe I missed it?) There was not much in the way of easily accessing transcribed output to and from shell scripts?
voice2txt command.wav | txt2intent
? Or the intent analyzation actually requires the sound data (what are the cases of the same phrase expressing different intent, or how do we even define / categorize intent in this context)>> Supported speech to text systems include:
>> CMU’s pocketsphinx
>> Dan Povey’s Kaldi
>> Mozilla’s DeepSpeech 0.6
>> Kyoto University’s Julius
In case you're not aware, those are all locally run (thus not sending data off, not sacrificing privacy as you mention)