Show HN: State-of-the-art German speech recognition in 284 lines of C++
github.com
github.com
wget "https://github.com/DeutscheKI/tevr-asr-tool/releases/downloa..."
wget "https://github.com/DeutscheKI/tevr-asr-tool/releases/downloa..."
wget "https://github.com/DeutscheKI/tevr-asr-tool/releases/downloa..."
wget "https://github.com/DeutscheKI/tevr-asr-tool/releases/downloa..."
wget "https://github.com/DeutscheKI/tevr-asr-tool/releases/downloa..."
cat tevr_asr_tool-1.0.0-Linux-x86_64.zip.00* > tevr_asr_tool-1.0.0-Linux-x86_64.zip
unzip tevr_asr_tool-1.0.0-Linux-x86_64.zip
sudo dpkg -i tevr_asr_tool-1.0.0-Linux-x86_64.deb
tevr_asr_tool --target_file=test_audio.wav
and then you'll be greeted with some TensorFlow Lite diagnostics, followed by the intermediate states of the beam-search decoder, followed by the hopefully correct transcription result.
And if that piques your curiosity, here's a short overview over the code: https://github.com/DeutscheKI/tevr-asr-tool#how-does-this-wo...
Also, the perplexity of provided ngram LM on CV test set is just 86 and most of 5-gram histories are already in the LM. This also suggests bias.
Anyway, it's roughly a 64% reduction for both wav2vec2 XLS-R and TEVR. So if your criticism that I overtrained the TEVR model turns out to be correct, then that would suggest that the Zimmermeister 2022 wav2vec2 XLS-R was equally over-trained, which would still make it a fair comparison w.r.t. the 16% relative improvement in WER.
Or are you suggesting that all wav2vec2 -derived AI models are strongly overtrained for CommonVoice? Because they seem to do very well on LibriSpeech and GigaSpeech, too.
Could you explain what you mean by "perplexity" here? Can you recommend a paper about it? I haven't read about that in any of the ASR papers I studied, so this sounds like an exciting new technique for me to learn :)
BTW, regardless of the metrics, this is the model that "works for me" in production.
BTW, BTW, it would be really helpful for research if Vosk could also publish a paper. As you can see, PapersWithCode.com currently doesn't list any Vosk WERs for CommonVoice German, despite the website reporting 11.99% for vosk-model-de-0.21.
> Also in Table 6, you see that Facebook's wav2vec 2.0 XLS-R went from 12.06% without LM to 4.38% with 5-gram LM.
It is probably Jonatas Grosman's model, not Facebook. Bias is a common sin for common voice trainers. Partially because they integrate Guttenberg texts into LM, partially because for some languages CV sentences intersect between train and test.
For comparison you can check Nemo model
https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nemo/models...
The improvement from LM is from 6.68 to 6.03 as expected.
> Zimmermeister 2022 wav2vec2 XLS-R was equally over-trained
Yes
> Or are you suggesting that all wav2vec2 -derived AI models are strongly overtrained for CommonVoice? Because they seem to do very well on LibriSpeech and GigaSpeech, too.
Not all the models are overtrained, I mainly complain about German ones. For example Spanish is reasonable:
https://huggingface.co/patrickvonplaten/wav2vec2-large-xlsr-...
> Could you explain what you mean by "perplexity" here? Can you recommend a paper about it? I haven't read about that in any of the ASR papers I studied, so this sounds like an exciting new technique for me to learn :)
Perplexity is a measure of LM quality. See here:
EVALUATION METRICS FOR LANGUAGE MODELS https://www.cs.cmu.edu/~roni/papers/eval-metrics-bntuw-9802....
Also for a recent perplexities of transformers see somethinglike
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context https://aclanthology.org/P19-1285.pdf
> BTW, regardless of the metrics, this is the model that "works for me" in production.
Sure, but it could work even better if you take more generic model.
>BTW, BTW, it would be really helpful for research if Vosk could also publish a paper. As you can see, PapersWithCode.com currently doesn't list any Vosk WERs for CommonVoice German, despite the website reporting 11.99% for vosk-model-de-0.21.
Great idea, we'll get there, thank you!
Thanks for the perplexity paper :) I'll go read that now.
This takes 2 weeks of A100 GPU time to train? Current GC spot pricing puts that at about $300.
Yes, but only for me in this specific case. That's because this German model was derived based on a model which was pre-trained for months on hundreds of thousands of hours of YouTube audio in various languages. So if I now train an English recognizer, I don't start from scratch. Instead, I start from a checkpoint which can already almost flawlessly recognize any human utterance. (Or all phonemes in the IPA alphabet, to be more precise)
The "training" then only learns the mapping between IPA phonemes and text notation.
I should play with this if I have some time. I've had the idea for a while to build a voice assistent which can switch modes or datasets while you are speaking. If you say "computer, play ...." for example it would load a recognizer that is specialized on song names. The idea is that you can mix English song names in a German prompt, and it will not be confused. Every voice assistent I know gets confused, presumably because they convert speech to plain text and only then act on the text.
I would say by now the generic recognizers are so good that this is becoming less and less useful. For example, this tool handles non-existing German words quite well.
That said, the tool has a "--data_folder_path" parameter where you can specify a different acoustic and language model.
BTW, I also want to build an offline voice assistant :)
That's how I got started on this journey. You might be interested in my next project, where I try to do offline real-time English recognition with a WebRTC API to make it easy for developers to connect my AI module with their own task logic. Here's the waiting list: https://madmimi.com/signups/f0da3b13840d40ce9e061cafea6280d5...
If you want real-time speech recognition with less than 0.5s of delay between speaking the word and it being fully recognized, then one needs to implement a different architecture. And that one is much more difficult and expensive to train that this one (which was already expensive).
That said, I want a fully offline and privacy-respecting voice assistant myself.
So attempting to build the AI for real-time streamed live English speech recognition will be my next project. I plan to ship it as an OpenGL-accelerated binary with WebRTC server so that others can easily combine my recognition with their logic. But it probably won't be free since I'm looking at >$100k in compute costs to build it. In any case, here's the waiting list: https://madmimi.com/signups/f0da3b13840d40ce9e061cafea6280d5...
If you don't mind, please email me at moin@deutscheki.de and explain in a bit more detail what the needs of that deaf community are. Maybe I can forward that to the right people to get my government to pay for the app that you wish to see developed.
I mean I agree with you, it certainly would increase life satisfaction for deaf people if they could "listen in" into conversations to know what others are gossiping about.
How about crowd funding it? Your previous work should be enough to convince people it's worth contributing to.
So I believe my best bet might be to partner up with a larger company who will pay for development and/or just charging users for a license. Nuance's Dragon Home is $200 and their Pro version is $500, so there's a lot of room for me to be cheaper while still reaching $100k in revenue with a realistic number of users.
Otherwise, perhaps you could ask LAION on #compute-allocation?
Comments like these are why "Show HN" can be so rewarding (despite all the pedantry about the submission title which I can't change anymore anyway).
Despite some snark - which I totally deserve, but I can't change the title anymore - I also received some very helpful advice, learned something new, and got introduced to people who plan to use this technology to help others. I can't think of any better outcome for me publishing my research.
Thanks :)
People just get confused by large numbers of similar syllables because we have to buffer them for more processing. I suspect a pure speech-to-text model doesn't need to worry so much about context and can just take the syllables one by one.
The result is that this model performs worse for highly repetitive words, just like humans do.
Training this pipeline was already quite expensive so I compared against all models and papers I could find online, but I couldn't afford to train a full new model just to check wav2vec2 with BPE.
That said, I did check against exhaustively allowing all 1-4 character tokens, which is pretty similar to BPE, and that performed worse in every situation.
As such, the AI works well with a variety of accents.
ldd myprogram should say "not a dynamic executable"
Are there any libraries for those of us in embedded?
$ lddtree build/tevr_asr_tool
tevr_asr_tool => build/tevr_asr_tool (interpreter => /lib64/ld-linux-x86-64.so.2)
libdl.so.2 => /lib/x86_64-linux-gnu/libdl.so.2
libm.so.6 => /lib/x86_64-linux-gnu/libm.so.6
librt.so.1 => /lib/x86_64-linux-gnu/librt.so.1
libstdc++.so.6 => /usr/lib/x86_64-linux-gnu/libstdc++.so.6
libgcc_s.so.1 => /lib/x86_64-linux-gnu/libgcc_s.so.1
libpthread.so.0 => /lib/x86_64-linux-gnu/libpthread.so.0
libc.so.6 => /lib/x86_64-linux-gnu/libc.so.6
ld-linux-x86-64.so.2 => /lib64/ld-linux-x86-64.so.2And how many lines can be re-used in a version that recognizes both German and English?
"284 lines of C++" is something you could fit on a small microcontroller. This isn't 284 lines of C++.
That said, I wrote "284 lines of C++" to indicate that this is compact enough for people to actually read and understand the source code. Also, compiling my implementation is super easy and straightforward ... something which can't be said for Kaldi, Vosk, or DeepSpeech.
If you try to read the CTC beam search decoder from Mozilla's DeepSpeech [1], that alone is about 2000 LOC in multiple files. If you try to read the pyctcdecode source that is used by HuggingFace [2], that's 1000+ LOC of Python.
But this implementation is all the client-side, i.e. the entire "native_client" folder hierarchy in DeepSpeech [3], narrowed down to a mere 284 lines.
Also, both DeepSpeech and HuggingFace Transformers use TensorFlow as a dependency, i.e. just like me. So in my opinion, it doesn't make sense to include TF in the LOC comparison if all the AI speech recognition systems use it. That would be like including libstdc++, too.
[1] https://github.com/mozilla/DeepSpeech/tree/master/native_cli...
[2] https://github.com/kensho-technologies/pyctcdecode
[3] https://github.com/mozilla/DeepSpeech/tree/master/native_cli...
Don't take this the wrong way, but I find that people with more knowledge of the subject tend to be more open about what they include, whereas people with less knowledge tend to do more gatekeeping. AI has a "moving goal post" issue that is notable enough to warrant a wikipedia page: https://en.wikipedia.org/wiki/AI_effect
Touting the "low number of lines" on a neural network project seems kind of silly, since the logic is encoded in the weights. Kind of like if I said "Doom in 5 lines of JS", but those 5 lines just downloaded and ran 130kb of WASM
./tevr_asr_tool
Lines of code began and persists as an absolutely awful way to measure anything.
But great hack anyway :)
import jsonThe number 284 means something to people who work in speech recognition - these are the people that know how much _they_ write when they try to compete with this library.
This number isn't meant for people who are disinterested (have no stake) in speech recognition.
Also, the C++ code is mostly a custom beam search decoder based on the research for my paper, so it's not like TensorFlow is doing all the heavy lifting here, because precisely that TEVR token decoder causes the relative 16% performance improvement of this speech recognition AI over others.
I wouldn't say the headline is "misleading" - no one is about to be fooled into thinking 300 lines of C++ could be capable of state-of-the-art speech recognition. The headline is squarely in the territory of "complete nonsense, but you can tell that without having to read further than the headline".
Anything you used in the process of generating it will be input that does count.
I posit this can be done in less than 284 lines of C++ while having an error rate equal to or better than the state-of-the-art for everyday speech.
Gentlemen, ready your putters…
As for the LOC count and excluding TensorFlow, I tried to explain my rationale here: https://news.ycombinator.com/item?id=32411566
Instead, to me, this reads, "Hey C++ fans, you don't need Python for nice things. Look at what you can do in 284 lines of code!"
I don't have C++ chops, so this is nontrivial for me. I appreciate OP sharing this!
That would be different at 2840 LOC or very very different at 284000 LOC.
By not making false assertions?
Judging by the comments the 284 lines is in addition to many hours of GPU training time plus some magic and a huge library. I didn’t even click the link because I knew it was 284 lines of plumbing code on top of something else.
I once wrote a entire renderman compatible renderer in a couple hundred lines of python — which was really a super dumb script which generated ctypes bindings from the header for an actual renderer (pixie if anyone is curious).
That's why I wrote "State-of-the-Art" at the beginning of the title. Because this is based on new research and it works better than previous research.
For example, you can get 90% of the way to a CSV parser in Python in one line (as `[line.split(",") for line in open("some.csv").readlines()]`). Should we consider it a false assertion to call that a "one liner" and if so, how should it properly be described?
I actually wrote a generator for a SVG parser/writer and not counting the xml library it came in at around close to 40k lines of code. Or maybe 70k lines, don’t recall just know it is a lot. If I split it up into one class per file it took something like 45 minutes to compile (it is a C-API extension) so I would dump all the code into a single file for faster compiles.
Haven’t looked at it in a while but the generator file is probably around 400 something lines. Certainly not going to claim its a validating SVG library in 400 odd lines of code.
Think I had some cockamamie scheme to make a SVG to grease pencil converter for blender and only got around to shaving that one yak.
Also, it's 3 characters too long to be a valid HN title.
And lastly, this does contain the parameters for a new AI model which is based on my research, so it's not all "glue logic" ;)
You cannot "outsource" the "core"?
And the size is relevant to people in the industry because DeepSpeech uses 2000+ LOCs for implementing their decoding, so this works better and is 10x less code.
Then you should have mentioned that in the title rather than the LOCs
It's also how you present it. If you said: "Building a TensorFlow backed German speech recognition system in only a few hundred lines of code." Then I feel like you're being more honest.
I’m inclined to accept that, too, but it is a moving goalpost and unfair when comparing between OSes and/or time periods. For example, consumer OSes have been shipping with speech recognition libraries for decades, and they’re getting better and better.
Mozilla's DeepSpeech is so large that you can't really read it and understand it. This one is 10x less code while recognition quality is 75% better (lower relative word error rate).
So this one is small enough that you can read the source code if you want to, while DeepSpeech is not.
From this perspective I can definitely get behind advertising the project with the LoC measurement. Subjectively, I still find this to be a bit of a "not telling the whole truth", however I also only ever toyed around with speech recognition ai.
I'm not saying this isn't an interesting project, but perhaps lead with something different?
-
ps: I skimmed the paper cited and what I wrote above is -not- correct. The project is not simply assembling a pipeline, it is claiming ('papers with code') an innovation:
"This paper presents TEVR, a speech recognition model designed to minimize the variation in token entropy w.r.t. to the language model. This takes advantage of the fact that if the language model will reliably and accurately predict a token anyway, then the acoustic model doesn't need to be accurate in recognizing it."
...
"We have shown that when combined with an appropriately tuned language model, the TEVR-enhanced model outperforms the best German automated speech recognition result from literature by a relative 44.85% reduction in word error rate. It also outperforms the best self-reported community model by a relative 16.89% reduction in word error rate."
The main innovation is that we prevent loss mis-allocation during training of the acoustic AI model by pre-weighting things with the loss of the language model. Or in short:
We don't train what you don't need to hear
If you want to play around with the TEVR tokenizer design, here's the source for that: https://huggingface.co/fxtentacle/tevr-token-entropy-predict...
> We don't train what you don't need to hear
This does sound a lot more interesting than the ~280 lines of code.
For people like my industry clients, on the other hand, "code that is easy to audit and easy to install" is a core feature. They don't care about the research, they just want to make audio files search-able.
What's the German equivalent of "How to Wreck a Nice Beach"?
(Alles okay. Kann noch fahren. Wiedersehen!)
I think "X in Y LoC" should be limited to where Y is the LoC to do the work, not the LoC to setup/interface with some other library. We're getting ever closer to "SQL DB in only 200 LoC" that simply forwards to sqlite or some such.
That's a custom-designed beam search decoder implemented in C++ and based on the research for my TEVR paper. It increases performance by a relative 16% reduction in word error rate.
Also, yeah, it's not 284 lines of C++ by any stretch of the imagination, but the main file is pretty readable, which is nice.
Yes, I didn't count my 3 dependencies KenLM, Wave, or TensorFlow, because those are used by pretty much all speech recognition projects. For comparing the complexity of my code to Mozilla's DeepSpeech, it makes sense to ignore the LOCs for shared dependencies.
> includes the entirety of tensorflow as a submodule
#include "absl/flags/parse.h"
#include "absl/flags/flag.h"
#include "absl/flags/usage.h"
#include "absl/flags/internal/commandlineflag.h"
#include "absl/flags/internal/private_handle_accessor.h"
#include "absl/flags/reflection.h"
#include "absl/flags/usage_config.h"
#include "absl/memory/memory.h"
#include "absl/strings/match.h"
#include "absl/strings/str_cat.h"
#include "absl/strings/string_view.h"
#include "tensorflow/lite/c/common.h"
#include "tensorflow/lite/delegates/hexagon/hexagon_delegate.h"
#include "tensorflow/lite/interpreter.h"
#include "tensorflow/lite/interpreter_builder.h"
#include "tensorflow/lite/kernels/kernel_util.h"
#include "tensorflow/lite/kernels/register.h"
#include "tensorflow/lite/model_builder.h"
#include "tensorflow/lite/testing/util.h"
#include "tensorflow/lite/tools/benchmark/benchmark_utils.h"
#include "tensorflow/lite/interpreter.h"
#include "tensorflow/lite/kernels/register.h"
#include "tensorflow/lite/model.h"
#include "tensorflow/lite/optional_debug_tools.h"
#include "kenlm/lm/ngram_query.hh"
#include "wave/file.h"
#include "tensorflow/lite/minimal_logging.h"