Bringing Google Live Transcribe's Speech Engine to Everyone
opensource.googleblog.com
opensource.googleblog.com
I'm not even saying it needs to name the people in the meeting. Just understand, contextually, if it is from "person 1" or "person 2." Then as it records associate it with that name.
Maybe this can help? But Google's existing APIs might be able to do this.
[1] https://arxiv.org/pdf/1603.03185.pdf (2016)
[2] https://arxiv.org/pdf/1811.06621.pdf (2018)
The models of Mozilla's DeepSpeech STT engine take 1.8 GB in compressed form: https://github.com/mozilla/DeepSpeech/releases/tag/v0.5.1
The main cause for the large size is the language model. They tried using different (smaller) language models, but they weren't as good.
It doesn't work very well in my experience.
So I decided it would be a good idea to open an issue in the linked repo, to find out what the costs would look like.
Turns out someone else already did that! https://github.com/google/live-transcribe-speech-engine/issu...
They have a fairly generous free plan.