I had to leave a lot out of this that I wish could have gone in, simply due to length constraints. In that regard, perhaps "The Ultimate Guide to Speech Recognition" wasn't be best choice of title. I'm sure that we'll be updating this article as time goes on, and Google's streaming API is something I want to make sure goes in it.
Also, something that was left out of the article was SpeechRecognition's listen_in_background method, which does solve this problem somewhat. My issue with it is that SpeechRecognition uses a somewhat crude RMS energy based VAD for detecting speech.
Thanks for your feedback!
I found LTSD pretty robust compared to simpler energy based things as long as you have a small chunk of background sound at the start. The LTSD implementation is largely from my friend Joao, so I can't take credit for the cool part, only the bugs
[0] https://gist.github.com/kastnerkyle/a3661d6be10a0ae9e01fd429...
One of the things I like in HNews is that there's often someone who can join discussion and add something to it.
Clearly it's worth to learn and discuss from experience of others to see broader spectrum. Thanks to the author for looking here and dropping few lines.