Another new experimental codec from xiph.org
xiphmont.dreamwidth.org
xiphmont.dreamwidth.org
My heavens; this sort of technology would enable, say, a robot to have a much better time understanding what a person is saying in a public place (as well as having possibilities for better chatting for people over the internet).
There are still[1] circumstances where one would have noise suppression as a pre-processing step to the speech detector. For instance when using microphone arrays, the frontend may use data from many microphones (2-20), perform adaptive Beamforming extract the cleanest possible mono speech signal. This can include a Vocal Activity Detection (VAD) as estimator, and multi-source aware noise reduction. The output is then ran through a standard mono detector, either on device (typically only Keyword Spotting) or sending to the cloud.
Strong neural networks in combination with microphone arrays is the reason why smart home devices like Alexa etc have become pretty decent (compared to what was feasible 10 years ago) at speech recognition when the speaker is far away.
1. Integrating more and more into functions into big neural networks and jointly optimizing the overall system is definitely a trend, and very actively researched.
Gaming.
What if you wanted to have a seemingly normal game, but when played, you discovered that the characters have seemingly infinite dialogue trees? Tens of thousands of hours of voice performances. What if you could code up a game especially because it became possible to do something like that? What if you had lines delivered in ten different tones of voice and the modulation of that was relevant to gameplay? So it'd seem a little like normal gaming voice acting, except that maybe you'd be getting nonverbal cues and not know why you were reacting differently or getting more tense, except the 'other people' in the game were acting subtly differently.
Yes. Yes there are lots of ways to use this type of thing (also, I'll note that game engines like using Xiph tech already, Godot uses .ogg already)
Absolutely there's a use case for this. Many.
Which is a risk to professional voice actors, but a huge opportunity for indie game developers.
Alternatively, low bandwidth also allows the use of noisier links, since effective bandwidth is the product of the channel width and SNR margin.
The military applications of either case are fairly obvious, but I could also imagine the utility for battery-powered devices (portable baby monitors?) if the processing itself is brought within tolerable power limits.