There’s also the new Portal hardware.
And so FB for instance can send some voice data to their servers and get a text output. And then FB can use text sentiment analysis to get further context about the message.
Sadly, most people don't have the speech data to train their own recognizers on large vocabulary systems, and that's even harder for languages that are not English. With exception of Google/Amazon/FB/Microsoft/Baidu/etc other people have to use the API's offered by the above companies to do high fidelity recognition. Which sucks because there is a cost to each recognition. You have to pay someone else to do it.
Whereas FB/Amazon/MS/Baidu/etc can do high fidelity recognition offline on large vocabulary and offer it as a service. THIS is why FB wants to make speech recognition systems.
Is the implication that offline Android recognition does not train on the owner's voice at all? I imagine a lot of phones these days are at least as powerful as the Pentium 200s used to train (successfully!) Dragon Dictate et al 20+ years ago.
Secondly, when I say "train", it is in a totally different context than how you seem to be using the term. You are using it in the context of adapting an acoustic model to a individual speaker to improve the performance. I am talking about building the initial model. Typical RNN or even convolution based algorithms require a lot of time and processing power to train. What's even harder to get than the processing power though is of course, data to train off of.
Secondly, the trained model itself is very big just for storage, and inference against the model is also resource intensive. This is why Android/Google maps/search/etc go out to the google's backend recognition servers for speech to text before falling back on the shitty (but relatively good) offline inline model (that may not even be using state of the art speech recognition techniques and may be using old school GMM based recognizers).
Finally, the large models trained on the backend servers using their distributed computing infrastructure are extremely more accurate than the shitty fallback model, so speaker dependent adaptations aren't necessary. If you can get very very good performance from a speaker independent model, why would you put the extra effort to make speaker dependent adaptations if the gain is very marginal? Not to mention the fact that speaker independent models are more useful in more situations and are extremely powerful. Google for instance can caption videos automatically using speech recognition, which is amazing. If the models were speaker dependent they wouldn't be able to do that. That's why the focus has been so much towards speaker independent models.
I totally disagree. Compared to Sphinx it is still lightyears better.
To wit I use it for my android based home automation voice recognition and even from a distance with background noise it still works about >90% accuracy. My original tests with Sphinx in a similar environment garnered about 30%.
I wonder if you could bootstrap a sizable speech dataset by trawling audio off YouTube and then using one of the really good cloud speech recognition services to label it. :)
Edit: It's even less [1]:
$0.00 1993 Member
$250.00 Non-Member
$125.00 Reduced-License
[1]: https://catalog.ldc.upenn.edu/LDC93S1There have been much better, larger datasets available for a long time, for example the Fisher English conversational telephone speech corpus was released in 2004 and contains ~1950h of transcribed speech. There are tons of other datasets in various languages and for various applications (conversational speech, broadcast transcription, etc.).
Facebook as research arm dedicated to play Go/Starcraft also, what do you think their reason is for doing that?
They can use a speech recognition system to transcribe videos, just like what Youtube is doing to improve ad targeting and recommendation. Why is this hard to understand?
Is that true? I don't have the stats, but I guess a very small percentage of yt videos were uploaded 'for profit'. Like 1% or less? Maybe much less.
From a software standpoint, this has never been proven. However, its super weird when it happens to you.
For example, I traveled to meet a coworker who was playing a mobile game I had never seen before and we talked about it. I never Googled it or anything like that.
Hours later I checked Instagram and the first ad was for the same mobile game. Coincidence?
Perhaps the game was simply advertised more in his city than my own?
Perhaps our phones being near each other prompted a "friend request suggestion" and then a took that to another level with installed apps?
Or just a coincidence and I am thinking too much about it. lol.
“Ah, but there’s an undetectable binary blob that gets linked in and called without being detected by anyone working on the code,” you say. In that case consider the battery life impact. The power consumption of the Facebook app compares favorably to its social media peers. Is everybody else also recording, encoding and encrypting all the time?
Such as?
Facebook has explicitly denied spying on people's conversations. I can't think of any situations where they have flat out lied about something like that, so I'm curious to hear what conspiracy theories they have proven true.
It’s a complete mischaracterisation that Facebook shared private messages with Spotify. By the same logic, you could say Google is sharing your emails with Apple when you access Gmail using the iOS Mail client.
Spotify was offering an integrated client to FB’s chat service. This UI integration was a market failure and was discontinued years ago. Of all the things wrong with Facebook, this wasn’t worth the noise.
Shadow profiles for example.
Scrolling through a social media feed is a heavy activity on a phone (which may be surprising because it seems passive). Constantly fetching more data from servers, decoding incoming images and videos in background threads, shuffling data to GPU-accessible buffers for fast scrolling, etc. — There’s a lot going on all the time when you’re scrolling mindlessly.
So what? Thousands of the same people have access to the full source code of everything underhanded Facebook do and it hasn’t stopped anything.