High-performance speech recognition with no supervision at all
ai.facebook.com
ai.facebook.com
[1] https://engineering.fb.com/2018/08/31/ai-research/unsupervis...
What if you trained with 960000 hours instead??
(basically instead of hello it says „o l l e h“)
I see they are using a GAN, and doing unsupervised training. But then they appear to compare their model to supervised-trained models.
How do they do this? Do they tack a supervised-trained model onto the end of their unsupervised model? I imagine they must do supervised training at some point, else how can they convert sounds to text?
Also : "The discriminator itself is also a neural network. We train it by feeding it the output of the generator as well as showing it real text from various sources that were phonemized."
Is the "real text from various sources that were phonemized" a manually labelized database? If Yes, that step is supervised, which makes the whole thing actually supervised to some extent
- Understanding for videos for content moderation, ranking and recommendations
- Captioning
- Assistant tech in Portal, Oculus and upcoming glasses
For example, in a video taken at the beach where someone says “wow that looks fun” (in any language) and there is a boat in the frame, you could serve them boat ads.
But go ahead, try it yourself! Have a conversation about having children with your loved one, next to your phones and your TV box, no online search required. Give it few hours and turn your Sling on, or browse some Amazon/Youtube. All of sudden you will see ads for products and companies you have never heard of, trying to sell you diapers or baby cribs. Where do you think it came from? Google, as of today, still is unable to read your mind.
1) we know that sending voice data to "HQ" costs power
2) we know that live transcription costs a huge wedge of power
3) we know that wakeword matching is quite power efficient. (see https://rhasspy.readthedocs.io/en/latest/wake-word/, https://github.com/MycroftAI/mycroft-precise)
So, in a quite room we know that to save power and data, devices won't be streaming data/listening. We can use that as a baseline for power and network usage.
Then we can start talking, measure that
then start saying watchwords and see what happens.
That aside, we know that its really expensive to listen & transcribe 24/7. Its far easier and cheaper to monitor your web activity and surmise your intentions from that. There are quite a few searches/website visits that strongly correlate to life events. Humans are not as unique and random as you may think, especially to a machine with perfect memory.
You not going to distinguish what people want to BUY from Google search as good as from conversation. When I google "Ferrari" it may mean I am looking for Ferrari wallpaper, Ferrari stats, Ferrari parts, or want to buy new Ferrari. When I have a conversation with someone about buying Ferrari, the conclusion IS I am ready to buy a Ferrari.
4A doesn't matter anyway, what matters is GDPR, it's not worth building a system that is not GDPR compliant even if you're outside the EU.
Maybe the computational benefits of such a simple algorithm outweighed the potentially bad clusters. Thoughts? I'll try to read the paper in more depth later in case they explain the choice and I missed it.
What a world!
Did you watch the video tweeted by Facebook's CTO?
https://twitter.com/schrep/status/1395766932104572928
The speed, accent, and word choices of the speaker would throw off previous state-of-the-art tech.
The amazing thing about this new unsupervised algorithm is they can throw unlimited amounts of unlabeled audio and text at a computer and wait for it to become great. It's not AGI but it is still amazing. Language detection is also already solved so there really is no labeling required.
We already have noise cancellation, good tiny batteries, fast wireless internet, and all the tech to learn spoken languages without supervision!
Autonomous driving has a much higher bar for accuracy.
https://pangeanic.com/knowledge/the-worst-translation-mistak...
Language has great ambiguity, depends heavily on time period, context and tone.
https://www.reddit.com/r/funny/comments/2chfge/you_need_some... That could lead to problems even without translation.
(Fwiw, if you read the comments at your second link, you'll find the image is a fake.)
I really don't know just what the hell my phone wants from me anymore.
They train this on examples of speaking and maybe it's more broad now.
Why would a search engine (1997) make glasses (2013)?
It literally was just because Sergey Brin thought it was cool.
Positive stories associated with the brand seem especially needed these days. Perhaps FB keeps other “rainy day” projects as well.
FB pays well and works on really interesting things but I just can't justify the moral price of working there.
I’m thinking language like this, but also labelled imagery for, for example, face detection, which works better on white people.
Has anyone attempted to create a way for people to create and donate labelled data to a dataset?
The idea is that the smart contract is a learning algorithm, and people can donate data to a public repository stored on a blockchain. The learned model is publicly available for everyone. People can also receive incentives for donating data in some implementations.
Good, high quality, wide coverage, labeled datasets are expensive to assemble. Most companies don't want to give them away. You can find a number from academia, though.