Microsoft demos breakthrough in real-time translated conversations
blogs.technet.com
blogs.technet.com
From that point of view the demo was misleading - we heard an English speaker talking naturally and heard clear responses, but the other side of the conversation was quite broken, although possible to understand.
If both sides were speaking in real world settings the demo would have been honest but far less impressive to casual watchers.
Despite this misdirection I am still impressed with the amount it did manage to translate from the long, naturally-spoken English sentences.
I hope they'll make a public beta soon, so that we can all try and see how this works in practice in the real world.
I tried an early version of the speech-to-text component (1 or 2 years ago?) myself using an off-the-shelf microphone with a model that was completely generic (none of me speaking in the corpus). It worked surprisingly very well.
If I'm correct this is all using the deep neural nets matching on phonemes and I'm really wondering how I could use this if I wanted to.
Is it something they license? Is it something that's perhaps not patented and would end up in CMU sphinx eventually?
They would have selected the best possible language pair for this demo, so we should expect it to have done very poorly if they picked, say, English<->Mandarin.
What stands out is that they didn't pick Spanish as the other language.
It's pretty universally agreed that Spanish[1] is the easiest language for English speakers to learn, and Portuguese is in the same ballpark[2], but German is significantly harder[3]. (And Russian, Chinese, and Arabic would be way to the right on an exponential graph.)
I'm guessing that machine translation of English<->German, for some reason, must be easier than English<->Spanish.
[1] There's a fairly authoritative study on this which I can't find it immediately.
[2] Just from personal experience I find that English<->Portuguese with Google Translate is astonishingly good (in either direction): http://brazilsense.com/index.php?title=Getting_by_with_just_...
[3] The difficulty of German vs Spanish is confirmed by an NSA (!) document that says that "Next to Vietnamese, German may be the most difficult for English-speaking students to learn for German has a difficult syntactical feature, the discontinuity of the predicate, which the others lack. Among French, Italian, and Spanish, there also seems to be only a slight difference in difficulty. It appears that these three are the easiest languages for English-speaking students to learn": http://www.nsa.gov/public_info/_files/cryptologic_spectrum/f...
> Among Vietnamese, German, French, Italian and Spanish, Vietnamese may be the most difficult... Next to Vietnamese, German may be the most difficult
This doesn't amount to being significantly harder, particularly in light of the statement (that you quoted) that there is just one feature that makes it harder than the other three.
Speaking from personal experience (native English speaker, no prior difference in exposure to the two languages, simultaneous study of the two, similar teacher quality and curriculum), I found French harder than German. Although my single anecdote doesn't prove German to be easier, surely it suggests that the one I found easier couldn't be significantly harder.
http://blogs.technet.com/b/next/archive/2012/11/08/microsoft...
*Well relatively, it would be super lame when compared to the kind of tech one would find in a child's toy aboard a starship.
It might be up there, but I have also heard things like Indonesian, which has the simplest grammatical structure. Chinese, speaking of morpho-syntax, is some ways as easy or difficult as English (tense, gender, and number are not more difficult than English, in my opinion, having studied Arabic to fluency and Chinese at the beginner level).
To get back on topic, topologies like this are good for focusing on which specific constructs will cause difficulty, but which is easiest to learn.
I imagine the vocabulary isn't as difficult for machines to process as grammar.
English shares no ancestry with French. However, in 1066 England was invaded by the Normans, leading to the entire aristocracy and upper classes speaking Norman French, causing a lot of French vocabulary to enter the English Language. Most of these words are for stuff in higher registers, though. The basic vocabulary in English is entirely Germanic (it's nigh-on impossible to write a sentence with only French words in English), while much of the more advanced or formal stuff is French (or Latin or Greek).
Afrikaans, Danish, Dutch, French, Italian, Norwegian, Portuguese, Romanian, Spanish, Swedish
[1] http://www.state.gov/m/fsi/
[2] http://web.archive.org/web/20071014005901/http://www.nvtc.go...
So what do you get in the end? This unimpressive video to native speakers and skeptical one to non-natives.
The fact that at the start of the demo, she asks him if he knows german and he said no should be a red flag for the impending bad demo.
- It's relatively easy to learn enough French to carry on a slow, clearly-enunciated conversation with a highly-cooperative speaker in a quiet room. Call it 350 hours of study, or B1 on the CEFRL scale: http://en.wikipedia.org/wiki/Common_European_Framework_of_Re...
- If you want to watch the French dub of [i]The Game of Thrones[/i] in a noisy room, and actually follow the plot twists, it's whole different game. Call it 1,000 to 2,000 hours of study and exposure, perhaps more for some people. If you want to listen to standup comedy, it's usually even worse.
In other words, it's surprisingly easy to establish basic communication, but you can break your heart trying to get really good.
Going by this experience, and by lots of experiments with Google Translate, I would predict that machine translation will suffer from similar challenges: It will be easy to get the point across if everybody cooperates and speaks clearly. But it will be a long time before you can speak idiomatically and casually, at full speed, and pretend the translation system isn't there.
Anyway, obviously there were a couple of mistakes by the speech recognition, and they were taking care to speak very clearly, but still, I'm impressed. The future looks bright.
I am guessing this is going to be the part of our future where language barrier slowly becomes the thing of the past. The most intriguing part of this demo was where Satya says that Machine learning technology gets better at previous languages as new languages are introduced. That in itself is an accidental revolution.
The translation was no better than, say Google Translate. A mashup of Siri + Google Translate wouldn't have been any worse, and both technologies exist for some years. This is, unfortunately, hardly a breakthrough. Rather a 24h API hackathon project ;)
I have seen those already and they don't work nearly as well as this one. I guess the lady was trying really hard to do not screw up the demo and that might actually reduce the wow effect for German speakers.
In other words, this has come a long way since "Dear aunt, let's set so double the killer delete select all", but I've seen enough "automated translation breakthroughs" to remain highly sceptical.
I feel like they should merge the two. The subtitles are a nice touch, but if it could capture the voice of the speaker it would be truly magical.
What's going on there? Someone must know, no?
Now, both Bing Translator and Google Translation use a technology called statistical machine translation: http://en.wikipedia.org/wiki/Statistical_machine_translation
It was about a girl that answers a call in english, talking to an arabic speaking client of her father.