Microsoft turns spoken English into spoken Mandarin – in the same voice
thenextweb.com
thenextweb.com
The economics of the issue is that a machine interpreter just has to be as good as a human interpreter at the same cost. That's a reachable target with today's computer technology. EVERY time I've heard someone else interpreting English or Chinese into the other language, I have heard mistakes, and I am chagrined to remember mistakes that I made over the years. We can't count on error-free machine interpretation between any pair of languages (human language is too ambiguous in many daily life cases for that), but if companies develop tested, validated software solutions for consecutive interpreting (what I usually did, and what is shown in the video) or simultaneous interpreting (the harder kind of interpreting in demand at the United Nations, where even in the best case it is not always done well), then those companies will be able to displace a lot of human professionals who rely on their language ability to make a living.
Right now a lot of interpreters in the United States make a lot of part-time income from gigs that involve suddenly getting telephone calls and joining in to interpret a telephone conversation in two languages. This is often necessary, for example, for physician interviews of patients in emergency rooms or pharmacist consultations with patients buying prescribed drugs (where I last saw a posted notice on how to access such an interpretation service). The IBM Watson project is already targeted at becoming an expert system for medical diagnosis, and patient care markets will surely provide a lot of income for further development of software interpretation between human languages.
It's still good for human beings to spend the time and effort to learn another human language (as so many HN participants have by learning English as a second language). That's a broadening experience and an intellectual delight. But just as riding horses is more a form of recreation these days than a basis for being employed, so too speaking another language will be a declining factor in seeking employment in the next decade.
Human translators are so expensive today that they are only used in situations where the translation has to be correct -- diplomacy, courtrooms, books, etc. Until a machine is much better than a human, these use cases won't switch to machine translation (similarly, self-driving cars won't be allowed until they are proven to be much safer than human drivers).
On the other hand, there's a large casual market for machine translations today for situations like reading foreign Web sites, chatting with people in different countries, reading Tweets in a different language, etc.
After watching this video, I'm fairly confident that a large part of the interpreting that I did could already be handled by this technology.
The idea is to teach people language at the same time as providing a real time translation service. Apparently if you multi-plex novices (and not at a bad rate) you get expert translation at a similar accuracy. The translators benefit by learning language, and the service is self supporting by proving translation.
He did an excellent TED talk on the subject:
http://www.ted.com/talks/luis_von_ahn_massive_scale_online_c...
A local entrepreneur is having some success with a remote system for interpreters that tries to replace the interpreter console and related expensive interpretation equipment:
Not sure how his has handled the strict no-lag requirements though, but they were doing trial runs in Washington and San Diego.
Professional translators and trained interpreters yes.
But thousands of people work solely (or assume the role periodically) as interpreters and translators for many more situations, mainly revolving around business.
Now, signing a business deal will still involve translator and a trained lawyer, but those other everyday cases, including showing a western partner around the Chinese offices, could switch to machine translation.
I disagree. It's only broadening in the additional people it allows you to commune with. Other than that, it's a waste of time.
Having to convert between languages (I'm a native speaker of English who lives in Germany) all the time is huge overhead, sort of like if every country had its own system of measurement, except that the overhead is incurred much more often, not just for measuring things.
There are subtler benefits - I speak three languages natively, and the part I love most about that is also the part that frustrates me most. Afrikaans is a modern language and has a relatively small vocabulary. Insulting someone therefore consists of creatively stringing together colourful combinations of everyday words, mixing in the odd English, Malaysian or Zulu word, and then spoken with a religious fervour that cracks you up. It's side-splittingly funny. What frustrates me is that no other language can do that, so I can't share that with my girlfriend, who is French.
German has similar mannerisms about it. It's naturally dry in a way that English, for all it's adjectives, can never hope to be.
To give you another perspective, I've always had a few deal-breakers I look for when meeting a girl I might otherwise be interested in. First, she needs to speak more than one language. Seconds, she needs to have spent a significant amount of time outside her own country. She must either be able to ski or snowboard. Last, she needs to be able to play a musical instrument. Beyond that I couldn't care less about race, job, creed, colour or any other bigoted perspective.
To know another language is to know another world.
In 1984, newspeak was all about removing ways to describe things the Party didn't want the public thinking about, so that in several generations of the language nobody would be able to comprehend dissent. And it actually has some real life foundations.
Just think of all the times a foreign language speaker says "I can't accurately represent this in whatever other tongue".
Note, I only know English, and after 6 years of Spanish throughout high school and college I have forgotten every wink of it. You really need to immerse yourself in a language to really dedicate it to memory, and constantly use it. Which is really a ton of overhead, that isn't really necessary, I agree.
German, for example, has no word for "silly", while English has no word for "gemütlichkeit".
Eliminating the overhead is worth any of these sacrifices, though.
Language structures our thoughts. It is claimed that Chinese are better at math because of their language structure. How you express things in a different language will help you more clearly see the concept in your own language. I learned more about English by learning German than I did in all my schooling. It improved my spelling as well because it made me understand the words I was using better.
http://fora.tv/2010/10/26/Lera_Boroditsky_How_Language_Shape...
It's VERY good. But skip the first 3 minutes of boring intro.
For broadening experience, see: http://en.wikipedia.org/wiki/Linguistic_relativity http://www.wired.com/wiredscience/2012/04/language-and-bias/
And for intellectual delight, see: http://www.amazon.com/Le-Ton-Beau-De-Marot/dp/0465086454 http://books.google.gr/books/about/Experiences_in_Translatio...
Unfortunately, even though it was posted three times to HN http://www.hnsearch.com/search#request/all&q=sex+machine... it never made the fron page.
Here is the my summary and comment: " Great talk. I don't know much about artificial neural networks (ANN) and even less about natural ones, but I have the feeling that I learnt a lot from this video.
If I understand correct, Hinton uses so many artificial neurons compared to the amount of learning data, that you would usually see an overfitting effect. However, his ANN's randomly shut of a substantial part (~50%) of the neurons during each learning iteration. He calls this "dropout". Therefore, a single ANN represents many different models. Most models never get trained, but they exist in the ANN, because they share their weights with the trained models. This learning method avoids over specializing and therefore improves robustness with respect to new data but it also allows for arbitrary combination of different models which tremendously enlarges the pool of testable models.
When using or testing these ANNs you also "dropout" neurons during every prediction. Practically, every rerun predicts a different result by using a different model. Afterwards, these results are averaged. The more results, the higher the chance, that the classification is correct.
Hinton argues, that our brains work in a similar way. This explains among other things a) Why are neurons firing in a random manner? It's an equivalent implementation to his "dropout" where only a part of the neurons is used at any given time. b) Why does spending more time on a decision improve the likely hood of success? Even though there might be more at work, his theory alone is able to explain the effect. The longer you think, the more models you test, simply by rerunning the prediction. The more such predictions the higher the chance, that the average prediction is correct.
To me, the latter also explains in an intuitive way, why the "wisdom of the crowds" works well when predicting events that many people have an, halfway sophisticated, understanding of. Examples are betting on sport events or movies box office success. As far as I know, no single expert beats the "wisdom of the crowd" in such cases.
What I would like to know is, how many, random model based predictions do you need until the improvement rate becomes insignificant? In other words, would humans act much smarter if they could afford more time to think about decisions? Put another way, does the "wisdom of the crowd" effect stem from the larger amount of combined neurons and the diversity of the available models that follows, or from the larger amount of predictions that are used to compute the average? How much less effective would the crowd be, if less people make more ("e.g. top 5") predictions or if the crowd was made up of few individuals which are cloned?
If the limiting factor for humans is the time to predict based on many different models and not the amount of neurons we have, this would have interesting implications. Once, a single computer would have sufficient complexity to compete with the human brain, you could merely build more of these computers and average there opinions to arrive at better conclusions that any human could [1]. Computers wouldn't be just faster than humans, they would be much smarter, too.
[1] I'm talking about brain like ANN implementations here. Obviously, we already use specialized software to predict complex events like weather, better than any single human could. But these are not general purpose machines. "
Did I about cover all the problems with groupthink and hivemind?
It's a shame though, but I hang out in new a lot, and I urge other HN'ers to hang out there too.
Upvote only interesting things
https://class.coursera.org/neuralnets-2012-001
He talks about a surprising amount of cutting edge achievements being made by deep neural networks just over the last few months.
I am a Ruby guy and I only marginally get in contact with their C++ code, but from what I learned so far this stuff is extremely memory and CPU hungry. It also depends on having been fed the right amounts of input. That's why Google Translate is so good. They have tons and tons of data from all the websites they parse, and in many cases the content can be obtained in different languages. Corporate pages are often translated paragraph by paragraph by humans which results in perfect raw data to train these algorithms. Also for example all documents that the European Parliament produces are translated into the languages of all member states.
Everything that has to do with translation has to do with context. I think the software right now is as smart as a six year old kid, except that it has a much bigger vocabulary. But if you say "The process has stalled. Let's kill it." it probably only makes sense if you know you are talking about computers.
It's hard to imagine that computers one day might really understand everything we say. But just by using Google Translate I think they really might. Это является удивительным. (I don't speak Russian. I hope I didn't insult anyone now. ;))
Actually this may be one of the reasons why Google's Japanese translations are so terrible. The why isn't really relevant here[0] (perhaps you already know anyway) but it there are times when the raw data becomes the most misleading.
[0] Obviously I still mean those actually translating by hand, not the companies which just throw all of their material into Google Translate and consider it a finely proofed document. There are plenty of the latter which makes for an amusing loop in the system.
I also disagree that having someone doing the translations on the webpage is a good source. One only needs to find Asian companies with direct translations to see how bad the human translations are. Just as those companies generally don't have good English writers, I think it is likely that American companies probably don't have outstanding Japanese / Chinese / et.al. writers.
A simple but famous case is "Watashi wa hamburger desu". The sentence has no subject, so with no context that would get translated as "I am a hamburger", but if you fill in a previously defined subject, it could be "I [order] a hamburger", "My [favorite food] is a hamburger", etc.
(concerning/as for) myself, (it's) hamburger.
On it's own if Bob says this, basically it comes out as "I (am) (a) hamburger".However, if Sally has just said something like "I'll have a salad...what about you Bob?" then it makes sense as Bob's order is the implied subject and it becomes "My order is hamburger." or "I'll have hamburger."
I know very little about linguistics but I think there are a bunch of other things that make Japanese-English difficult to translate via software as well.
There is the whole aspect of culture embedded in it. あなた could mean "you" or something like "dear/sweetie" depending on the context. There also the question of how to translate "you" (etc) in English text to Japanese as you have to consider politeness etc. If you are just translating a business web page it's probably safe to stick with polite forms, but if you are translating say the dialogue in a TV show you want to preserve the tone of the characters.
In terms of voice recognition, Japanese seems to have a lot of homophones to me when compared to English. It may just be my imagination, but here are some I ran into recently:
舶,錘, 頭, 摘む, 積む, 詰む, and 紡錘 are all pronounced つむ and mean completely different things. Or 六, 碌, and 録 are pronounced ろく. 上, 神, 紙, 髪, and 加味 are all pronounced かみ.
I seem to run into things like that regularly; when just hearing it spoken you need the context to figure out what they mean.
This can be mostly solved by context. There are very few situations in normal speech where you'd hear "kami" and not know if they're talking about 神 (god) or 髪 (hair). Also, it's not particularly hard to code that knowledge. E.g. try かみにいのる (pray to god) and かみのけをきった (cut hair) on Google Translate. It will suggest the correct kanji in both cases.
Anyway, I'm not a native Japanese speaker, but I find the whole homophone thing a bit overrated. As far I as can recall the only pair of homophones that cause trouble in normal speech are 科学/化学 (both pronounced kagaku, meaning science/chemistry) and 私立/市立 (shiritsu, private/municipal).
> This can be mostly solved by context. T
Right, as I said. It's not too bad, but it's easier when you can just translate word for word.
かみにいのる gives me "pray to bite" on Google translate; as you say, it suggests the right kanji...but that's precisely my point. It needs you to disambiguate for it to be sure.
I'm not saying this is an insurmountable problem, I'm contrasting the difficulty.
> There are very few situations in normal speech where you'd hear "kami" and not know if they're talking about 神 (god) or 髪 (hair).
I ran into it recently in music. Babymetal has a song that starts:
伝説の黒髪を華麗に乱し
When you listen to the song, it'd be easy to momentarily think she might be saying "black god" or "black paper" since while the pronunciation wouldn't be identical, it's pretty close. Since I'm human, I figured out pretty quickly what she is saying...but in the equivalent English phrase there's no issue there...it's "black hair" or "black paper".
This is admittedly not "normal speech", but I could see it popping up there too.
I've seen confusion over 神/髪 in other situations too, though those were deliberately puns so probably don't count, but demonstrate it's possible to have situations where it's at least somewhat ambiguous.
> I find the whole homophone thing a bit overrated
I'm sure it's exaggerated to me because my Japanese is pretty atrocious, but I think my point is valid: any time you have homophones in a language it makes things more difficult to set up a system that listens to speech and translates. Japanese seems to have more homophones than English, and if that's true it is proportionally more difficult to translate in that regard.
Fair enough. I'm not claiming there's no homophone ambiguity either, just that it's a relatively easier problem compared to, say, the stuff Microsoft is doing.
Yeah, when I say "normal speech" I don't include pop music lyrics.
This can backfire. I remember hearing that sometimes/back in the day "Baile Átha Cliath" (the Irish for "Dublin", the capital city of Ireland) would sometimes get translated as "London" the capital of the UK. This is due to Google Translate trying to match up Laws in Ireland (in the Irish language) with UK laws (which would be very similar or potentially based on the same original law). However in the Irish law "Baile Átha Cliath" would be replaced with "London".
Here's an example of it: http://translate.google.com/#ga/en/L%C3%A1%20alainn%20inniu%...
I wonder if there's some feedback loop caused by websites that used google translate itself to offer the alternative versions :)
Best would be parliamentary speeches with transcipts, and the closed captioning for national news programs. The main constraint is storage space/computational power.
In college I had studied Japanese and a friend introduced me to the anime cartoon Initial D. His copy had the original Japanese with English subtitles, and so I could assess the translation to some degree -- it was very good. On Netflix you can watch Initial D, but after 2 minutes I had to turn it off because the English dubbing really failed to capture the characters.
As someone noted in this thread, the presenter's synthesized voice in the linked video doesn't seem to reflect his own. If he could have said something like "Wo hui shua putonghua" and had the machine output say the same, it might have been more convincing.
The speaker says "to take in much more data" but it gets parsed by the speech-to-text as "to take it much more data" which is such an unlikely phrase I can't really work out why it's not auto-corrected.
The phrase provided doesn't appear to be in either Google's nor Bing's web indexes. Typing "to take i" in to either Google or Bing's search box produces a hit for "to take in" as the most likely match; and within milliseconds.
Similarly (and ironically) with "about one error out of" being parsed as "about one air out of".
That he goes on to say that they use statistical techniques and phrase analysis for the translation makes this sort of error all the more intriguing, why isn't that same statistical approach weeding out these sorts of errors.
Nonetheless an impressive demonstration.
- Because grammar is typically more expressive, and dependent upon a concept that otherwise may not exist in words. Thus statistical grammar models and context checkers would be much more volatile to generating nonsense from user input (along the lines of the Sokal hoax) or restricting output to a range of acceptable models (giving the machine its own voice in a sense). That leads to the second thing...
- It kills freedom and creativity (or at least, how we receive it). Imagine comedy routines in stoic deadpan. Perfunctory exchanges in formal constructions (and vice versa). Obviously you can avoid all of these situations if you wanted to, but in that case it should probably be saved for those special occasions. It could probably help a lot of businessmen wanting to write their statements and messages in shorthand without spewing boilerplate text. But it's potentially damaging to every child or student who is still finding out how they want to express themselves in the given context.
Note: I think it is fair to assume that grammar checking would include the ability to reformulate or generate text that obeys the relevant models. Spell checkers suggest spellings, grammar checkers have to suggest fixes and changes as well, and if we want to get any further than Win98 era Word it will probably have to have a plain old fix-it generator as well.
You are right about the source quality though.
If we can get to the point of having handheld devices that can accomplish live translation of spoken word, what exactly is the point of different languages anymore?
Also, bravo to Microsoft; I'll remove my jaw from the floor after I watch your video a second time.
So if we can all talk to each other across the world in real time, and we can all understand each other because of this technology, what exactly is the point of different languages anymore?
Trying to understand or think like an American is a completely different experience from trying to think like a Chinese which, respectively, is completely different from thinking like a German. These modes of thinking not only make each culture/country/peoples unique, it actually facilitates various strengths.
Primarily, the Chinese language actually allows humans to remember and store more information in short term memory when compared to more Western/romantic languages. Most humans can store 7 (+- 2) bits of information (bits being defined as one contextual idea: 25 + 23 are 3 bits, Picasso's Mona Lisa is 1 bit).
The Chinese language actually facilitates math/short-term memory because many things are spoken/read/written as one contextual idea. When remembering a large number (602-112-5097 for example), English speakers tend to remember this number as: Area code = 1 contextual bits Each individual number = 7 contextual bits This happens because the English language separates non-related digits into individual ideas.
In Chinese however, masses of digits are written and spoken as one long contextual idea. Similarly, in memory, these long numbers are more easily stored as one contextual bit and take up "less" space.
Conversely, English (and to some degree romantic languages) happen to have lots of descriptors (what we call adjectives and adverbs). This, respectively, is one of the reasons why English speakers tend to be more creative with how they express themselves.
Does this mean that we'll eventually tend towards an universal language that implements all the good points of current languages? Perhaps.
Personally, I enjoy the uniqueness of each language by itself.
Not always.
http://en.wikipedia.org/wiki/Diglossia
http://en.wikipedia.org/wiki/Creole_language
A bunch of things. First and foremost: cultural representation. There are many things that are simple words in Chinese that have no English equivalent. The values and traditions of a culture are subtly communicated via its language, and even if we can instantaneously translate the literal meaning of what is being spoken, subtext will be lost.
This is why at a high level (beyond "where is the restroom") translation is a highly involved field.
Secondly is the usability of this system. Even if it is 100% accurate and immediate you'll just end up with the UN problem: communication becomes asynchronous because you need to wait for the translator. You'd say one thing, the other person would listen to your translator. He'll say something back, and you get to listen to his translator. It's a hell of a lot better than nothing, but there is still a tremendous advantage to being able to converse in real-time.
True. This is why I will be teaching my children Basque and Basque only.
New languages might have originally emerged due to separation, but that doesn't mean lack of separation will cause already existing languages to go away.
I'd love some comparison - that doesn't sound like the same voice to me (awfully close to the 'standard' computer voice, IMO), but some of it is crummy recording quality, and showing the flexibility would go a long way toward convincing me.
Seems like my years of learning Chinese and living in China are about to become useless...
Actually I stopped learning Chinese and living in China because I discovered the following. They were learning English faster then I could learn Chinese, and I only needed to know enough to let them know I wasn't culturally insensitive.
Not sure where the quote was but it went along the lines of "Don't try to talk in their language, because you will make a hash of it and they will have the advantage."
Are Chinese computer voices just flat-out worse than English ones, or is there something specific?
1. is anyone here fluent in mandarin to assess the quality of the output?
It would be even cooler if they created a distribution of possible sound frequency for each syllables in both English and Chinese, and determined where in the distribution his speech pattern lies, and transfer the "ranking" in the distribution. Hence you get a subjective transformation instead of a objective one. :)
Very cool stuff.
There is no possible way to confuse that robot voice with the speaker. Technology like you suggest is a long way away from the consumer market.
Now it seems the time may be too late. Rats.
EDIT: Oh, I see. You think I think I had this idea first and out of nowhere. Wrong.
My comment was a bit of bittersweet admiration at a very specific success in implementation. Can we just assume that the average person on HN knows wtf a universal translator is, and not every time someone says, "I had this very idea" that they are making accusations that company/person X is stealing the idea away? Christ. Not a single word was spoken that this idea was "stolen" or "grabbed away". Nor were words employed signifying that "I thought of this first and completely independently".
I was saying it looks like it's too late to be the person who gets to be first to demonstrate it. And by "very idea" I meant, quite specifically, how cool it would be to be able to translate my voice into another language in near-real-time. And by "Rats", I meant "Oh snap, looks like others have been working on doing that same thing and impressively pulled the damn thing off."
If you spent a moment looking at my comment history, you'd see I am usually quite clear in stating what I mean directly. Had I meant to imply an idea was "grabbed away from [me] before I could create a startup", I'd have used those exact words. Instead, I said that I'd had this very idea years ago and only recently thought about how the tech might be there today to make it happen, and maybe make for an interesting startup. I didn't even try to imply it would be my startup.
Nonetheless, point taken. I'll ensure I more specifically couch any future statements about having ideas with the qualifier that I am, in fact, not implying that I had it first or independently in the whole of human history.
Anyway, my apologies either way. Your comment read a tad snarky with the "grabbed away before you could create a startup" bit. I was left thinking, "Oh great. Relegated to the company of some kid who thinks his startup idea was stolen." My reply could have been tempered a bit more.
I didn't hear the inflections in his voice superimposed on the Chinese voice, so it is just modeling his voice characteristics and reflecting that into output voice. From what I understand the voice modeling happened off line, not in real time, so it is not nearly as sexy as if it was happening as he spoke. In other words, a nice touch, but I don't see it as revolutionary (unless I'm missing something).
it is pretty neat.
English: "I hope you enjoy the rest of the presentations today" http://translate.google.com/#en/zh-TW/I%20hope%20you%20enjoy... Translated into Chinese: "I hope you enjoy rest (as in take a break)/resting introduction/presentation"
Bing translator is much better using the same English input: http://www.bing.com/translator, the Chinese that comes out is grammatically correct in Chinese. The reverse translation comes out correct as well.
But yes, seems that if the translator is good enough, it would be simple to reproduce the process from the video.
雞肉和米飯 is a really odd way of saying chicken and rice... if you google the phrase, the search results don't come up as good as the way native speakers would expect:
google 雞肉和米飯: https://www.google.com/search?q=%E9%9B%9E%E8%82%89%E5%92%8C%...
google 雞肉飯 (the more normal way in chinese "chicken meat, rice"): https://www.google.com/search?q=%E9%9B%9E%E8%82%89%E9%A3%AF&...
google "poulet et riz" https://www.google.com/search?hl=en&tbo=d&rlz=1C1CHF...
Hmm... they are all chicken, but still pretty different from each other, especially if talking to a foodie!
Maybe give M$ another 15 years, but for now, the best way to learn a language is to do it through the social method.