Microsoft Research uses Kinect to translate between spoken and sign languages
thenextweb.com
thenextweb.com
First issue: The signing they are doing in the video is very simple and limited. Not just in the number of signs, but in the complexity of what they are doing. Their face remains perfectly neutral. They move slowly and carefully. They are signing individual signs one at a time. That's not how a signed language works!
Signing without facial expression is like speaking in complete monotone, making no eye contact with the person you're speaking to. You can probably get your point across, but you're not having a real conversation, and you're missing something important. The lack of flow and motion to the signing is also awkward as hell. In short, the people in those videos are not actually fluent in sign, they're trained actors who have memorized certain motions.
Second issue: There are concepts in sign language that have no equivalent in English, or many spoken languages.
One example- in ASL you have variables/registers, in the programming sense. You can sign "John", then point to your left. From now on in this conversation, if you point left, you mean John. If while you speak, you want to describe John moving to your right, you can do that (explained next). Now pointing right means "John".
How might you 'move' John to your right? You could have a single finger pointed up, like John was a finger puppet, and walk him to the right. If he got into a car, you could have that single finger get into another specific hand form that generally means 'vehicle'. Then you could draw a tree with your free hand, and drive the car into the tree. You just told the story about John driving into a tree. There were no specific nouns or verbs used. This is how a lot of conversations work in ASL.
The point I'm trying to illustrate above is that description and conversations happen very differently in signed languages than they do in English. It isn't that the words come in a different order, it's that the idea you are trying to communicate is explained in a completely different way. You're comparing apples and coconuts.
Third issue: every city has its own dialect. Signed languages aren't written down, there isn't much global media of people signing, and the result is that dialects change constantly. Drive to any other major city, entire words have changed. Talk to a family across the street, they might have different signs for some words. Sure, the core language is the same, but you've got a lot of different nouns and verbs.
Going even deeper, some people here have already mentioned the difference between ASL and Signed English. Some people sign English words one at a time in the order of English grammar. Some people go into the 'pure' ASL realm, as I described above, and use very complex concepts not used in English. It's not a choice, they'll have simply learned ASL that way. Most people are somewhere in between the two extremes, and it's a continuous domain, not discrete, with a huge variety.
I won't discourage research into this, and I think what they're doing is awesome- I considered doing a master's thesis on this exact idea. I just want to put a nice big disclaimer here that this isn't useful yet, and there's a long way to go.
Edit: Wall of text much mabbo? Woah man, take a breather.
So I see this implementation as the same, its part of the long road. First, start with the basics. Prove the machine can recognize what you intend. Then step it up. From there you can broaden it other dialects and what not.
Still you cannot help but cheer them on
> While this is clearly a massive achievement, there is still a huge amount of work ahead. It currently takes five people to establish the recognition patterns for just one word.
I don't get it. They're saying they need five (only five!) seperate people to train the thing, and then it works pretty well for everyone? That seems to be working pretty darn well to me, I don't see the problem or 'huge amount of work'.
Probably something lost in the journalism.
I know nothing about CSL, but ASL phonology is rather complicated. There is a lot of assimilation, for example, and the meaning of signs can change based on the spacial orientation of consecutive signs.
My old linguistics professor told me about a joke someone made at an ASL conference. He was signing with a colleague, and talking about his progress improving his ASL. She then begins making the sign for "improve":
http://lifeprint.com/asl101/pages-signs/i/improve.htm
But instead of making contact with her upper arm, she shifts her right arm downwards and repeats the starting position of the sign. This is the ASL equivalent of a "...NOT" joke, as in, "your signing has improved... NOT!"
The degree of space between starting and ending positions in the "improve" sign can articulate different degrees of improvement.
While I suspect this kind of nuance isn't yet translatable by this project's sign-recognition software, I would really love to see what kind of progress they're making in this area.
That being said, the FCC still pays $6 a minute for VRS services.
Besides this is a research project. Just goes off to show the potential a full body recognition platform such as Kinect can have.
Basically ASL (and other forms of sign language) use constructs and tense that don't necessarily translate well into specific words without lots of experience.
A few examples: The signs for 'wish' and 'hungry' look the same. The difference is the facial expression.
If you wanted to sign, "All we want to do is eat your brains", the order would most likely be: 'We want do-what? Your brains eat.'
There is another form of signing called 'Signed English.' It borrows most of ASL's signs but puts them in standard english order. It has been a bit controversial in the past, though.
You also run into odd colloquialisms or regional signs. I'm not quite sure what to call them. A quick example is the word 'grass.' I learned the sign for grass but, for some reason, that sign means 'truck' where I live currently.
Anyway, I'm rambling. If you have any specific questions I can try to answer them.
edit Watch this youtube video with captions enabled. The video is captioned in english and in "ASL": http://www.youtube.com/watch?v=UQYjZc7gKXc
Sign is incredibly conceptual and has a large spatial component. This is definitely a cool step though.
ASL can be very expressive and can be an art form itself. There's a neat thing in the deaf culture called 'ABC Stories.'
http://www.youtube.com/watch?v=wBUdzGH6WbU
They are stories that are told using the alphabet finger spelling signs.
Does anyone know what techniques they are using to accomplish this? I did the changes in Active Appearance Models, converted them into 'sound files', then ran them through Microsoft's HTK Markov Model system to get my stuff.
In theory, they might be doing a similar thing.