Let's take an example: the sign for "send email" in ASL (http://www.lifeprint.com/asl101/pages-signs/e/email.htm). If I point at you at the end of the sign it could mean "I will send you an email." If I point at myself it could mean "You should send me an email" or "Did you send me an email" depending on my facial expression. If I point off into space it could mean "I'm sending an email." If I start by pointing at you and then end the sign by pointing off into space it could mean "You should send an email." So your translator AI needs not only to understand the facial expressions and movements of the signer, but also the spacial relationships of everyone in the conversation. And that is just one aspect of the difficulty - there are many other features of sign languages that are just as hard to translate.
Perhaps this is the sort of thing that future AI systems could do. But it is quite complex.
Speech to text was "solved" a long time ago but I've seen it take many years to become as usable as it has recently. And it still regularly is frustrating to use for me!
Not every deaf person uses sign language and there are many different sign languages in the world. American Sign Language(ASL) is but one of these languages.
Specifically, sign languages are not visual representations of existing languages (e.g. ASL and English) but completely different languages altogether.
I'm sure other signed languages have rough equivalents in their regional areas as well. I know Mexico has a few different signed languages, though I only have passing familiarity with 1 of their signed languages and it's definitely not a representation of Spanish.