Not an expert in either technology but I would imagine a lot of the difficulties of speech to text (bad recording conditions, variety of accents & pronunciation differences, etc) also have analogues in computer recognition of hand signs (camera alignment is bad, cut off, lighting is bad, someone's hand signs are lazily performed or slightly different than textbook ASL, etc).
Speech to text was "solved" a long time ago but I've seen it take many years to become as usable as it has recently. And it still regularly is frustrating to use for me!