1. Tons of diverse data sets (real world)
2. Solution for Noise - Either de-noise and train OR train with noise.
There are lots of extra challenges that voice recognition problem have to solve which is not common with other deep learning problems:
1. Pitch
2. Speed of conversation
3. Accents (can be solved with more data, I think)
4. Real time inference (low latency)
5. On the edge (i.e. Offline on mobile devices)