We can already synthesize the speech itself to the point it sounds real, although it's very computationally intensive at present. Infusing the speech with emotion remains very difficult. Some day voice actors will probably just license their likeness along with a sample of emotive recordings.
After that, one of the few remaining holy grails is generative dialog for NPCs that's both believable, dynamic in relation to the game world, and don't sound like glorified ELIZA bot output.
Neural nets can probably be used for level design today. One use case might be creating an entire urban environment complete with residential interiors, where the artists don't have to slave over each individual apartment for it to be believable. Imagine playing a war game where the levels look like there's actually people living there.