Using NeuralNets to make smooth character animations
techcrunch.com
techcrunch.com
but maybe in the future we'll have fully NN autoencoded games!
In many games the user wants to see close to immediate response to their input, e.g. moving the stick to turn or pressing a button to jump but that inherently conflicts with the desire to smoothly transition between animations.
In reality, your brain starts planning an upcoming turn or jump and moving your body accordingly some time before reaching the physical obstacle you are avoiding. Control schemes for games generally don't allow for this kind of anticipatory input but rather you press the jump button at the moment you want to jump. This leaves no time to smoothly blend the animation - by the time the game receives the input it should already have started the blend some time ago.
I spent ages tuning the code that determined how high to place the dog's centroid above the ground. I gained a lot of respect for the amount of geometry calculations a real dog has to do in order to keep its face out of the floor when running. The code needed to start reacting to a probable landscape collision hundreds of milliseconds before it happened in order to prevent the need for unnaturally large forces.
Obviously the thing to do is to predict user input. Capacitive sensors within buttons to try to detect when your finger is above or resting on a key? Or just hall effect analog buttons, so you can tell when the button starts to go down before it reaches the bottom of the depress action.
If you play a game like Street Fighter 2 you'll note that the animations for many of the basic attacks are extremely brief, something like 11 frames in duration. This may further be reduced by animation cancelling, which allows a new animation to override the currently playing one, speeding things up.
I'm generally skeptical of the ability to speed up the frame rate in order to smooth things out anyway. Players tend to use certain key frames as cues to help them properly time their attacks. By running at a higher frame rate, you reduce the temporal distinctiveness of these key frames. At any rate, I'd like to see and judge the result for myself.
Imagine this system used to implement rock climbing. The player is simply pressing the "W" key to go up, but depending on their character's skill level the speed and manner in which traversal occurs could be very different.
I really hope Bethesda is paying attention to this, because open world games could use it more than anything. In Skyrim you climb mountains by mashing the jump button. Far from ideal, and tends to break immersion.
The tech could even be used to power melee animations some day, perhaps entire fighting styles.
It would likely have secondary effects of easing the burden on animators, allowing them to concentrate on more important aspects of the game, like facial interactions.
(Looking at you, BioWare.)
We can already synthesize the speech itself to the point it sounds real, although it's very computationally intensive at present. Infusing the speech with emotion remains very difficult. Some day voice actors will probably just license their likeness along with a sample of emotive recordings.
After that, one of the few remaining holy grails is generative dialog for NPCs that's both believable, dynamic in relation to the game world, and don't sound like glorified ELIZA bot output.
Neural nets can probably be used for level design today. One use case might be creating an entire urban environment complete with residential interiors, where the artists don't have to slave over each individual apartment for it to be believable. Imagine playing a war game where the levels look like there's actually people living there.
Clearly early days, but still, the promise is there.
So far, correct me if I am wrong, the state of the art prior to this would be something like Unity's Mecanim, where you define a state machine for animation transitions, which would interpolate animations to save you work.
Very handy for character locomotion, but only gets you so far. I'm not sure how applicable offline processing is to this though, since if you have a large number of different types of obstacles to be traversed, you'd have to bake >= that number of animations, and your animation state machine would be enormous. The end goal might be to do that in real time, but I don't think anyone is going to seriously suggest running a neural net to calculate animation positions while the game is running. Maybe if you had some smarts in the level designer you could determine the potential pool of animations needed based on placed geometry, and build the animations at packaging time along with the state machine.
If that is still not enough, they can also try to look at ways of producing a "good enough" effect with a smaller network.
Then, it can also be limited to only some characters, the ones that you are more likely to pay attention to.
Now, the MS Kinect does some intense processing behind the scenes (random forests) to capture your motion in real time. Yet you can still run games with decent performance. I think a neural network to adjust animations is not too dissimilar in terms of performance cost.
http://theorangeduck.com/media/uploads/other_stuff/phasefunc...
http://www.gameanim.com/2016/05/03/motion-matching-ubisofts-...
I would always trust actual concrete implementation rather than a research project though... unless they have source and demo available.