Just watched the video, but is it correct to assume the inputs to this are:
- Manually specifying layers in the image by depth.
- Manually specifying all the objects you'd want to animate.
- Providing an animation direction and type.
Then it generates an animation?