The state of AI for hand-drawn animation inbetweening
yosefk.com
yosefk.com
This is one of the most overlooked problems in generative AI. It seems so trivial, but in fact, it is quite difficult. The difficulty arises because of the non-linearity that is expected in any natural motion.
In fact, the author has highlighted all the possible difficulties of this problem in a much better manner.
I started with some simple implementation by trying to move segments around the image using some segmentation mask + ROI. That strategy didn't work out, probably because of some mathematical bug or data insufficiency data. I suspect the later.
The whole idea was to draw a segmentation mask on the target image, then draw lines that represent motion and give options to insert keyframes for the lines.
Imagine you are drawing a curve from A to A. You divide the curve into A, A_1, A_2... B.
Now, given the input of segmentation mask, motion curve, and whole image, we train some model to only move the ROI according to the motion curve and keyframe.
The problem with this approach is in sampling the keyframe and matching consistencies --making sure RoI represents same object-- across subsequent keyframes.
If we are able to solve some form of consistency, this method might be able to give enough constraints to generate viable results.
> But what’s even more impressive - extremely impressive - is that the system decided that the body would go up before going back down between these two poses! (Which is why it’s having trouble with the right arm in the first place! A feature matching system wouldn’t have this problem, because it wouldn’t realize that in the middle position, the body would go up, and the right arm would have to be somewhere. Struggling with things not visible in either input keyframe is a good problem to have - it’s evidence of knowing these things exist, which demonstrates quite the capabilities!).... This system clearly learned a lot about three-dimensional real-world movement behind the 2D images it’s asked to interpolate between.
I think that's an awfully strong conclusion to draw from this paper - the authors certainly don't make that claim. The "null hypothesis" should be that most generative video models have a ton of yoga instruction videos shot very similarly to the example shown, and here the AI is simply repeating similar frames from similar videos. Since this most likely wouldn't generalize to yoga videos shot at a skew angle, it's hard to conclude that the system learned anything about 3D real-world movement. Maybe it did! But the study authors didn't come to that conclusion, and since their technique is actually model-independent, they wouldn't be in a good position to check for data contamination / overfitting / etc. The authors seem to think the value of their work is that generative video AI is by default past->future but you can do future->past without changing the underlying model, and use that to smooth out issues in interpolation. I just don't think there's any rational basis for generalizing this to understanding 3D space itself.
This isn't a criticism of the paper - the work seems clever but the paper is not very detailed and they haven't released the code yet. And my complaint is only a minor editorial comment on an otherwise excellent writeup. But I am wondering if the author might have been bedazzled by a few impressive results.
You can also see much more 3D ish things in the paper, with 2 angles of a room and video created moving the camera between them. Of course in some sense it adds to my point without detracting from yours...
There's a general issue with generative AI drawing "a horse riding an astronaut" - art generators still struggle to do this because they just can't generalize to odd scenarios. I strongly suspect this method has a similar issue with "interpolate the frames of this yoga video with a moving handheld camera." AFAIK these systems are not capable of learning how 3D people move when they do yoga: they learn what 2D yoga instructional videos look like, and only incidentally pick up detailed (but non-generalizable) facts about 3D motion.
I don't have more than a fuzzy idea of how to implement this, but it seems to me that key frames _should_ be interchangeable with in between frames, so you want to train it so that if you start with key frames and generate in-between frames, and then run the in-between frames through the ai, it should regenerate the keyframes.
(I guess we're used to machines and people struggling at opposite things so this is counter counter intuitive, or something...)
Animation key frames are not interchangeable with inbetween frames since the former try to show the most body parts in "extreme" positions though it's not always possible for all parts due to so called overlapping action. This is not to say you can't generate plausible "extremes" from inbetweens; acting wise key frames definitely have the most weight.
AI being good at stills is true, though it takes a lot of prompting and cherry picking quite often; most results I get out of naively prompting the most famous models are outright terrifying.
Think of this in terms of constraints. An image from scratch has self consistency constraints (this part of the image has to be consistent with that part) and it may have semantic constraints (if it has to match a prompt). An animation also has the self consistency constraints, but also has to be consistent with other entire images! The fact that the images are close in some semantic space helps, but all the tiny details become so important to get precisely correct in a new way.
Like, if a model has some weird gap where it knows how to make an arm at 45 degrees and 60 degrees, but not 47, then that's fine for from-scratch generation. It'll just make one like it knows how (or more precisely, like it models as naturally likely). Same with any other weird quirks of what it thinks is good (naturally likely): It can just adjust to something that still matches the semantics but fits into the model's quirks. No such luck when now you need to get details like "47 degrees" correct. It's just a little harder without some training or modeling insight into how an arm at 45 degrees and 47 degrees are really "basically the same" (or just that much more data, so that you lose the weird bumps in the likelihood).
I wouldn't be surprised if "just that much more data" ends up being the answer, given the volume of video data on the internet, and the wide applicability of video generation (and hence intense research in the area).
These things combined mean less information to learn a more difficult world model.
Training on a wireframe model seems like it would be easier, since there are plenty of wireframe animations out there (at least for humans) you could use and remove in-between frames to try inferring them.
And AI has less room to hallucinate - it is more a kind of interpolation - even if in this short curt example, the AI still "morphs" instead of cleanly transitioning.
The real animation work and talent is in keyframes, not in the inbetweening.
I don't know if that's really a good pathway to become a key animator - how many inbetweeners are there for one key animator?
As to how much poor quality in-betweening hurts the performance to the audience is a complicated discussion. Animation that is _very_ bad can often be well accepted if other factors compensate (voice acting, design, direction, etc.)
A good in-betweener is not simply interpolating between the keys. For hand drawn animation at least, there's a lot more going on than that.
We'll leave out any discussion of breakdowns here. For one it's a difficult concept, much more difficult than 'tweening to explain. The other is that different animators will give different opinions on what a breakdown is or does.
I will say, though, I think that properly tagged breakdown drawings could significantly improve the performance of ai generated in-betweens.
Anyone who is seriously interested in the process should read the late, great Richard William's book, _The Animator's Survival Kit_. This is especially true for those who want to "augment" the process with machine learning. The book is very readable, even for non-artists. And he gets into the nitty gritty of what makes a good performance, and the mechanics behind it.
Edit: Another good resource, and relevant to 3D animation as well, is Raf Anzovin's _Just To Do Something Bad_ blog. He has many posts on what he calls "ephemeral rigging" that are absolutely fascinating. Be aware that the information is diffused through out the blog and not presented in a form for teaching. His opinions are fairly controversial in the field. But I think he is onto something. (https://www.justtodosomethingbad.com/)
The thing about animation is that it is not about interpolation. It's about the spacing between drawings. The methods developed by animators were not at all mathematical, but something that they felt out by experimentation (trial and error).
The math that does enter into it are directly related to the frame rates. If animation had started in modern times, with frame rates of 30 fps or 60 fps, it would have been a very different animal. And much harder!
At 12 fpt or 24 fps you have a very limited range of "eases" that can be done. So while eases do figure into it, its the arcs, the articulation, and the perceived mass of the parts of the character that make it seem alive. Looking only at the contours and the in-betweens misses all the action.
An awareness of the graphic nature of the drawings, the stylizations of figures and faces are also critical. Cartooning is its own artform and it is tied directly to the way human brains make sense of what the eye sees. Getting more realistic often takes you further from your destination.
Storytelling is also a core part of good animation. Making a character seem to think and react, like it is alive can be done by a good animator. But you won't get there by imitating the real world directly. Rotoscoping has very limited use in good character animation and storytelling. It's all about abstracting out what the brain feels is important and what it expects. You can get away with murder if you caricature the right details.
When I've worked with training new animators, one of the points I stress is that it is articulation and the perceived mass of the character that really sells a performance. The best art style in the world is nearly useless if the viewer doesn't buy into the notion that they are watching a thinking person reacting with a physical body to events in an interactive world.
My feeling is that you will get further if you build articualated rigs and teach the ai to make it move. 2D or 3D. There is footage of tiny AI driven robots in a Google eperiement that are learning to play soccor. The ai is learning to make them move and solve problems (running around the soccar field and scoring goals.) Very natural looking behavior (animation!) develops almost automatically from that.
Trying to solve the problem by dealing with lines, contours, and interpolation seems very far away from the important parts of animtion.
Just my two cents worth.
Get a copy of the Williams book, it's on Amazon. Read his thoughts, he explains things much better, and more entertainingly, than I do. Sharpen up your pencil and start making some simple walks. Simple stick figures and tube people work just fine. And you may find that you enjoy the art form. Even if you don't become an animator yourself, the exercises will deepen your appreciation and understanding of the art form.
I hope to avoid building rigs because they're, well, rigs. Much nicer and more flexible to control things though drawing than a rig which has a bunch of limitations and then there are issues with hair/cloth/water/etc. What can be done without a rig is another question but the methods I reviewed in this post are not the most that can be done for sure.
I'm not just being cute when I say that. The problems the AI in the examples was having have distinct analogs with the problems human animators have. Arcs are a problem, as is the notion that in-betweens are mostly about interpolation.
As I said, timing and articulation are at the heart of most kinds of animation. Even very stylized animation must be aware of this, if not being a slave to it. Imagination and expression are important, but first the audience has to _believe_.
The discussion about converting to a vector format was an interesting diversion. I've been experimenting with using potrace from inkscape to migrate raster images into SVG and then use animation libraries inside the browser to morph them, and this idea seems like it shares some concepts.
One of my favorite films is A Scanner Darkly, and that used a technique called rotoscoping which I recall was a combination of hand tracing animation and computers then augmenting it, or vice versa. It sounded similar. The Wikipedia page talks about the director Richard Linklater and also the MIT professor Bob Sabiston who pioneered that derivative digital technique. It was fun to read that.
Technically it's "interpolated rotoscoping" using a custom tool called Rotoshop, which takes vector shapes drawn over footage then smoothly animates between the frames giving a distinct dream-like look to it.
Rotoscoping is where you work to a traditional animation framerate drawing over live action but each frame is a new drawing and doesn't have the signature shimmery look Scanner Darkly and Waking Life so I think it's worth pointing out the distinction.
I know hand-drawn 2D is its own beast, but what's your thought on using 3D datasets for handling the occlusion problem? There's so much motion-capture data out there -- obviously almost none of it has the punchiness and appeal of hand-drawn 2D, but feels like there could be something there. I haven't done any temporally-consistent image gen, just playing around with StableDiffusion for stills, but the ControlNets that make use of OpenPose are decent.
3D is on my mind here because the Spiderverse movies seemed like the first demonstration of how to really blend the two styles. I know they did some bespoke ML to help their animators out by adding those little crease-lines to a face as someone smiles... pretty sure they were generating 3d splines however, not raster data.
Anyway, I'm saving the RSS feed, hope to hear more about this in the future!
I sort of hope you can handle occlusion based on learning 2D training data similarly to the video interpolation paper cited at the end. If 3D is necessary, it's Not Good for 2D animation...
AI for 3D animation is big in its own right; these puppets have 1 billion controllers and are not easy for humans to animate. I didn't look into it deeply because I like 2D more. (I learned 3D modeling and animation a bit, just to learn that I don't really like it...)
I wonder if it would be beneficial to train on lots of static views of the character too - not just the frames - so that permanent features like the face gets learned as a chunk of adjacent pixels, so when you go to make a walking animation, the relatively low amount of training data on moving legs in comparison to the high repetition of faces would cause only the legs to blur unpredictably, where the faces would be more in tact - the overall result might be a clearer looking animation.
I'm not surprised that using off the shelf diffusion models or multi-modal transformer models trained primarily on still images would lead to this level of quality, but I am surprised if these results are from models trained specifically for this task on large amounts of animation data.
One problem with diffusion and video is that diffusion training is data hungry and video data is big. A lot of approaches you see have some way to tackle this at their core.
But also, AI today is like 80s PCs in some sense: both clearly the way of the future and clumsy/goofy, especially when juxtaposed with the triumphalism you tend to hear all around
why can't this basic idea be applied to simple 2d animation over two decades later?