Human Motion Diffusion Model
github.com
github.com
My gut says that the quality could be rapidly improved without changing the underlying design at all.
The real issue with this, I think, is that motion capture for humans is already widely available and provides much higher fidelity and control than text. Unless I'm misreading the paper badly, this model was trained on exactly such data. Blending between multiple animations through motion capture is also well-understood.
So while the results are impressive, the practical gains seem very marginal. I think perhaps that the equivalent to "inpainting" (as mentioned in the text) and "style transfer" would be the big gain here? If we could use this to retarget animations to different body plans (child, adult, space monster) quickly, or for smarter interpolation between human-authored keyframes, I could see that being a much-desired tool.
So if they’re training on the standard corpuses of motion capture data available and even mixing in their own, they likely won’t have fingers to base data on.
But it's cumbersome to put on and take off, and to operate, especially when working alone. While I'm in pretty good shape, there's heaps of movements (eg. martial arts, swordplay, firing/reloading a gun etc) that would probably look silly if I performed them. I can see this being very handy at least for prototyping animations at the very least.
Replacing finger bone positions is pretty trival in Blender as well using the Pose Library feature so the lack of finger data isn't that much of a big deal.
Either way, the shoulder and elbow joints don’t move much during rope skipping, and no matter how it’s recorded, the motion capture data that was used for training should reflect that.
My best guess is that the model has picked up on the tiny arm motions that are present in rope skipping and wildly exaggerated them for some reason.
Not sure if this diffusion is the answer but some smarter way to extrapolate existing animation to work with bodies and objects that are 10-20% different and look natural.
Zork and QWOP's lovechild? :-)
I never imagined ML or stable diffusion would be the answer here, but now I wonder if it will be? So much stable diffusion stuff has so far felt to me like just toys to play with, but modeling movement seems like it could make a gigantic difference in videogames and animation generally.
Compare this model with how Jordan Mechner digitized human movement for his game Prince of Persia in the 80's:
Try to find the email addresses of the authors of the scientific papers that go along with the models, there's usually a good chance that someone will answer if your request is reasonable.