MDM: Human Motion Diffusion Model
guytevet.github.io
guytevet.github.io
"MotionCLIP: Exposing Human Motion Generation to CLIP Space"
https://arxiv.org/abs/2203.08063
Would anyone be able to explain how the two techniques are related?
What are the training data? Sounds like they would need tons of diverse labelled motion capture data.
robotics also needs instant reaction time, which i don't think this kind of thing is good at
maybe something like a rough pre-simulation step for further rapid prototyping work? can't imagine the practical application..
disclaimer: i have no idea what i'm talking about, and am just going through the motions, much like the diffusion model under discussion
Without that, you have a much bigger solution space to search for the optimal approach. And you likely wouldn't even want that optional solution anyways: if your humanoid robots have sufficiently strong actors, optional solution for, to pick a simple example, navigating a staircase would likely be vertically scaling up the handrails through the slot in the middle. That might save staircase capacity and a tiny amount of battery (better aerodynamics!) but it would certainly ruin any humanoid qualities.
On the other hand, if you take another step back and think of the query input to the diffusion model, perhaps a text "human walking up to a higher floor" in our stair-climbing example, then you are absolutely right, that would be much closer to what we'd call "conscious" than I had ever imagined!
(my usual model of technical consciousness is an entirely different beast, the "simulate before act" approach with the additional requirement that the simulation includes a "self" and "peers" and the simulation includes rules/mechanisms that put them in the same category)
The robotic challenge in those tasks is not insufficient Actuation capabilities or Control capabilities, but insufficient sensory Proprioception.
Sure you can compensate for those in Control or adding Vision guidance but ultimately what you needs is sensors/transducers that are able to mimic touch and pressure and proprioception in a way that approximates what exists in the animal kingdom.
I was also wondering about the possibility of generating, say, an anthropomorphized cartoon otter that can be animated using a model trained on both otter and human motion to produce a result that is something in between.
It could reduce the workload for producing animated stories by one or more orders of magnitude sometime in the possibly not-too-distant future.
And _then_ stable diffusion came out. The effectiveness of diffusion models for text to image is an idea that has been floating around a little longer than you might think.
Katherine's work on clip-guided-diffusion over the `guided-diffusion` ImageNet checkpoints was effectively the first time the public got to see what text-to-image via diffusion instead of purely transformer-based solutions (like in DALLE1/dalle-mini) would look like. And it happened well before GLIDE was published (and gets a mention/citation).
The CompVis team (Blattman, Rombach, etc.) has been able to not just compete, but surpass (in some ways - it's nuanced) the work of the big American research labs (OpenAI in particular) with solid novel research. Their research on `VQGAN` outperformed the Autoencoder from the DALLE-1 paper, and they've been competing directly in the vision space ever since.
Incredibly talented people.
The conference for reference
Which as a writer quite excites me, even though it'll probably be quite bad at the beginning, with a flood of terrible products similar to the flood of Unity games.
Besides, BD is doing fine with classical optimal control theory for Atlas' locomotion.