It’s hard to tell from the video, but I wouldn’t be surprised if they used their face capture animation to drive wrinkle maps[1] on their MetaHuman models. This is something that other apps don’t offer out of the box, but which can significantly increase realism.
The MetaHumans drop in models are the thing that makes this magic imo. It’s not the motion capture tech so much as the complete pipeline to produce game assets.
I agree that the automatic integration with MetaHuman is the main benefit.
I’m guessing that is for when you want to move your body around and capture the facial expressions associated with doing so. You need some sort of rig to keep the camera in your face if you’re moving your face.
https://www.youtube.com/watch?v=lXZhgkNFGfM
The detail on the new Metahuman face capture is visibly better and a smoother capture pipeline has been a valuable goal for a some time. Noisy mocap and re-rigging animations to share them between models, takes time.
[Quote from article] The algorithm uses a "semantic space solution" that Mastilovic said guarantees the resulting animation "will always work the same in any face logic... it just doesn't break when you move it onto something else."
- Whole-clip solve instead of noisy frame-by-frame streaming - Automatically generating a face model to calibrate the rig - Solving from raw sensor data instead of Apple ARKit pose data
VFX studios typically develop proprietary solutions a few years in advance, so it can be hard to say what is truly "new and different."
In the past, it was common to use video capture with reference markings on the actor so you could figure out how they were moving in three diminsions.
From this demo, it looks like they are directly leveraging the 3D data from the iPhone's Lidar sensor in addition to the video camera data.