Shape of Motion: 4D Reconstruction from a Single Video
shape-of-motion.github.io
shape-of-motion.github.io
We've been waiting 4 years. I just don't understand what is taking so long.
Even at a low resolution, the difference is night and day. Even with a very small window, this is a leap forward for VR immersion. Why in the hell is no one using it?!?
As they state, it requires a 46 camera rig to record, immense computational power to compress the data into something usable, and a 1 Gbs data stream, which means there is nearly zero commercial use, and certainly not widespread use, at those specs.
Oh, and it's covered in a patent minefield. You can check the author names to see. Pretty much every paper from Google gets covered in them, so be wary of implementing anything you read in their papers without a serious patent search. This one took me 10 seconds.
Not everything whizzbang you see in a paper or demo is not being sold as consumer products because of ignorance. There is always a reason, usually one you have not discovered yet.
I can't say I've had any good experiences with VR video... it's incredibly natural for gaming, but for scripted video having my head control the camera dilutes the experience for me: you either lose a key part of the language of cinema, or you make people sick by periodically taking control away from them.
The absolute best application of VR video I can imagine is immersive theatre, and I think with videogame engines are accomplishing that very well with modern performance capture.
Is there an experience you'd recommend that you think could change my mind?
Beyond that, though, there is a lot of potential just past the horizon. Check out the demo for the paper I linked. Sure, you can't move around very far, but you can move enough to be part of the scene, unlike the usual tripod video setup.
One challenge here is dedicated hardware support, running any sort of video processing/decoding on the CPU is unfeasible (GPU less so), particularly if the context is a battery-powered device (Quest 3, Apple Vision Pro), with high resolution and framerate (6K-8K, 60fps).
That doesn't even get into the logistics of capturing, serving, and software supporting this video data either
Even with its relatively low resolution, the demo blew me away. Given a choice between current 8K@120 content and something equivalent to this 4 year old demo, I would choose the latter hands down.
Getting that kind of look around in a video scene would be really engaging. A bit different than VR or watching in The Sphere, with the engagement being that there are still things right out of view you have to pan the camera for
It might be interesting for one or two movies specifically built around the feature, but otherwise it would be a gimmick no one would care for. For games, sure, but movies are a different experience.
Maybe you could have it work as a documentary (good luck getting a bored social group to go for that) or a virtual tour, but we already have 3D interactions of those too.
We’ve had tons of movie viewing experiments and ultimately always go back to the tried and true 2D screen, with the bolder ideas being relegated mostly to the domain of theme park gimmicks. Which are interesting in their own right, but don’t survive on their own.
> how do you deal with cuts and changes of scenery?
the same way the games did it. by doing nothing special at all and retaining the same functionality. it really depends on how this 4D reconstruction works before I could say it uniquely adversely affects the experience
for the most part what's interesting to me is that the overhead costs seem low enough not to care about random things big studios did at great expense with no way to justify the market appeal. its either a portfolio piece or 1,000 monhtly users supporting my lifestyle indefinitely.
(Funny there is a VR mod for Monster Hunter Rise which makes me think just how fun Monster Hunter VR would be)
There are exceptions but not many. It takes a ton more work to make a scene look good from any arbitrary viewpoint.
> we utilize a comprehensive set of data-driven priors, including monocular depth maps
> Our method relies on off-the-shelf methods, e.g., mono-depth estimation, which can be incorrect.
They also had the "Immerge," which was a ~1m diameter, hexagonal array of 2D cameras. They got the 4D data from having a 2D (spatially distributed) array of 2D samples (each camera's views). It's under sampled, because they threw out most of the light, but using 3D as a prior for reconstructing missing rays is generally pretty effective.
But I also understand a lot of what they demoed at first was smoke and mirrors, plus a lot of traditional 3D VFX workflows. Still impressive as hell at the time, it's just that the tech has progressed significantly since ~2018.
I've seen open source software for plentopic images which might be able to generate a point cloud but I've only gotten one good shot of the Lytro which was similar to a shot I took with this crazy lens
https://7artisans.store/products/7artisans-50mm-f0-95-large-...
But w/ time should it be 7?
You can move a virtual camera 3-dimensionally within the scene at any individual frame (x, y, z), and also move the scene through its animation to play the animation forwards and backwards (in other words, you move the camera through the 'time' axis).
So, like the scrubber in any video? Doesn’t feel like that warrants the 4D moniker. Which is not to say you’re not right, I think you are and that’s what they mean, but it that being the case it feels more buzzword than anything.
You are correct though, they both serve the same function.
However, I’ve yet to see a video player that lets you reposition the camera as if you were using photo mode in a video game. That’s (essentially) what this thing offers.
and synthesizing a (limited) model with 3 spatial dimensions, plus 1 time dimension.
3D over time is colloquially called "4D;" though we don't call video "3D" by analogy as the term binds strongly to its purely spatial use.
Colloquially, meaning “used in or characteristic of familiar and informal conversation” 4D films have a definition, and that ain’t it. 3D over time is colloquially still referred as 3D, as evidenced by decades of 3D blockbusters.
I've been interested in the state of the art in that domain myself, having thousands of 2D videos I've shot which I would love to see "spatialized" well, someday.
There's the 2D frame and the time dimension. Then there's the structural information conveyed by motion, parallax, scene composition, camera movement, etc.
That's why there's the 180 rule, amongst other things.
Algorithms can take a video and turn it into a 4D volume. As can our brains.