Experiment: Can 3D improve AI video consistency?
backdroptech.github.io
backdroptech.github.io
Overall, it worked. No sudden changes in proportions, clothing, or style. Still, there are some limitations, especially with fine details.
We’re looking into whether this could be useful as a tool and would love to hear what you think: Has anyone experimented with 3D + AI generation for images or video? or sees a better way to approach this?
Demo and details in the blog: https://backdroptech.github.io/3d-to-video/
1. The users with experience and patience for this are slim. Blocking out a scene and the animation are tough, and the users with this skill and inclination are using Blender and ComfyUI already or are submitting renders to RunwayML V2V. We're still too early for AI auto rigging and animation to work, though those technologies will make this approach easier.
2. AI video users want I2V quality and predictability, not unpredictable V2V style transfer. You need control over the exact look and feel of the starting frame as well as the animation. If you can't get this, the renders are useless.
3. One of the advantages of AI video is that it can animate things a human animator cannot easily do. Non-humanoids, crowd movements, explosions, etc.
Basically, this requires deep integration with a new class of video model.
You'll find that even with this technology perfected, it fits into a comprehensive suite of tools that AI video creators will use. They will still lean on I2V for most shots and V2V compositing for other shots.
I've done a lot of hands-on interviews and demos. Steve May, various studios, schools, etc. Steve kind of negged me and told me there are bigger players working on this. My guess was Odyssey Systems at the time, but they turned out to be working on something else.
I do think this is a valuable technology, but there's a tremendous amount of work to do to make it work.
1. I assume you mean patience when setting up a 3D scene? That’s definitely a factor, but it’s getting easier with image-to-3D tools, and AI can even assist with object placement to speed things up.
2.Yeah, predictability is key. Our approach is about making it easier to generate high-quality, consistent images, which are then fed into video models—rather than relying on direct video-to-video style transfer, which can be more chaotic.
3. Agreed! AI can animate things that traditional methods struggle with, but consistency is still a challenge. This workflow helps strike a balance between AI flexibility and user control.
The quality is identical to human animators (because: same tools, same process).
Just the cost is lower (although training the AI workers is a new cost).
The only companies that can do this are those that have a very strong workflow, because AI workers operate on individual steps, NOT the entire workflow.
2D animation already has this, so it's easier to adapt to AI workers than 3D (for that reason).
I could see applying changes at the 3D model level which wouldn't be directly accessible if it was only an internal representation.
For example, you can attach a LoRA specifically to one object and run a separate workflow just for that, giving you way more control. That’s a big difference—Gaussian Splatting doesn’t naturally lend itself to object-level edits since everything is blended into the same representation.
But the whole premise of AI video is that you're basically directing, and the model does all the hard parts for you.
That makes it a lot more useful.
Is it training a model based on 3d model? Is it doing img2img? is doing video2video using 3d video/image as source?
Article is showing off stuff but is light on explanation of what is going on here.
Also, the person is bearded. Hard to notice face changes there. I want to see more demos but instead of a 3d like render, do a realistic render without beard etc.
edit: Plus these are very short clips. 3D should theoretically help, but this article could have been better.
edit2: in the anime example, the girls face is definitely changing (+lips are not correct). It feels like you are doing img2video here?
This really increased the quality of results for me.
The complexity and cost likely increase approximately with the third power. Does the result justify this effort?"
By passing a structured 3D-rendered image to the AI, you increase the chances of the video model generating exactly what you want, rather than relying on unpredictable, frame-by-frame AI improvisation.
Maybe add "digital" too...
But I have very recent first hand experience of creating a video for our startup's Facebook post with Minimax image-to-video inference, from an image of our animated avatar character.
...And yes, first the videos were bad quality with lots of inconsistencies, but after adding "animated" to the prompt, in front of the "man" word, the result was pretty great already on the first try! Which I then ended up even using. (you can even check it here if interested https://fb.watch/xRC-fptexM/)
Perhaps it should be self-evident, but still, it was not to me. :)
Edit. I guess my point was also that the animated character in the video ended up being somewhat 3D as well.
Your experience with Minimax sounds cool! Adding 'animated' to the prompt helping consistency makes sense—AI models often struggle with structure, so any guidance helps.
I mean, unless its a wild magic sci-fi movie, it makes sense that once a character is known in separation to its background, that the background doesn't change with a hands movement of the character. And the character cannot move beyond earthy physics?
Is that what this is?