Don’t know the exact details but I imagine the further from the original input image the more the system needs to make up stuff. Same why generative video models are limited to a few seconds. It will improve
We’re just getting started with 3D and incentives for it to improve are strong