The model is essentially doing nothing but dreaming.
I suspect that anything that looks like familiar 3D-rendering limitations is probably a result of the training dataset simply containing a lot of actual 3D-rendered content.
We can't tell a model to dream everything except extra fingers, false perspective, and 3D-rendering compromises.