Is 3d a different problem, or a similar one but considerably harder? I'd expect the data encoding (vertices vs pixels) to change a bit about it but I'm not familiar enough to know.
Secondly, there's vastly more labeled image data in the world than 3D data, so creating a CLMP (contrastive language and mesh pairing) model is harder.
It's very late but I may be able to give a much better answer on more of the nuances of 3D generation tomorrow.
I can imagine you'd have the problem of stray floating voxels then, which isn't as noticeable when it happens with 2D pixels.