It can learn anything you have data for.
Heck, we do it with geospatial data already, generating segmentation vectors. Why not 3D?
It can learn anything you have data for.
Heck, we do it with geospatial data already, generating segmentation vectors. Why not 3D?
Not in theory, but the level of complexity is way higher and the amount of data available is much smaller.
Compare bitmaps to this: https://fossies.org/linux/blender/doc/blender_file_format/my...
But we haven't gotten diffusion working well for text/code, so generating long files is a problem.
I'm not experienced enough to validate their claims, but I love the choice of languages to evaluate on:
> Python, Bash and Excel conditional formatting rules.
A 3D scene is vastly more complex, and the way you consume it is tangential to the rendering of it we use to interpret. It is a collection of arbitrary data structures.
We’ll need a new approach for this kind of problem
> A 3D scene is vastly more complex
3D scenes, in fact, are also data, numbers and tokens. (Well, numbers, but so are tokens.)
Not at all the same as fixed sized arrays representing images.
In fact, the data structures of a 3D scene can be serialized as text, and a properly trained text gen system could generate such a representation directly, though that's probably not the best route to decent text-to-3d.
Serializing 3D models as text is not going to work for negligibly non trivial circumstances.