LLMs are fundamentally one dimensional which works fine when you're generating next tokens for text which because that's a 1D problem.
I do wonder how much progress we could make on a problem like this with a 3D transformer architecture.
I do wonder how much progress we could make on a problem like this with a 3D transformer architecture.
Even if you could do what you're suggesting with an LLM (I have my doubts) this result would be a mesh or 3D pixel grid or something, yes?
This is terrible for interoperability and it's the opposite of what mainstream CAD packages do.