Point-E: Point cloud diffusion for 3D model synthesis
github.com
github.com
Seems super fast, some are saying 600x faster [0], than than the version made off of Google's paper. But it is a little less accurate. Point clouds are less useful but some on Reddit and the authors have tools to try to convert to meshes [1][2]. It does feel like stable diffusion level generation of good 3d assets is right around the corner. It will be interesting to see which tech wins out, whether it's some variant of depth estimation like sd2 and non ai tools can do, object spinning/multi angle view like Google's tool does, or whatever this tool does.
[0] https://twitter.com/DrJimFan/status/1605175485897625602?t=H_...
[1] https://www.reddit.com/r/StableDiffusion/comments/zqq1ha/ope...
[2] https://github.com/openai/point-e/blob/main/point_e/examples...
It’s a fun demo. Worth to note that on mobile it didn’t include any button to download the generated point cloud data itself, at least not that I could find. Might be the same on desktop also.
Additionally I think the amount of time taken depends on the amount of visitors. I had to wait about 7 minutes for it to finish.
From [1]. Seems like there is a pattern of "AI asked to generate final results with only final results to learn from, immediately asked for the apple in the picture" in AI generators. I suppose lack of specialization in application domains of NNs is a deliberate design choice for these high-profile projects, in a vague hope of simulating emergent behaviors as seen in the nature and avoiding to be another expert system(while being one!), but that attitude seems limiting usefulness, here and again.
People developing these models are very aware of what 3D workflow is like.
The issue is that image->point cloud training data is very easy to get, whereas image or point cloud -> clean 3d mesh training data is very hard to get in unconstrained domains.
Generating point clouds is where the state of the art is now. That doesn't mean that the whole field isn't entirely aware that text->3d mesh unlocks many more capabilities.
Without knowing a dang thing about AI, it feels like the problem moreso lies in:
1. Math related to topology: vertices, faces, edges, tri vs quad etc
2. Different topologies for the same object are better for different use cases. Rendering, skinning, morphing, physics etc all have different optimal topologies, and the definition of optimal varies based on workflow and scene specifics or even the human who has skills based on certain topological preferences. In other words, I'm not sure how much of 3D workflows are standardized even -- getting the topological data for workflows is no easy task, and it's not super usable until the model output can plug right into a workflow and the existing DCC ecosystem.
text2img generates a static asset, text2mesh is far more interesting beyond just the static rendering part which is where mesh topology becomes a big sticking point.
I believe there are three problems:
* There isn't software that generates point clouds from video games. This should be solvable but AFAIK hasn't been done yet.
* The diversity of models in video games is much lower than the real world
* Games use a bunch of techniques to reduce the poly count while making assets look like they are high poly (eg texture mapping). It's unclear what should be generated here.
Take a look at the field of NeRFs (Neural Radiance Fields: https://datagen.tech/guides/synthetic-data/neural-radiance-f...) for more on this subject.
As always this goes to show that if you can't be the first, be the loudest. OpenAI has the most well oiled media machine I've seen in awhile.
Seeing the waves of publicity OpenAI gets with every new release, I think we're seeing a new model for big-tech AI research groups. It isn't enough to just hire world-class research talent that publish area-defining papers. There has to be a commensurate investment in media to publicize the research. Obviously, if you don't have the research, you have nothing to market. But it should say something that OpenAI prioritizes great design, communication, and publicity in addition to the world-class research team. It wouldn't surprise me if we see Google AI / DeepMind / FAIR / double-down with their own investments to expand the media presence of their AI orgs.
https://arxiv.org/abs/2212.08751 (via https://news.ycombinator.com/item?id=34060986)
https://techcrunch.com/2022/12/20/openai-releases-point-e-an... (via https://news.ycombinator.com/item?id=34069231)
https://twitter.com/drjimfan/status/1605175485897625602 (via https://news.ycombinator.com/item?id=34068271)
(but no meaningful comments at those other threads)
3D sensors are slowly but surely becoming more common. The iPhone Pro series has one, and AR hardware designs tend to include these capabilities. So this model synthesis seems a bit ahead of the curve, in a good way.
I think the premise of this is text-to-3D, and that because it's quicker generations you don't really need anything besides a GPU to start playing around with it.
Please correct me if I am wrong.
The pointcloudtomesh notebook seems to be be able to output something could be converted for 3d printing purposes.
I haven’t yet attempted to do so, but that does seem like an exciting and general purpose use case.
We made some midjourney lamps—and then printed them! Pretty cool.
I found out about Kaedim a few weeks ago and when I saw this repo, it came to my mind as well.
Specifically for guiding generation of the mesh from a possibly AI-generated point cloud (PTC). E.g. using manual contraints on an mostly automatic quad (re-)mesher ran as a post process on the triangle soup obtained from meshing the original, AI-generated PTC.
I.e.:
1. AI-generate PTC from image(s).
2. Auto-generate triangle mesh via marching cubes or whatever from PTC.
3. Quad re-mesh with mesh-guided automatic constraint discovery (think edges, corners etc.).
4. Manual edit quad-mesher constraints.
5. Quad re-mesh.
That would explain their pricing which seems a tad too high for a fully automatic solution. $600 for 30 models x 10 iterations. I.e. each iteration would cost $2.
Or maybe it's just so niche this is simply because of number of users for now and indeed fully automatic.
Curious to hear what other people involved in 3D and cloud compute think.
[1] https://www.youtube.com/watch?v=dnwDPLLzzEUI wonder how much of the research paper is written by ChatGPT.