MeshGPT: Generating triangle meshes with decoder-only transformers
nihalsid.github.io
nihalsid.github.io
"We first learn a vocabulary of latent quantized embeddings, using graph convolutions, which inform these embeddings of the local mesh geometry and topology. These embeddings are sequenced and decoded into triangles by a decoder, ensuring that they can effectively reconstruct the mesh."
This idea is simply beautiful and so obvious in hindsight.
"To define the tokens to generate, we consider a practical approach to represent a mesh M for autoregressive generation: a sequence of triangles."
More from paper. Just so cool!
Quantized embeddings are just that, but you introduce some discrete structure into the NN, such that the representations there are not continuous. A typical way to do this these days is to learn a codebook VQ-VAE style. Basically, we take some intermediate continuous representation learned in the normal way, and replace it in the forward pass with the nearest "quantized" code from our codebook. It biases the learning since we can't differentiate through it, and we just pretend like we didn't take the quantization step, but it seems to work well. There's a lot more that can be said about why one might want to do this, the value of discrete vs continuous representations, efficiency, modularity, etc...
Conceptually I understand embedding quantization, and I have some hint of why it works for things like WAV2VEC - human phonemes are (somewhat) finite so forcing the representation to be finite makes sense - but I feel like there’s a level of detail that I’m missing regarding whats really going on and when quantisation helps/harms that I haven’t been able to gleam from papers.
But really it's only really useful if you absolutely need to have a discrete embedding space for some sort of downstream usage. VQVAEs can be difficult to get to converge, they have problems stemming from the approximation of the gradient like codebook collapse
The difference from a "normal" convolution is that you can consider arbitrary connectivity of the graph (rather than the usual connectivity induced by a regular Euclidian grid), but the underlying idea is the same: to calculate the result of the operation at any single place (i.e., node), you need to perform a linear operation over that place (i.e., node) and its neighbourhood (i.e., connected nodes), the same way that (e.g.) in a convolutional neural network, you calculate the value of a pixel by considering its value and that of its neighbours, when performing a convolution.
What do I think is really compelling in this field (given that it's my profession)?
This has me star-struck lately -- 3D meshing from a single image, a very large 3D reconstruction model trained on millions of all kinds of 3D models... https://yiconghong.me/LRM/
Do we have strong evidence that other models don't scale or have we just put more time into transformers?
Convolutional resnets look to scale on vision and language: (cv) https://arxiv.org/abs/2301.00808, (cv) https://arxiv.org/abs/2110.00476, (nlp) https://github.com/HazyResearch/safari
MLPs also seem to scale: (cv) https://arxiv.org/abs/2105.01601, (cv) https://arxiv.org/abs/2105.03404
I mean I don't see a strong reason to turn away from attention as well but I also don't think anyone's thrown a billion parameter MLP or Conv model at a problem. We've put a lot of work into attention, transformers, and scaling these. Thousands of papers each year! Definitely don't see that for other architectures. The ResNet Strikes back paper is a great paper for one reason being that it should remind us all to not get lost in the hype and that our advancements are coupled. We learned a lot of training techniques since the original ResNet days and pushing those to ResNets also makes them a lot better and really closes the gaps. At least in vision (where I research). It is easy to railroad in research where we have publish or perish and hype driven reviewing.
A competent modeler can make these types of meshes in under 5 minutes, and you still need to seed the generation with polys.
I imagine the next step will be to have the seed generation controlled by an LLM, and to start adding image models to the autoregressive parts of the architecture.
Then we might see truly mobile game-ready assets!
I don't think this general complaint about AI workflows is that useful. Most people are not a competent <insert job here>. Most people don't know a competent <insert job here> or can't afford to hire one. Even something that takes longer than a professional do at worse quality for many things is better than _nothing_ which is the realistic alternative for most people who would use something like this.
Maybe not to you, but it's useful if you're in these fields professionally, though. The difference between a neat hobbyist toolkit and a professional toolkit has gigantic financial implications, even if the difference is minimal to "most people."
2. those "hobbyists" in all examples are in fact professionals now. That's why they could scale up.
Secondly, the number of hobbyists only matters if you're talking about hobbyists that develop the technology-- not hobbyists that use the technology. Until those tools are good enough, you could have every hobbyist on the planet collectively attempting to make a Disney-quality character model with tools that aren't capable of doing so and it wouldn't get much closer to the requisite result than a single hobbyist doing the same.
May be relevant in the long run, but it'll probably be 5+ years before this is commercially available. And it won't be cheap either, so out of the range of said people who can't hire a competent <insert job here>
That's why a lot of this stuff is pitched to companies with competent people instead of offered as a general product to download.
I think you should look at the progress of image, text, and video generation over the past 12 months and re-asses your timeline prediction.
But this stuff trickles down to the public very slowly. Because indies aren't a good audience to sell what is likely an expensive tech that is focused on mid-large scale production.
>Most people are not a competent <insert job here>. Most people don't know a competent <insert job here> or can't afford to hire one.
emphasis mine. Affordability doesn't have much to do with capabilities, but it is a strong factor to consider for an indie dev. Devs in fields (games, VFX) that don't traditionally pay well to begin with.
>the relevance of your argument to the reality of the scene is not clear.
I feel it's clear if you're following the conversation chain. Is there something you'd like me to clarify?
Practical applications for a medium-large studio is very different from the practical applications of a solo/small dev team. That's all I'm really getting at. There is all kinds of cool tech from 2010 that still isn't viable for indies but is probably used at every AAA studio, so there's precedent I'm basing this on.
I honestly think we'll get there within 18 months.
My skepticism is whether the technique described here will be the basis of what people will be using in ~2 years to replace their low level static 3d asset generation.
There are several techniques out there, leveraging different sources of data right now. This looks like a step in the right direction, but who knows.
At the moment, making 3d models is a lot of skilled, monotonous work, especially for stuff like scene furniture. I guess I'd be pretty happy if some of that work could be automated away, and I'm pretty confident that there's no point automating away the remainder, for the same reason you don't want ChatGPT writing your screenplay.
I just checked the timestamps on my Dall-E Mini generated images. They're dated June 2022
This is what people were doing on commodity hardware back then:
https://cdn-uploads.huggingface.co/production/uploads/165537...
This is what people are doing on commodity hardware now:
https://civitai.com/images/3853761
I'm not even going to try to predict what we'll be able to do in 2 years time; even when accounting for the current GenAI hype/bubble!
Sweet. Can you point me to these modelers who work on-demand and bill for their time in 5 minute increments? I’d love to be able to just pay $1-2 per model and get custom <whatever> dropped into my game when I need it.
Sure, you can probably demo your skills on one such model, but to do it consistently non-stop is a fantasy.
I really believe that the more experienced you are in a particular use case, the more use you can get out of an ML model.
Unfortunately, it's those very same people that seem to be the most resistant to adopting this without really giving it the practice required to get somewhere useful with it. I suppose part of the problem is we expect it to be a magic wand. But it's really just the new PhotoShop, or Blender, or Microsoft Word, or PowerPoint ...
Most people open those apps, click mindleslly for a bit, promptly leave never to return. And so it is with "AI".
The pipeline problem also exists: if you need to still have the skillsets you build up through learning the craft, you still need to have avenues to learn the craft--and the people who already have will get old eventually.
There's a golden path towards a better future for everybody out of this, but a lot of swamps to drive into instead without careful forethought.
Being able to model something - is way different from being able to do it in the least amount of triangles and/or without losing details.
It's not about competent modellers, any more than SD is for expert artists.
It's about giving tools to the non-experts. And also about freeing up those competent modellers to work on more interesting things than the 10,000 chair variants needed for future AAA games. They can work on making unique and interesting characters instead, or novel futuristic models that aren't in the training set and require real imagination combined with their expertise.
Prompt wizard hands work off to the finisher/detailer.
It'll boost productivity and lead to higher quality finished content. And you'll be able to spot when a production - whether video game or movie - lacks a finisher (relying just on generation by prompt). The objects won't have that higher tier level of realism or originality.
Or flipping burgers at McDonald's!
There are only so many games that the market can support, and in those, only so many unique characters[0] that are required. We're pretty much at saturation already.
[0]Not to mention that if AI can generate chairs, from what we have seen from Dall-E & SDXL, it can generatte characters too. Less great than human generated ones? Sure, but it's clear that big boys like Bethesda and Activision do not care.
As they are generated, variations are much easier to come by easier, than buying a couple asset packs.
For context, building from scratch in a 3D pipeline requires you to wear a lot of different hats (modeling, materials, lighting, framing, animating, ect). It costs a lot of time to get to not only learn these hats but also use them together. The individual complexity of those skill sets makes it difficult to experiment and play around, which is how people learn with software.
The shortcut is using premade assets or addons. For instance, being able to use the Source game assets in Source Filmmaker combined with SFM using a familiar game engine makes it easy to build an intuition with the workflow. This makes Source Filmmaker accessible and its why theres so much content out there made with it. So if you have gaps in your skillset or need to save time, you'll buy/use premade assets. This comes at a cost of control, but that's always been the tradeoff between building what you want and building with what you have.
Just like GPT and DALL-E built a bridge between building what you want and building with what you have, a high fidelity GPT for the 3D pipeline would make that world so much more accessible and would bring the kind of attention NLE video editing got in the post-Youtube world. If I could describe in text and/or generate an image of a scene I want and have a GPT create the objects, model them, generate textures, and place them in the scene, I could suddenly just open blender, describe a scene, and just experimenting with shooting in it, as if I was playing in a sandbox FPS game.
I'm not sure if MeshGPT is the ChatGPT of the 3D pipeline, but I do think this is kind of content generation is the conduit for the DALL-E of video that so many people are terrified and/or excited for.
My wife is passionate about film/TV production and VFX.
She's currently in school for this but is concerned about the difficulty of landing a job afterwards.
Do you have any recommendations on breaking into the industry without work experience?
I think producer roles are a little bit less ultra competitive / scarce as they are actually jobs jobs where you have to use excel and planning and budgeting.
Being a producer means being on the phone all the time, negotiating, haggling, finding solutions where they don’t seem to exist.
Be it in TV, advertising or somewhere in the media space, the common rule is that producers are mostly actually terrible at their jobs, that’s my experience in London. So if she’s really good and really dedicated and learns the job of everyone on set, I’d say she has a shot.
The real secret to being good in filmmaking is learning everyone else’s job. Toyota Production System says if you want to run a production line you have to know how it works.
If she wants to do VFX production she could start doing her own test scenes, learning basics in nuke and Blender, even understanding the role of Houdini and how that works.
If she does that - any company will be lucky to have her.
Another massive benefit is composability. If the model can generate a cup and a table, it also knows how to generate a cup on a table.
Think of all the complex gears and machine parts this could generate in the blink of an eye, while being relevant to the project - rotated and positioned exact where you want it. Very similar to how GitHub Copilot works.
We're safe for now but we should learn how to leverage the new tech.
Some are safe for several years (3-5), that's it. During that time it's going to wreck the bottom tiers of employees and progressively move up the ladder.
GPT and the equivalent will be extraordinary at programming five years out. It will end up being a trivially easy task for AI in hindsight (15-20 years out), not a difficult task.
Have you seen how far things like MidJourney, Dalle, Stable Diffusion have come in just a year or two? It's moving extremely fast. They've gone from generating stick figures to realistic photographs in two years.
Being exceptional at programming isn't hard. Being exceptional at listening to bullshit requirements for 5 hours a day is.
Doesn’t apply too much to mesh generation but was certainly the case in image gen. Mistakes that wouldn’t fly for a human artist (hands) were just accepted as part of AIgen.
So these areas are much less strict about precision than coding. Making these tools much more capable are replacing artists in some tasks than CoPilot is for coders atm.
IMO the last time a major tech advance was visible was Davy Jones on the Pirates films. That was a fully photorealistic animated character that was plausible as a hero character in a major feature. That was a breakthrough. After that a lot of refinement and speeding up.
This is different. I have some positivity about it, but it’s getting hard to keep track of everything that’s going on tbh. Every week it’s a new application and every few months it’s some quantum leap.
Like others said, Midjourney and DallE are essentially photorealistic.
It seems to me that the next step is generative AI creating better and better assets.
And then of course you have video generation which is happening as well…
What you want for asset creation is not photorealism, but style and concept transfer, multimodal controllability (text alone is terrible at expressing artistic intent), and tooling. And tooling isn't something that is developed quickly (although there were several rapid breakthroughs in the past, for example ZBrush).
Most of the fancy demos you hear about sound good on paper, but don't really go anywhere. Academia is throwing shit at the wall to see what sticks, this is its purpose, especially when practice is running ahead of theory. It's similar to building airplanes before figuring out aerodynamics (which happened long ago): watching a heavier-than-air thing fly is amazing, until you realize it's not very practical in the current form, or might even kill its brave inventor who tried to fly it.
If you look at the field closely, most of the progress in visual generative tooling happens in the open source community; people are trying to figure out what works in real use and what doesn't. Little is being done in big houses, at least publicly and for now, as they're more interested in a DC-3 than a Caproni Ca.60. The change is really incremental and gradual, similarly to the current mature state of 3D. Paradigms are different but they are both highly technical and depend on academic progress. Once it matures, it's going to become another skill-demanding field.
The idea that somehow “AI isn’t art directable” is one I keep hearing, but I remain unconvinced this is somehow an unsolvable problem.
The idea that AIgen is unusable at the moment for professional work doesn’t hold up to my experience since I now regularly use Photoshop’s gen feature.
I can keep writing prompts for DE3 or similar until it gives me something like what I want, but the problem is, there are often subtle but important mistakes in many images that are generated.
I think it's really good at portraits of people, but for anything requiring complex lighting, representation of real world situations or events, I don't think it's ready yet, unless we're ready to just write prompts, click buttons and just accept what we receive in return.
Midjourney already has tools that allow you to select parts of the image to regenerate with new prompts, Photoshop-style. The tools are being built, even if a bit slowly, to make these things useful.
I could totally see creating Matte paintings through Midjourney for indie filmmaking soon, and for tiny budget films using a video generative tool to make let’s say zombies in the distance seems within reach now or very soon. Slowly for some kind of VFX I think AI will start being able to replace the human element.
>The idea that somehow “AI isn’t art directable” is one I keep hearing, but I remain unconvinced this is somehow an unsolvable problem.
That's not my point. AI can be perfectly directable and usable, just not in the specific form DE3/MJ do it. Text prompts alone don't have enough semantic capacity to guide it for useful purposes, and the tools they have (img2img, basic in/outpainting) aren't enough for production.
In contrast, Stable Diffusion has a myriad of non-textual tools around it right now - style/concept/object transfer of all sorts, live painting, skeleton-based character posing, neural rendering, conceptual sliders that can be created at will, lighting control, video rotoscoping, etc. And plugins for existing digital painting and 3D software leveraging all this witchcraft.
All this is extremely experimental and janky right now. It will be figured out in the upcoming years, though. (if only community's brains weren't deep fried by porn...) This is exactly the sort of tooling the industry needs to get shit done.
edit: Seems like mesh completion is the main input-output method, not just a neat feature.
NO. Details are mostly like icing on top of the cake. Sure, good details make good art but it is not always the case. True and beautiful art requires form + shape. What you are saying is something visually appealing. So, the reason why diffusion models feel so bland is because they are good with details but do not have precise forms and shape. Nowadays they are getting better, however, it still remains an issue.
Form + shape > details is something they teach in Art 101.
> Inspired by recent advances in powerful large language models, we adopt a sequence-based approach to autoregressively generate triangle meshes as sequences of triangles.
It's only inspired by LLMs
LLMs are autoregressive sequence models where the "role" of the graph convolutional encoder here is filled by a BPE tokenizer (also a learned model, just a much simpler one than the model used here). That this works implies that you can probably port this idea to other domains by designing clever codecs which map their feature space into discrete token sequences, similarly.
(Everything is feature engineering if you squint hard enough.)
It looks like the input is itself a 3D mesh? So the model is doing "shape completion" (e.g. they show generating a chair from just some legs)... or possibly generating "variations" when the input shape is more complete?
But I guess it's a starting point... maybe you could use another model that does worse quality text-to-mesh as the input and get something more crisp and coherent from this one.
Diffusion models could be used to generate textures.
Mark is right and so so early.
edit: Oh, _that_ Mark? lol okay
edit edit: Maybe credit Lecun or something? Mark going all in on the metaverse was definitely not because he somehow predicted deep learning would take off. Even the people who trained the earliest models weren't sure how well it would work.
Another application: take the output of your gaussian splatter or diffusion model and run it through MeshGPT. Instant usable assets with clean topology from text.
I don't think people here realize how are we inching to automating the automation itself, and the programmers who will be able to make a living out of this will be a tiny fraction of those who can make a living out of it today.
So much more refreshing than the dense abstract, intro, results paper style.
Indie games already seems pretty derivative these days. I think this tech will kill them in mid-term as big companies use them.
User name checks out.
It feels like a countdown until every creative in the videogame industry is automated.
It’s no different than saying “these home kitchen appliances are really gonna kill off the restaurant industry.”
AI creating video games would drastically increase the volume of games available in the market. This surge in supply could make it harder for indie games to stand out, especially if AI-generated games are of high quality or novelty. It could also lead to even more indie saturation( the average indie makes less than 1000 dollars).
As the market expectations shift, I think most indie development dies unless you are already rich or basically have patronage from rich clients.
The games market has been in the same place as the rest of the arts for some time now: if you want to be noticed, you have to mount a bit of a production around it, add layers of design effort, and find a marketing funnel for that particular audience. The days of just making a Pong clone passed in the 1970's.
What technology has done to the arts, historically, is add either more precision or more repeatability. The relationship to production and arts as a business maps to what kinds of capital-and-labor-intensive endeavors leverage the tech.
Photographs didn't end painting, they ended painting as the ideal of precisely representational art. In the classical era, just before the tech was good enough to switch, painting was a process of carefully staging a scene with actors and sketching it using a camera obscura to trace details, then transferring the result to your canvas. Afterwards, the exact scene could be generated precisely in a photo, and so a more candid, informal method became possible both through using photographs directly and using them as reference. As well, exact copies of photographs could be manufactured. What changed was that you had a repeatable way of getting a precise result, and so getting the precision or the product itself became uninteresting. But what happened next was that movies and comics were invented, and they brought us back to a place of needing production: staged scenes, large quantities of film or illustration, etc.
With generative AI, you are getting a clip art tool - a highly repeatable way of getting a generic result. If you want the design to be specific, you still have to stage it with a photograph, model it as a scene, or draw it yourself using illustration techniques.
And so the next step in the marketplace is simply in finding the approach to a production that will be differentiating with AI - the equivalent of movies to photography. This collapses not the indie space - because they never could afford productions to begin with - but existing modes of mobile gaming, because they were leveraging the old production framework. Nobody has need of microtransaction cosmetics if they can generate the look they want.
The chaining of various AI's and the feedback loops between are accelerating far beyond what people think it is.
Just yesterday major breakthroughs were released on stable diffusion video. It's the pace and categorical type of these breakthroughs that represent a paradigm shift, never seen before in the creative fields.
pacman was recreated just AI
Far more people prefer playing games than making them.
We'll probably see a new boom of indie games instead. Don't forget, a large part of what makes the gaming experience unique is the narrative elements, gameplay, and aesthetics - none of which are easily replaceable.
This empowers indie studios to hit a faster pace on one of the most painful areas of indie game dev: asset generation (or at least for me as a solo dev hobbyist).
The ai game framework handles the full game creation pipeline.
It's really going to speed up my pipeline to not have to pipe all of my meshes into a procgen library with a million little mesh modifiers hooked up to drivers. Instead, I can just pop all of my meshes into a folder, train the network on them, and then start asking it for other stuff in that style, knowing that I won't have to re-topo or otherwise screw with the stuff it makes, unless I'm looking for more creative influence.
Of course, until it's all the way to that point, I'm still better served by the procgen; but I'm very excited by how quickly this is coming together! Hopefully by next year's Unreal showcase, they'll be talking about their new "Asset Generator" feature.
We can create radiance fields with photogrammetry, but IMO we need much better algorithms for transforming these into high quality triangle meshes that are usable in lower triangle budget media like games.
I suppose though there is a case for AI models for example doing what nanite does entirely algorithmically and research like this paper may come in handy there.
Whilst it might not be the solution I'm waiting for, I can now see it as possible. If an AI model can handle traingles, it might handle edge loops and NURBS curves.
What I really appreciate about this is that they took the concept (transformers) and applied it in a quite different-from-usual domain. Thinking outside of the (triangulated) box!
So maybe the next step is something like CLIP, but for meshes? CLuMP?