Diffusion for World Modeling
diamond-wm.github.io
diamond-wm.github.io
Maybe all of those are clues that parts of the human subconscious mind operate pretty close to the principles behind diffusion models.
Our brains love retroactively altering fuzzy memories.
On LSD, I was swimming in my friend’s pool (for six hours…) amazed at all the patterns on his pool tiles underwater. I couldn’t get enough. Every tile had a different sophisticated pattern.
The next day I went back to his place (sober) and commented on how cool his pool tiles were. He had nfi what I was talking about.
I walk out to the pool and sure enough it’s just a grid of small featureless white tiles. Upon closer inspection they have a slight grain to them. I guess my brain was connecting the dots on the grain and creating patterns.
It was quite a trip to be so wrong about reality.
Not really related to your claim I guess but I haven’t thought of this story in 10 years and don’t want to delete it.
That being said, your reality will influence your dreams if you're exposed to some things enough. I used to play minecraft on a really bad PC back in the day, and in my lucid dreams I used to encounter the same slow chunk loading as I saw in the game.
One of the more annoying parts of growing older was that I stopped lucid dreaming.
The online, continuous and lossy version of this problem is more like how our memory works and still largely unsolved.
I think these models lack a world model, something with strong spatial reasoning and continuity expectations that animals have.
Of course that's probably learned too.
This is what a big tech company was doing in 2015.
The same stuff at industrial scale à la large LLMs would be absolutely mind blowing.
5 million frames of video data with corresponding accelerometer data, and you get this for genuine photorealism.
And I thought I was fairly explicit about video data, but just in case that's ambiguous: the stuff you record with your phone camera set to video mode, synchronised with the accelerometer data instead of player keyboard inputs.
As for output, with the model as it currently stands, I'd expect a 24h training video at 60fps to be "photorealisic and with similar weird hallucinations". Which is still interesting, even without combining this with a control net like Stable Diffusion can do.
You see, there are these things called "proof of concept"s that are meant to not be a product, but instead show off capabilities.
Counterstrike is an example, meant to show off complex capabilities. It is not meant to show how the useful thing of these models is to literally recreate counterstrike.
If that capability is possible, then it could be possible to take 100 examples of seperate world models that exist, and then combine those world models together in interesting ways.
Combining together world models is an obvious next step (IE, not showed off in this proof of concept. But it is a logical/plausible future capability).
Having multiple world models combined together in new and interesting ways, is almost like creating an entirely new world model, even though thats not exactly the same.
I don't have a problem with you using this with a Camera and the real world or your creations ; but I do have a problem when people are able to use someone's work and use these blend and call them original.
It's just as I would have taken a 3d model static mesh, apply some blend on it and call it my own.
No it's f* not.
Thats undecided by the courts. Everyone is training using other people's data right now though, and very few companies are even being sued (let alone have finished the multi-year process to actually be punished for it).
> Or does that mean I should close off all of my creations
When you put something out there, you should expect that other people are going to use those creations. Almost certainly in many ways that you don't approve over.
Don't release something publicly if you don't want other people to use it.
> but I do have a problem
Your anger isn't particularly useful. What are you going to do about it? This particular proof of concept model was made with 1 single GPU running for 2 weeks. You can't stop that.
> It's just as I would have taken a 3d model static mesh, apply some blend on it and call it my own.
Something that I am sure many people are doing all the time, already. Transforming and using other people's content is as old as the internet. AI does little to change that.
We all thought Sports-Game $CURRENT_YEAR or CoD/MW MicroTransaction-Garbage was bottom of the barrel, a whole new world of fresh hell awaits us!
You have it all wrong. It's not going to be AAA studios doing this. Instead, it will be randos who have a cool/fun concept that they want to try out with a couple friends.
Sure, most will be bad. But there might be some gems in there that go viral.
Making it easier for regular people to experiment with games is a good thing.
Or even better, what if this allows people to throw something together to see if a game mechanic is fun, and once it's tested out, then they can make the game for "real".
Allowing quicker experimentation times is also a good thing.
The rate of progress has been mind-blowing indeed.
We sure live in interesting times!
I can how this can already be used to generate realistic physics approximations in a game engine. You create a bunch of snippets of gameplay using a much heavier and realistic physics engine (perhaps even CGI). The model learn to approximate the physics and boom, now you have a lightweight physics engine. Perhaps you can even have several that are specialized (e.g. one for smoke dynamics, one for explosions, ...). Even if it allucinates, wouldn't be worse than the physics bugs that are so common in games.
Also, games introduce randomness in a controlled way so users don't get frustrated by it appearing in unexpected places. I don't want characters to randomly appear and disappear. It's fine if bullet trajectory varies more randomly as they get further away.
People play and enjoy many games with varying levels of randomness as a fundamental component, some even professionally (poker, stock market). This could be made such a game.
Which also means an off-the-shelf deterministic engine.
Wanted to jump on the platform? Wow too bad, model had a little day-dream and now you’re floating somewhere. Wanted to peak a corner? Oops too bad, model had a moment and now the wall curves differently. Trying to pick a shot-well you’d definitely hit the other player if the model wasn’t making physics go out the window. Inconsistently.
Consistent bugs you can anticipate and play/work around, random ones you can’t. Just look at pretty much any speed running community for games before 1995.
Say goodbye to any real competitive scene with random unfixable potentially one off bugs.
Real world usage will be probably different, and maybe even unexpected by the authors of this research.
Or from another angle the end-user is a game developer trying to actually work with this kind of technology, which is just a nightmarish prospect. Nobody in the industry is asking for a game engine that runs entirely on vibes and dream logic, gamedev is already chaotic enough when everything is laid out in explicit code and data.
I’m no fan of academia, but it undeniably produces useful and meaningful knowledge regularly.
Of course they don't. Stuff like this is a proof of concept.
If they had a product that worked, they wouldn't be in academia. Instead, they would leave the world of research and create a multi billion dollar company.
Almost by definition, anything in academia isn't going to be productized, because if it was, then the researchers would just stop researching and make a bunch of money selling the product to consumers.
Such research is still useful for society, though, as it means that someone else can spend the millions and millions of dollars making a better version and then selling that.
I don't know about that. Physic bugs are common, but you can prioritize and fix the worst (gamebreaking) ones. If you have a blackbox model, it becomes much harder to do that.
An obvious next step towards a more playable game is to add state vector to the inputs of the model: it is easier to learn to render the world from pixels + state vectors than from pixels alone.
Then it depends what we want to do. If we want normal Counter Strike gameplay but with new graphics, we can keep existing CS game server and train only the rendering part.
If you want to make Dream-Counter-Strike where rules are more bendable then you might want to train state update model...
lightweight, but producing several hundred watts of heat.
If you look at the metnet papers similar approaches have been successful for weather prediction where models have better results and are order of magnitude faster at inference time compared to numerical models. Of course, training will be time consuming and resources intensive, but it's something you can do once and then ship it to users when they install / update the game.
How would a "function approximation" of Newtonian physics, with billions of parameters, be cheaper to compute?
It seems like this would both be more expensive and less correct than a proper physics simulation.
But it seems extremely promising for fiery explosions, smoke, and especially water. Anything with dynamics that are essentially complex.
Also for lighting -- both to get things like skin right with subsurface scattering, as well as global ray-traced lighting.
You can train specific lightweight models for these things, and they important thing is that their output is correct at the macro level. E.g., a tree should be casting a shadow that looks like the right shadow at the right angle for that type of tree and its types of leaves and general shape. Nobody cares if each individual leaf shadow corresponds to an individual leaf 10 feet above or is just hallucinated.
Does it respects/builds some kind of game map in the process or is it just a bizarre psychedelic dream walk experience where you cannot go back the same place twice and space dimensions are just funny? Is a game map finite?
I'd be interested to see what happens if you look down at your feet for a while, then back up. If the ground looks the same everywhere, do you come up in a random place?
But I am just guessing, and I haven't tried it yet.
This is similar to LLM-based RPGs I've played, where you can pick up a sword and put it in your empty bag, and then pull out a loaf of bread and eat it.
Mondays
https://worldmodels.github.io/
Just want to point that out.
The World Models paper is still one of the most amazing papers I've ever read. And I just really keep wanting to show that, in case people really don't see that, many in-the-know... knew.
So far it's only been used to train a scene from photographs from multiple angles and rebuild it volumetrically by adjusting densities in a point-cloud.
But it might be possible to train a model on multiple different scenes, and perform diffusion on a random point cloud to generate new scenes.
Rendering a point cloud in real time is also very efficient, so it could be used to create insanely realistic game worlds instead of polygonal geometry.
It seems someone already thought of that: https://ar5iv.labs.arxiv.org/html/2311.11221
I was suggesting a more modest approach, I guess, one where the reverse-denoising process involves picking and placing existing 3D assets, e.g., those in GTA 5, so that the process is actually building a plausible map, using those 3D assets, but on the fly...
Turn your car right and a plausible street decorated with buildings, trees and people is dreamt up by the algorithm. All the lighting and physics would still be done in-engine, with stable diffusion acting as a dynamic map creator, with an inherent knowledge of how to decorate a street with a plausible mix of assets.
I suppose it could form the basis of a procedurally generated game world where, given the same random seed, it could generate whole cities or landscapes that would be the same on each player's machine. Just an idea...
Diffusion based generators will do everything soon. And in every style imaginable.
We'll probably solve the energy issue in time.
Not sure how efficient that would be though, and would only work for assets like teapots and whatnot, not whole game maps say.
A point cloud is basically a 3D texture of colors and densities, so a raymarching algorithm can traverse it adding densities it collides with to find the final fragment color. That's how realistic fog and clouds are rendered in games nowadays, and it's very fast, except they use a noise function instead of a scene model.
That's not how I'm familiar with it. As I know it[1], a point cloud is literally that, a collection of individual points, that represents an object scene.
While what you describe is like the scalar field[2] I mentioned, each position in space has some value. You can render them directly like you say, I was thinking to extract geometry a level-set method could be interesting.
It's typically not done at the pixel level, but at the "latent space" level of e.g. a VAE. The image generation is done in this space, which has fewer outputs than the pixels of the final image, and then converted to the pixels using the VAE.
Ah, okay, so the work is done at a different level of abstraction, didn't know that. But I guess it's still a pixel-related abstraction, and it is converted back to pixels to generate the final image?
I suppose in my proposed (and probably implausible) algorithm, that different level of abstraction might be loosely analogous to collections of related game engine assets that are often used together, so that the denoising algorithm might be effectively saying things like "we'll put some building-related assets here-ish, and some park-related flora assets over here...", and then that gets crystallised in to actual placement of individual assets in the post-processing step.
The VAE isn't really pixel-level, it's semantic-level. The most significant bits in the encoding are like "how light or dark is the image" and then towards the other end bits represent more niche things like "if it's an image of a person, make them wear glasses". This is way more efficient than using raw pixels because it's so heavily compressed, there's less data. This was one of the big breakthroughs of stable diffusion compared to previous efforts like disco diffusion that work on the pixel level.
The VAE encodes and decodes images automatically. It's not something that's written, it's trained to understand the semantics of the images in the same way other neural nets are.
Focusing on pixel level generation is the right approach I think. The somewhat noisy output will be improved upon probably in a short timeframe. Now that they proved with Doom (https://gamengen.github.io/) and this that it's possible, probably more research is happening currently to nail the correct architecture to scale this to HD and minimal hallucination. It happened with videos alredy so we should see a similar level breakthrough soon.
For example https://github.com/NVlabs/CTG
Edit: fixed link
Image models are NOT denoised at the pixel level - diffusion happens in latent space. This was one of the big breakthroughs that made all of this work well.
There's a model for encoding/decoding between pixels and latent space. Latent space is able to encode whatever concepts it needs in whichever of its dimensions it needs, and is generally lower dimensional than pixel space. So we get a noisy latent space, denoise it using the diffusion model, then use the other model (variational autoencoder) to decode into pixel space.
It seems decent in short bursts. As it goes on it quite quickly loses detail and the weapon has a tendency to devolve into colorful garbage. I would also like to point out that none of the videos show what happens when you walk into a wall. It doesn't handle it very gracefully.
The lack of temporal consistency will still make it feel pretty dreamlike, but it won't matter that much, because the base is consistent and it will look amazing.
I think the model does not have to know anything about the functionality. It can just dream up what is most probable to happen based on the training data.
Here's a demo from 2021 doing something like that: https://www.youtube.com/watch?v=3rYosbwXm1w
https://www.reddit.com/r/aivideo/comments/1fx6zdr/gta_iv_wit...
I wonder if this type of AI upscaling could eventually also fix things like slightly janky animations, but I guess that would be pretty hard without predetermined input and some form of look ahead.
Limiting character motion to only allow correct, natural movement would introduce a strange kind of input lag.
https://www.nvidia.com/en-us/on-demand/session/gtcspring22-d...
You hit the nail on the head.
Curious, since this is a strong loop old frame + input -> new frame, What happens if a non-CS image is used to start it off? Or a map the model has never seen. Will the model play ball, or will it drift back to known CS maps?
Is that was vision-language models already do? Somehow all of the language should be grounded in the world model. For models like Gemini that can answer questions about video, it must have some level of this grounding already.
I don't understand how this stuff works, but compressing everything to one dimension as in a language model for processing seems inefficient. The reason our language is serial is because we can only make one sound at a time.
But suppose the "game" trained on was a structural engineering tool. The user asks about some scenario for a structure and somehow that language is converted to an input visualization of the "game state". Maybe some constraints to be solved for are encoded also somehow as part of that initial state.
Then when it's solved (by an agent trained through reinforcement learning that uses each dreamed game state as input?), the result "game state" is converted somehow back into language and combined with the original user query to provide an answer.
But if I understand properly, the biggest utility of this is that there is a network that understands how the world works, and that part of the network can be utilized for predicting useful actions or maybe answering questions etc. ?
Alternative as of last year there are now purely diffusion based text decoder models
I'm not sure why the title says it was trained on 2x4090 either as I can't see this on either the linked page or the paper. The paper mentions a GPU year of 4090 compute was used to train the Atari model.
https://github.com/eloialonso/diamond/tree/csgo?tab=readme-o...
The part about the continuous control still seems weird to me though. If anyone understands that then very interested to hear more.
After reading the comments I can assume that if you play outside of the scope it was trained on, the game loses its functionality.
Nevertheless, R&D for a good cause is something we all admire.
I give it the same five years before there are games entirely indistinguishable from reality, and I don’t just mean graphical fidelity - there’s no reason that the same or another model couldn’t provide limitless physics - bust a hole through that wall, set fire to this refrigerator, whatever.
There are fundamental limitations with what are in the end all essentially neural nets; there is no understanding, only prediction. Prediction alone is not enough to emulate reality, which is why for example genuinely self-driving cars have not, and will not, emerge. A fundamental advance in AI technology will be required for that, something which leads to genuine intelligence, and we are no closer to that than ever we were.
Instead of rendering the final shaded and textured pixels, the engine would output just the material IDs, motion vectors, and similar "meta" data that would normally be the inputs into a real-time shader.
The AI can use this as inputs to render a photorealistic output. It can be trained using offline-rendered "ground-truth" raytraced scenes. Potentially, video labelled in a similar way could be used to give it a flair of realism.
This is already what NVIDIA DLSS and similar AI upscaling tech uses. The obvious next step is not just to upscale rendered scenes, but to do the rendering itself.
Have you considered how you'd tell the difference between a prediction and understanding in practice?
I have no idea what this means.
> I have no idea what this means.
You can throw a ball up in the air and predict that it will fall again and bounce. You have no understanding of mass, gravity, acceleration, momentum, impulse, elasticity...
You can press a button that makes an Uber car appear in reality and take you home. You have no understanding of apps, operating systems, radio, internet, roads, wheels, internal combustion engines, driving, GPS, maps...
This confusion of understanding and prediction affects a lot of people who use technology in a "machine-like" way, purely instrumental and utilitarian... "how does this get me what I want immediately?"
You can take any complex reality and deflate it, abstract it, reduce it down to a mere set of predictions that preserve all the utility for a narrow task (in this case visual facsimile) but strip away all depth of meaning. The models, of both the system and the internal working model of the user are flattened. In this sense "AI" is probably the greatest assault on actual knowledge since the book burning under totalitarian regimes of the mid 20th century.
No I would answer that it is indeed understanding, to upend your "guess" (prediction) and so prove that while you think you can "predict" the next answer you lack understanding of what the argument is really about :)
and in case you hadn’t noticed, that kind of uncomprehending slopthink has been going on for a lot longer than the AI fad
Hooke's law was pure curve-fitting. Hooke definitely did not understand the "why". And yet we don't consider that bad physics.
Newton's laws can be derived from curve fitting. How is that different from "understanding"?
People have an intuitive understanding of motion - we see it literally every day, we throw objects, etc.
And yet it took literally thousands of years since discovery of mathematics (geometry, etc.) to formulate a concept of force, momentum, etc.
Ancient Greek mathematicians could do integration, so they were not lacking mathematical sophistication. And yet their understanding of motion was so primitive:
Aristotle, an extremely smart man, was muttering something about "violent" and "natural" motion: https://en.wikipedia.org/wiki/Newton%27s_laws_of_motion#Anti...
People started to understand the conservation of quantity of motion only in 17th century.
So we have two possibilities:
* everyone until 17th century was dumb af (despite being able to do quite impressive calculations)
* scientific discovery is really a heuristic-driven search process where people try various things until they find a good fit
I.e. millions of people were somehow failing to understand motion for literally thousands of years until they collected enough assertions about motion that they were able to formulate the rule of conservation, test it, and confirm it fits. And only then it became understanding.
You can literally see conservation of momentum on a billiard table: you "violently" hit one ball, it hits other balls and they start to move, but slower, etc. So you really transfer something from one ball to the rest. And yet people could not see it for thousands of years.
What this shows is that there's nothing fundamental about understanding: it's just a sense of familiarity, it is a sense that your model fits well. Under the hood it's all prediction and curve fitting.
We literally have prediction hardware in our brains: cerebellum has specialized cells which can predict, e.g. motion. So people with damaged cerebellum have impaired movement: they still can move, but their movement are not precise. When do you think we find specialized understanding cells in the human brain?
Ad-hoc heuristics are not the same thing as understanding. It took formal reasoning for humans to actually understand motion, of a type that modern AI does not use. There is something fundamental about understanding that no amount of familiarity can substitute for. Modern AI can gain enormous amounts of familiarity but still fail to understand, e.g. this Counter-Strike simulator not knowing what happens when the player walks into a wall.
There's no understanding. It's just a formula which matches the observations. It also matches our intuition (a heavier object is hard to move, etc), and you feel this connection as understanding.
Centuries later people found that conservation laws are linked to symmetries. But again, it's not some fundamental truth, it's just a link between two concepts.
LLM can link two concepts too. So why do you believe that LLM cannot understand?
I middle school I did extremely well in physics classes - I could solve complex problems which my classmates couldn't because I could visualize the physical process (e.g. motion of an object) and link that to formulas. This means I understood it, right?
Years later I thought "But what *is* motion, fundamentally?". I grabbed Landau-Lifshitz mechanics textbook. How do they define motion? Apparently, bodies move in a way to minimize some integral. They can derive the rest from it. But it doesn't explain what a motion is. Some of the best physicists in the world cannot define it.
So I don't think there's anything to understanding except feeling of connection between different things. "X is like Y except for Z".
1. Finding a comprehensive explanation
2. Having a comprehensive explanation which is usable
99.999% people on Earth do not discover any new laws, so I don't think you use #1 as a fundamental deficiency of LLMs.
And nobody is saying that just training a LLM produces understanding of new phenomena. That's a strawman.
The thesis is that a more powerful LLM together with more software, more models, etc, can potentially discover something new. That's not observed yet. But I'd say it would be weird if LLM can match capabilities of average folk but never match Newton. It's not like Newton's brain is fundamentally different.
Also worth noting that formulas can be discovered by enumeration. E.g. `m * v` should not be particularly hard to discover. And the fact that it took people centuries implies that that's what happened: people tried different formulas until they found one which works. It doesn't have to be some fancy Newton magic.
An incomplete understanding is no understanding at all, and I would argue that we can only predict, and we can certainly emulate reality, otherwise we would not be able to function within it. A toddler can emulate reality, anticipate causality - and they certainly can’t be said to be in possession of a robust grand unified theory.
Given a model which can generate the game view in ~real time and a model which can generate the models and textures, why would you ever use the first option, apart from a cool tech demo? I'm sure there's space for new dreamy games where invisible space behind you transforms when you turn around, but for other genres... why? Destructible environment has been possible for quite a while, but once you allow that everywhere, you can get games into unplayable state. They need to be designed around that mechanic to work well: Noita, Worms, Teardown, etc. I don't believe the "limitless physics" would matter after a few minutes.
You're basically saying that game development would need to do the work twice: step 1: develop a fully functional game, step 2: spend ridiculous effort (in terms of time and compute) on training a model to emulate the game in a half-baked fashion.
It's a solution looking for a problem.
And this would also be an exciting route to go at remastering old games. I‘d pay a lot to play NFS Porsche again, with photorealism. Or imagine Command & Conquer Red Alert, „rendered“ with such a model.
You can drop in low-res textures and have AI tools upscale them. Models can be replaced, as well as lighting and the best part: it's all under your control. You're not at the merci of obscure training material that might or might not result in a consistent look-and-feel. More knobs, more control, less compute required.
It's the thing Miyazaki is on about with his famous quote.
I think it's possible specific engine components could be ML-driven in the future, like graphics or NPC interactions. This is already happening to a certain degree.
Now, I don't think it's impossible for an ML model to run an entire game. I just don't think making + running your game in a predictive ML model will ever be more effective than making a game the normal way.
If ML has any place in games it's for specific subsystems which don't need absolute precision, NPC behaviour, character animation, refining the output of a renderer, that kind of thing.
(The latest CS removed support for macOS)
when trying to run on a mac it only plays in a very small window, how could this be configured?
The next step to create training data is a real human with a bodycam. There is only the need to connect the real body movement (step forward, turning left, etc) to typical keyboard and mouse game control events, to feed them into the model, too.
I think that is what the devs here are dreaming about.
I love being a horse in the 1900s that automobilewill never take off /s
Games are just an accessible and easy to replicate context to work in, they're not an end goal or target application.
The research is about AI agents interacting with and creating world models. Such world models could just as well be alien environments - i.e. the kind of stuff an interstellar and even interplanetary probe would need to be able to do, as two-way communication over large distances is impractical.
https://notebooklm.google.com/notebook/a240cb12-8ca1-41b4-ab... (7m59s)
As always, it's not actually much of a technical deep dive, but gives a quite decent overview of the pieces involved, and its applications.