You just don't like AI.
It can be used for training agents, prototyping, video generation, and is quite possibly a glimpse of a whole new type of entertainment or a new way to create video games.
What's the point of the massive amount of money spent on video games in general? Or all of the energy spent moving people back and forth to an office? Or expensive meals at restaurants? Or trillions in weaponry? Or television shows or movies?
I feel like sharing early closed-source blog-posts is part of the research process. I'm sure someone in this thread has thought of a use case that the Google team missed. Open/closed source arguments here feel premature IMO.
This is just a marketing fluff piece that does not benefit anyone and is ego stroking at best.
I still think things like this are important, and at least give folks a bit of time to ideate on what will be possible in a few years. Of course having the model or architecture on hand would be nice, but I'm not holding that against Google here.
Noone knows yet. AI technology like this is closer to scientific research than it is to product development. AI is basically new magic, and people are in a "discovery" phase where we are still trying to figure out what is possible. Nothing of value was immediately created when they discovered DNA. Productization came much later when it was combined with other technologies to fit a particular use case.
To have this perspective you must believe that this will never get better than it currently is, its limitations will never be fixed, and it will never lead to any other applications. I don't know how people can continue to look at these things with such a lack of imagination given the pace of progress in the field.
You could use real video games to do this but I guess there'd be a risk of over-fitting; maybe it would learn too precisely what a staircase looks like in Minecraft, but fail to generalize that to the staircase in your home. If they can simulate dream worlds (as well as, presumably, worlds from real photos), then they can train their agents this way.
This would only be training high-level decision policies (ie, WASD inputs). For something like a robot, lower level motor control loops would still be needed to execute those commands.
Of course you could just do your training in the real world directly, because it already has coherence and plenty of environmental variety. But the learning process involves lots of learning from failure, and that would probably be even more expensive than this expensive simulator.
Despite the claims I don't think it does much to help with AI safety. It can help avoid hilarious disasters of an AI-in-training crashing a speedboat onto the riverbank, but I don't think there's much here that helps with the deeper problems of value-alignment. This also seems like an effective way to train robo-killbots who perceive the world as a dreamlike first-person shooter.
The major difference being the former scales very poorly for generating training data compared to the latter. Genie 2 is not even a video game and has worse fidelity that video games, the upside is it probably scales even better than video games for generating training scenarios. If you want androids in teal life, Genie 2 (or similar systems) is how you bootstrap the agent AI. The training pipeline will be: raw video -> Genie 2 -> game engine with rules -> physical robot
One of those arrows is not like the others
Any model would have to succeed in one stage before it can proceed to the next one.
Self driving vars have cameras as part of their sensor suite, and have models to make sense of sensor data. Video will help with perception and classification (understanding the world) with no agency needed. Game-playing will help with planning, execution, and evaluation. Both functions are necessary, and those that come after rely on earlier capabilities
For example, let's say that the car is approaching an intersection, and suddenly sees a puddle on the road to the left getting brighter - a visual world model like this might extrapolate a scenario that the brightness is the result of a car moving towards the intersection assigning this some probability, and signing another probably to a scenario that it's just a flickering headlight, and the car would then decide whether and how much to slow down.
In this example there is a sensor, but it definitely doesn't tell the robot "exactly what is there", and while we could try to write rules about what it should do, the Bitter Lesson tells us it's better to just let it create its own model.
Page 8 of the Genie 1 paper: https://arxiv.org/abs/2402.15391
That's nice. I'm not completely disabled, but I am disabled, and I very much would appreciate them, as my capability to do things over the longer term is very much not going to go in the direction of improving. As it is, there are a lot of things I now rely on people for, that at one time, I did not.
Whilst I recognise its probably not going to happen in a time span that is useful to me, I do wish it could, so that I could be less of a burden on those around me, and maintain a relative level of independence.
Unless you have a young/quick death, there's a really good chance you will be, too.
Prompt: Here's a blueprint of my new house and a photo of my existing furniture. Show me some interior design options.
Dreaming and sleeping is incredibly expensive, we spend 33% of our "availability" on average asleep.
This kind of work is a step toward building similar tools for general AI agents (IMHO).
All motivation to make further games is removed because now "somebody" can spool out a 3D adventure game instantly with a line of text. It implies you'll waste a year of your time, and just before release, out pops the dramatically better AI product to steal away all further business and reduce all time you've spent to meaningless. Everybody then waits indefinitely for the "endgame gear." https://xkcd.com/989/
Exactly like LLMs and image generators almost completely took away all business for normal writers and normal painters, because now all managers want is the AI. Now there's endless articles about how "somebody" prefers "AI" for every task. Now the market won't invest in anything unless it has "AI" in the name. Now people idiotically add "AI" to everything just to have the investment.
asked the same thing a while back, and the answers boiled down to "somehow helps RL agents train". but how exactly? no clue
[edited out some barbs I wrote because I find some comments on this website REALLY annoying]