Anyone that can read English could do that though.
Anyone that can read English could do that though.
You.. literally can? I have no idea what 90% of the people here are saying, it's like they've never even used one of these models before.
Will it effectively create an internal model describing world objects and how they interact with each other, persist that so it doesn't get lost when it's context window gets filled up, then after it has sufficiently complete knowledge of the fundamentals after the tutorial levels successfully apply that model by making plans to solve the puzzles and execute them by clicking the right coordinates tied to the visual feedback?
I highly doubt it. To me it often just looks like people are defining narrow search spaces (e.g by having all of the task complexity pre-digested by the harness design), pointing a brute force engine at them, spending 20 thousand dollars in compute and then saying "hey look, it can do anything!".
When we access the API, we don't get to train the model, we just do inference on the already trained model.
It’s an interesting challenge though. I might start to tackle it by having the model write its own tool program(s) to play the game. It’s possible that the model could choose that strategy itself from a high level prompt alone.
1. It’s too slow for real-time games. To play mario, you’d need to step frame by frame like a TAS. I don’t know if Gruntz has real-time elements or not.
2. It will be expensive. You won’t get very far with a Plus subscription.
The models likely already have some knowledge on game objectives unless the game is really obscure, so it should do a decent job. It can figure out details of the mechanics along the way.