The Elite Four/Champion was a non-issue in comparison especially when you have a lv. 81 Blastoise.
See: https://www.lesswrong.com/posts/7mqp8uRnnPdbBzJZE/is-gemini-...
[1] https://community.aws/content/2gbBSofaMK7IDUev2wcUbqQXTK6/ca...
Claude got stuck reasoning its way through one of the more complex puzzle areas. Gemini took a while on it also, but made it through. I don't that difference can be fully attributed up to the harnesses.
Obviously, the best thing to do would be to run a SxS in the same harness of the two models. Maybe that will happen?
Basically, the gane being conpleted by gemini was in an inferior category (however minuscule) of experiment.
I get it though. People demanded these types of changes in the CPP twitch chat, because the pain of watching the model fail in slow motion is simply too much.
these models are trained on a static task, text generation, which is to say the state they are operating in does not change as they operate. but now that they are out we are implicitly demanding they do dynamic tasks like coding, navigation, operating in a market, or playing games. this are tasks where your state changes as you operate
an example would be that as these models predict the next word, the ground truth of any further words doesnt change. if it misinterprets the word bank in the sentence "i went to the bank" as a river bank rather than a financial bank, the later ground truth wont change, if it was talking about the visit to the financial bank before, it will still be talking about that regardless of the model's misinterpretation. But if a model takes a wrong turn on the road, or makes a weird buy in the stock market, the environment will react and change and suddenly, what it should have done as the n+1th move before isnt the right move anymore, it needs to figure out a route of the freeway first, or deal with the FOMO bullrush it caused by mistakenly buying alot of stock
we need to push against these limits to set the stage for the next evolution of AI, RL based models that are trained in dynamic reactive environments in the first place
llm trained to do few step thing. pokemon test whether llm can do many step thing. many step thing very important.
Successfully navigating through Pokemon to accomplish a goal (beating the game) requires a completely different approach, one that much more accurately mirrors the way you navigate and goal set in real world environments. That's why it's an important and interesting test of AI performance.
Pokemon is interesting because it's a test of whether these models can solve long time horizon tasks.
That's it.
Basically, the model has to keep some notes about its overall goals and current progress. Then the context window has to be seeded with the relevant sections from these notes to accomplish sub goals that help with the completion of the overall goal (beat the game).
The interesting part here is whether the models can even do this. A single context window isn't even close to sufficient to store all the things the model has done to drive the next action, so you have to figure out alternate methods and see if the model itself is smart enough to maintain coherency using those methods.
My man, ChatGPT is the sixth most visited website in the world right now.
What counts as a killer app to you? Can you name one?
a bunch of people think that something like chatgpt is a killer app, and they know it when they see it. you assert that it obviously is not, so clearly the above intuition isn't working for the purposes of discussion.
instead, someone should define the term so that we know what we're talking about, and i offer you the ability to do it so that the frame of the discussion can be favorable to your point of view. but you are also not willing to do that, so how do you expect to convince anyone of your viewpoint?
It is a dismissive rhetorical device to prove a wrong point on an internet forum such as this that has nothing to do with reality.