What if we set GPT-4 free in Minecraft?
twitter.com
twitter.com
>You are a helpful assistant that tells me the next immediate task to do in Minecraft. My ultimate goal is to discover as many diverse things as possible, accomplish as many diverse tasks as possible and become the best Minecraft player in the world.
>8) Tasks that require information beyond the player's status to verify should be avoided. For instance, "Placing 4 torches" and "Dig a 2x1x2 hole" are not ideal since they require visual confirmation from the screen. All the placing, building, planting, and trading tasks should be avoided. Do not propose task starting with these keywords
>7) Use `exploreUntil(bot, direction, maxDistance, callback)` when you cannot find something. You should frequently call this before mining blocks or killing mobs. You should select a direction at random every time instead of constantly using (1, 0, 1).
>9) Do not write infinite loops or recursive functions.
You can really imagine the sorts of pitfalls the agent fell into that induced the authors to add these stipulations.
https://github.com/MineDojo/Voyager/blob/main/voyager/contro...
This function keeps a global count of how many times it's been called to mine a block that doesn't exist nearby, and will warn the bot that it should try exploring first instead.
There's nothing necessarily wrong with all that - it's an important research question to understand how much hand-holding the agent needs to be able to do these sorts of tasks. But readers should be aware that it's hardly dropping an agent into the world and leaving it to its own devices.
did GPT-4 just solve the halting problem?
IMO, in biological systems the explore vs exploit tradeoff is pretty analogous to the halting problem. There doesn’t seem to be an optimal general “solution” to it.
Once an organism is very familiar with its environment it’ll approximate near-optimal trade offs, which would suggest it’s just using heuristics.
For example, it’s trivial to write a program that searches by brute force for a cycle that would disprove the Collatz conjecture, and halts if it finds one. No human knows whether that program will halt.
Write a program that, for a positive integer, runs the Collatz process on it (if even: halve it; if odd, multiply by three and add one; repeat).
If the process results in a 1, move on to the next number and repeat.
If the process produces a number it produced previously, halt. (Worried you’ll need unbounded state for this part? Use tortoise+hare, it’s fine).
This program halts if and only if the Collatz conjecture is false.
Now, the Collatz conjecture might not be unprovable. But for now no human can tell you whether or not that program halts.
I ended up using it to compare single core performance on any windows machine, because the timestamped logging was deterministic. Rewrote it in python and still use it these days.
That’s a common misconception. The brain is not like a computer. The brain can’t store and execute programs.
Can't it? The brain is absolutely capable of emulating a turing complete system such as running a program.
I mean, sure you can reason about a portion of code on your screen but there is no way you could emulate anything without some visual support.
There is no way your short term memory could store the program and the variables. The human brain can barely remember 5 to 10 words for several seconds and you have no control on your long term memory.
So yes your brain can somehow emulate a computer if you give him a pen, a sheet of paper paper, time, and a lot of sugar. But that’s not because it’s functioning like a computer but rather because you learnt how a computer work.
if(going_to_halt)
dont();And obviously trivial to avoid recursion too.
As soon as you add a 2nd layer of loops however, you reach Turing-completeness and suddenly the halting problem becomes unsolvable.
-------
So you don't need to deny jmp/loops. You just need to deny _nested_ loops. And... find that old paper I read like 15 years ago to figure out the details to discriminate halt/not-halt in the single-layer loop language.
while (state !== STOP):
read = TAPE[index]
[state, write, dir] = TABLE[read, state]
TAPE[index] = write
index = index + dirE.g. this Ruby program is undecidable:
gets
This one using a hypothetical gets guaranteed to eventually halt is still undecidable: while getsWithTimeout != "halt"
endFor (;;) { print(“hahaha you didn’t say the magic word!”); }
In another loop would also prevent termination without program shutdowns.
To clarify, when I said "for loop", I meant what is sometimes called a "counted for loop" (or simply "counted loop"): there is an (maximum, if you allow early exit) iteration count that is computed before executing the loop and can't increase later.
In C syntax, it is for (int i = 0, e = ...; i < e; ++i) { ... } and the body of the loop is not allowed to change the value of either i or e.
Edit: actually I may have been unclear. When I said "for loops are guaranteed to terminate", in the context of the discussion, I meant "if the only kind of loops you allow are (counted) for loops in a language where loop-free expressions are guaranteed to terminate, you get a language where all expressions are guarantee to terminate". So loops can contain other loops, as long as they are all of the "counted for loop" kind.
The Halting Problem is a truly interesting result, and for the most part uninteresting in practice.
while (isFamousMathProblemWeDontKnowWhetherItHoldsForAllTheIntegers(n)) { n++ }
Say your Turing machine is searching for a contradiction in ZFC. It enumerates all the possible proofs, and checks if they are valid and they prove 0=1. You can prove in ZFC that you can't prove that it halts, nor you can prove that it doesn't halt.
Now your Turing machine won't halt, according to ZFC. (It can halt in practice, if ZFC has a contradiction.) Same for any other undecidable programs.
GGP was talking about programs halting, not halting, and "something else", which GGP called "undecidable".
There are programs that ZFC cannot prove if they halt or not. For example, searching for a contradiction in ZFS is such a program. Its undecidability means that there is no sequence of axioms of ZFC that ends with "this program halts" or "this program does not halt".
But then this program does not halt, since its halting would mean that ZFC can prove that this program halts.
In any case, programs can halt or not halt. There are programs that ZFC cannot prove that they halt or not, but those programs do not halt.
That step is not logically consistent. It could halt, it's just that ZFC can't prove it will ever do so. For example, the program which computes the 8000th busy beaver number halts (per definition of the busy beaver function), but is undecidable in ZFC[1].
An undecidable program can halt. An undecidable program can run forever. But whatever axiom system is being used to decide that can't prove which (without running the program, potentially for an infinite number of steps).
I don't believe that there exist such a program.
What I don't believe that there exist a program that is a proper counter-example, that is, its halting is undecidable in ZFC, and it halts. Exactly because what I wrote earlier.
The problem is that you dont know if the checker that you use to detect if a program halts will itself halt on your given input or continue forever. But given a sound checker with some assumption, one can find non-halting programs which wont be detected by the checker using the diagonalization trick.
*: Trivial in a mathematician's sense
The Minecraft videos are impressive.
Nethack (https://www.nethack.org/) has been used for AI development in the past and more recently:
http://shelf2.library.cmu.edu/Tech/9997774.pdf
https://portfolios.cs.earlham.edu/wp-content/uploads/2018/12...
https://arxiv.org/abs/2211.00539
https://proceedings.neurips.cc/paper/2020/hash/569ff987c643b...
https://github.com/facebookresearch/nle
https://ojs.aaai.org/index.php/AIIDE/article/view/12923
I am curious how well Voyager would do in Nethack.
As long as they're all still "special" single-purpose systems (LLM is about processing and responding to language input for example, CV / Computer Vision models specialize in operating on visual or image inputs, etc.), that's all they'll ever be, no matter how good they get at pretending they're more.
It’s very good at understanding text but by itself it can’t think. Turning text into commands it knows is doable.
You just scroll down a tiny bit on the twitter page and get this nice video and summary from the author.
> Voyager has 3 key components:
> 1) An iterative prompting mechanism that incorporates game feedback, execution errors, and self-verification to refine programs;
> 2) A skill library of code to store & retrieve complex behaviors;
> 3) An automatic curriculum to maximize exploration.
> 9) Do not write infinite loops or recursive functions.
> Sometimes GPT-4 will write an infinite loop that runs forever.
I'd like to see a visual/language model/AI that learns to play minecraft as an actual inhabitant of the game. i.e. processing visual input, recognising objects, working out whats going on, learning how to move around. Learning how to make food and avoid monsters. It would be an 'Embodied AI' within the world of Minecraft.
The language part would allow us to talk to this being. You could ask it things like:
"Do you prefer to make a house, or dig a cave?"
"What are your hopes for the future?"
"Is there a recent achievement you are particularly proud of?"
etc
There is already an AI VTuber, Neurosama, trained to "play" games, including Minecraft (also OSU! and Among Us.)
I don't think she's learned lava is bad yet.
Perform self-reflection every time damage is taken (the iterations of self-reflection can depend on % health lost). Something like "Please output as a list of general guidelines / best-practices that any Minecraft player can use in the future. For each guideline add a risk profile that is introduced if the guideline is neglected. Based on the following game log over the last 100 steps, what should the Minecraft player have done differently to avoid taking damage? \n <game log>"
And keep storing those guidelines over multiple iterations. If there start to become too many best-practices, ask GPT-4 to "Please output as a list of general guidelines / best-practices that any Minecraft player can use in the future. For each guideline add a risk profile that is introduced if the guideline is neglected. Take the guidelines below and summarize them. Prioritize retaining more detail about the guidelines with a higher risk profile, and do merge guidelines if possible and appropriate. \n <old guidelines>"
Given the number of times I've taken damage in Minecraft from a Creeper that has seemingly just appeared behind me, I expect this feedback loop to pretty quickly end up with a bot that does a full surroundings scan after every action to make sure there's no Creepers around.
A big text corpus gives quite a huge amount of context.
SD also got better through this
Because the researchers found it to be an interesting problem. Using an AI to beat Minecraft has been an active benchmark in the ML world for some time now.
> Who wants this?
Me, as well as the many others who have liked and shared the paper.
> To what end?
Furthering our ability to create autonomous agents. If we can get it to work in Minecraft, that's one step closer to getting it to work in real life.