Minecraft diamond challenge leaves AI creators stumped
bbc.co.uk
bbc.co.uk
If anyone had pass the bar this challenge set, they'd have basically solved the biggest hurdle in general AI, and would probably have received a call from John Carmack.
That is very big in AI right now, to try and remove the need for the "human in the loop" and learn complex tasks from 0 background knowledge.
Obviously, like you point out, this is very difficult to do and it requires lots and lots of data and compute (no, more than that. Really lots). As the article reports, this eliminates any hopes of "democratising" machine learning and makes advancing the state of the art a game that only large corporations can play with any hope of winning.
All the problems solved by fine-tuning image net for example, instead of training from zero
Identifying objects in a 3D environment is useful across several domains and it shouldn't be limited to Minecraft
That's not a very useful goal. I can trivially make a game that is impossible for such an AI to win. It would require you to enter a 1024 bit code [0] and if you are correct you win. For a speed runner the time it takes to beat the game is merely limited by their typing speed because the code is always the same but an AI with no prior knowledge can only brute force the solution. How does the speed runner win in the first place? They ask the game developer and then post the code on a wiki.
[0] an euphemism for minecrafts crafting system
The way you'd train a neural network with that footage is comparing the recording of the neural network playing the game to the recording of the speedrunner, and scoring it based on the difference in the footage.
However it wouldn't really be learning "from scratch" and "end-to-end" anymore, since you're essentially teaching it to imitate and play like a human player. Also your network would likely be just blindly executing a series of learned actions, without any real understanding of the game. If you confronted it from some situation that's not part of some speedrun it memorized, it would probably be at a loss. Pretty useless.
In the example above, at no point the neural network would be watching the video footage. The reference footage is just used for (automatically) scoring how well it does during training. It's not fed into the network itself in any shape or form.
You can perhaps design a game that is difficult to solve by AI (though I wouldn't bet on it). The question with AI is usually not what it can't do but what it _can_ do. If a certain system fails at your task but performs well in a range of other tasks that are considered useful and challenging, then it's worth consideration. Especially so if there is no system that can perform well on that one task.
Setting up a stupid method of getting to the goal doesn't disqualify said goal.
In a strong sense, this implies ML is not "intelligent" in any conventional sense of the word. What little we have of understanding of learning and intelligence in the scope of humans strongly implies that drawing parallels with solutions to other tasks, building on previous knowledge, is the key to intelligence.
Perhaps the purest expression of this is in pedagogy of math. Piaget and Brissiaud are some of the scholars who I think really grokked this. They brought forward the concept of compression as essential for learning. E.g. if you have learned but not yet compressed the operation of addition, you will have a very hard time learning multiplication. After addition has been compressed such that you don't have to spend any effort doing it, you are ready to learn multiplication.
Maybe it would be worthwile to explore if there are AI techniques that can learn simple tasks first, then use that knowledge to solve harder tasks?
All the data and the training process together are very clunky in comparison.
IE, an inability to be taught by others is a serious limitation to say the least.
Or, when we see “one” example of a new animal (for example), do we see a single picture? Or do we see it moving for about a second? Effectively multiple images? And, having previously learned a general model for physics, stereoscopic vision, and lighting, does that one second of motion gives us the size, shape, and a guess at the musculature of the animal?
If you strip those advantages, how many training examples do we need? The examples of people blind from birth or from extreme youth who regain vision in adulthood[0] implies that we learn a lot in our youth about how to see, which we can share between recognition tasks we learn later. Keeping those basic skills — avoiding “catastrophic forgetting” — still seems to be a rare and advanced feature in AI, but it’s not impossible.
There is still a lot we don’t know about how our own minds work. We might be closer than we realise, or much further.
[0] https://en.m.wikipedia.org/wiki/Mike_May_(skier) amongst others. I never did find the name of the guy who learned to see from touching a monkey statue…
This demands the mastery of language which is a very high bar. Most animals haven't achieved this, so I think if you expect to get a computer to learn this way before being able to reason as well before you consider it having intelligence, you would also have consider most animals not having intelligence as well.
Then avoid language as medium. Many animals can learn from being shown an action once and apply it to themselves.
I think you are out of luck already. Even with transfer learning, retraining BERT to fit a larger dataset takes ages even on an RTX 8000. A single state-of-art experiment/training run at Facebook now costs 7 figures. The age of individual SOTA ML researcher/developer is over.
To me it shows we have already picked the easy fruits and technology will not continue to evolve exponentially.
I don’t know why people have so much faith in ML and AI. Considering what a CPU actually does, shouldn’t the default position be one of skepticism?
It’d be truly amazing if you can make a machine intelligent in the human sense by merely coming up with a clever algorithm.
Considering what a neuron actually does, shouldn't the default position be that the brain is a computer?
Maybe biologist come up with the next big break through. I mean, if you want to match an organically grown organism ...
It’s just easier, engineering wise, to start over from scratch on each iteration. And it makes it easier to evaluate how each change improves the outcome or not.
Well...not so much if you consider that machines are orders of magnitude faster than humans in any other task, including processing and recording data, a.k.a. learning.
so 4 days to train their AI and not 4 days of training data.. on the data which was "60 million frames of recorded human player data"
You misread it. It's around 300 hours of training data. It's 4 days to absorb those 300 hours.
I bet a human child who has never even seen a computer could learn to do this in far less than 300 hours. I suspect even 1 hour would be enough.
I don't know if there's a starting tutorial in the game now that explains some basic crafting recipes to get started with (there better be!), but unless the AI can understand it, or you're explicitly programming in the recipes as data, good luck.
[1] https://gitlab.aicrowd.com/minerl/minerl-resources/blob/mast...
But I dont think the problem is to get Minecraft agents that can make diamond swords ... it's to use only ML to make agents, which means throwing away all that. That's fine.
I guess they needed to add tutorials to better accomodate the wave of younger people who picked up the game in the mid 10's, but I actually appreciate Notch's decisions (or series of happy accidents that culminated in a highly addictive game).
Alphastar SCII bot has been using much more resources and time than this to train, so maybe there is one of the reasons no entrant has achieved the goal yet.
> A relatively small Minecraft dataset, with 60 million frames of recorded human player data, was also made available to entrants to train their systems.
It will be interesting to see if these artificial and somewhat arbitrary constraints (although I get that the idea is to restrain it to resources that are somewhat realistically available to a single individual without organizational backing today) will either cripple this challenge or in the end yield some innovative results because the entrants will have to devise algorithms that use much less data and resources than what has been traditionally required to get SOTA results.
Additionally it is also unclear whether the way humans learn to play this game is actually using a smaller or a much bigger dataset to learn from. Sure, a human can learn to play it in 20 minutes, but that's after 9-10 years of other pretraining of seing, understanding and operating in the 3D physical world performing various tasks, getting compressed knowledge from other people by watching and listening to them... Maybe that would be an interesting challenge - to still constrain the final model to 1 GPU for 1 day, but at least allow the model to pretrain on arbitrary similar data, if it is not sourced directly from minecraft or any clones.
Maybe there is a reason they had this restriction? I can only think of allowing the winning AI to be ready for end users whic h generally have a single GPU?
The restricted training resources are just part of the challenge. They point out that a human child can learn the necessary steps in minutes by watching someone else do it, so they wanted to see if anyone could make a computer learn it with relatively limited resources.
They're just biasing competitors towards efficient solutions.
It's probably worthwhile to note that AlphaStar is trying to become as skilled as the strongest human players, whereas in this case it's more of a binary "is capable of getting diamonds" thing, they don't need to be world-class diamond miners.
But also, the machine learning algorithms we have nowadays are in many ways superhuman. Seeing how much can be learned with these constraints is interesting as well.
- must submit source code - must submit source code - must submit source code
There's your problem right there. It's not that 'creators' are stumped. It's that, given how much it would be worth pitching the same closed source to an investment group or solution seeker directly, no 'creator' with a sufficiently advanced 'new' or capable system would ever submit their source code to one of these competitions. This goes for all competitions that require participants to submit their source code. It's what you you seek or your backers after-all which is valued far more than the potential prize you're doling out. Thus, your business model. The question is always, will someone who can develop such an advanced system be dumb enough to part ways with their IP for such low value? I think not.
So, you can pretty much throw any conclusions made from any such competitions in the trash. This goes for even bigger ones by bigger names. You're going to get what you'd intelligently expect : a small sampling of the same ol' same ol' approach. Don't expect any novel submissions. Don't expect any surprises. So, what's the point of this? I'm speaking about the whole site 'aicrowd' btw and any other group that organizes a 'submit your code' competition ...
I’ve written bots for MMO’s and the hardest part for a task like this is fighting the API the bot has access to. I suspect I could even solve this challenge 10% of the time with a blind bot that just did some brownian motion. Unless you never run into trees, every other resources is pretty easy to just dig down for.
The only thing that would make this hard is the thing that makes it hard for a human player which is knowing you’re on the diamond level without using the xyz debugging info.
That’s a really good point, lava pools start at y=10 so mining around 11 is both safe and in the diamond zone. The question is if there’s enough training data to reinforce an agent so that once it finds a (underground) lava pool it should start mining at that level.
Something else that makes this much harder than it is for a human: there's no sound! You can't hear the nearby lava or water when digging down.
Check out the information in the observation space-- nothing about sound in here.
http://minerl.io/docs/environments/index.html#minerlobtaindi...
But maybe a computer could do a better job than a human at looking at the clouds before digging, estimating how high they are, then keeping track of how many levels down we've dug. I'm pretty sure the clouds are at a fixed altitude.
The rule is actually quite vague, which is not very surprising, as it seems quite hard to define what domain knowledge is allowed and what isn't without having lots of loopholes.
The easiest way to code this with traditional methods would be to still use deep learning for the image recognition part of it. The input to the agent at each step includes an array of numbers representing the pixels on the screen. So it doesn't get to see the 3d terrain directly; it sees a projection into 2d, and needs to recognize the 3d terrain and objects. And from there it needs to synthesize that into a map of the situation-- for example, the presence of lava, water, cliffs, and hostile monsters.
Doing that kind of object recognition without deep learning would be pretty onerous. The strategy and planning parts of this, though, I agree would be comparatively straightforward. Further, even if it started with a representation of the actual terrain blocks in a 3d model, the synthesis step into a representation of what the overall situation is-- that also is much more easily done with deep learning than with traditional methods.
Back in the day I wrote a highly effective bot for playing Star Wars Galaxies which would run missions unattended. It was created exclusively with hand-coded routines and a big nested state machine. It included a hand-written OCR library for reading the compass coordinates and accepting missions from the mission terminal in a particular direction. It knew a pre-recorded path for navigating in and out of town to and from the mission terminal. It used the little red dots on the radar scanner to determine the presence of nearby hostile creatures and mission targets. I ended up selling this system to a gold farming outfit in China; it was only available for sale to the public for a handful of days.
I bring this up because this bot for SWG didn't contain any modern machine learning or deep learning techniques, and it worked great. But if it weren't for the coordinates displayed on the radar (used for navigating through the maze of town back to the mission terminal), along with easily discernible red dots on the radar indicating enemy presence (used for knowing which direction to face during combat), I'd have been at a loss for the complex object recognition / situation recognition necessary to turn this into a task solvable by a straightforward nested state machine.
Seriously, I'm starting to think turn based strategy games like Civilization, etc, will be the last to get any attention. Why? We need good AIs for these games and they're not yet solved as far as I know. Furthermore, you don't have to model the human interface by limiting actions per second like with DotA or StarCraft. Yes, we solved Go, but turn based strategy games on PC are more complicated. Seems like a worthy area to research.
Does anyone know of any work on games like these?
I don't think the answer is technical. A good AI for a game like Civilization would have knowledge of human social concepts like honor and vindictiveness, which are both concepts problematic to pin down. The definitions of concepts like these are fluid, they change with the times and societies in question. But that's not to say any possible setting for these concepts is equally valid when trying to emulate human behavior. Some settings will seem unrealistic, like a mustache twirling villain or a hyper-rational pointy-eared alien. Neither make for good AIs if you're shooting for human-like behavior.
Humans can certainly individually craft fictional personalities that seem realistic; authors do it all the time. But from a gameplay perspective that tactic falls flat; you end up with a limited set of personalities the player becomes familiar with and the game consequently loses relay value. In the Civilization games, every player knows that Gandhi has a short temper and likes to launch nukes. Rather than crafting individual personalities, how does a game designer define a function that returns personalities with a realistic distribution? How can a game designer define such a function if the parameters aren't truly known by science and the artists who craft individual personalities are going off wishy-washy metrics like gut instinct and artistic intuition?
This is a pretty wild assumption to say the least.
These games have parameters for loyalty, forgiveness, propensity towards warmongering, and more. That's part of what these games are. If you took away that philosophy towards AI design, you'd no longer have a Civilization game.
As for tactical AI (which is only one component of Civilization), an optimal play from the AI probably isn't what most players desire, nor is optimal play with randomly imposed inefficiencies. Real humans don't run countries optimally, so an AI that runs a Civilization nation optimally won't ring true. You want an AI that rarely makes mistakes that a human wouldn't make, but frequently makes mistakes that a human would make.
Making a good AI for a game like Civ is not nearly as straight forward as creating an AI for a game like Chess or Go because players have different expectations from the games.
No one's demanding optimal play from AI. Meanwhile I'm sure pretty much no one's happy with idiotic AI, which is what we have at the moment. I already gave an example: "smashing waves and waves of units into a walled city with +30 strength." Any human player familiar with the rules won't sink 100 city-turns worth of production over 20 turns into attacking an unconquerable city, losing everything while shaving 20% off the wall, then repeat that for another 40 turns; that's idiotic, and that's exactly how the AI behaves alarmingly often. Do you enjoy dealing with this kind of behavior? I bet you don't, unless you only derive pleasure from crushing AIs (which I do enjoy, but the satisfaction from the lack of challenge only lasts so long).
Currently game difficulty in Civ is basically defined by how much of a head start AI players have and how much they cheat (which in theory is a consistent advantage throughout the game but in reality matters less and less once the human player starts conquering), so once you overcome your early disadvantages the late game becomes boring.
> Making a good AI for a game like Civ is not nearly as straight forward...
I never said it's easy to create a great AI for Civ. However, it shouldn't be hard to outdo the existing, idiotic AI by a huge margin (while preserving all the traits like nuke-happiness) if, say, DeepMind decides to put some resources into it.
I don't think difficulty in analyzing the odds of engagements like that is the reason good AI for Civilization is elusive.
You're not though, you're just trying to build a bot that can win the game.
(Obviously the ideal is an AI that plays like a human and can hold its own, but a playstyle similar to humans is definitely more important than ability to win.)
After I learned the basics my own opinion is that Civilization is no longer a fun game, due to the poor AI.
We can probably agree that finding fun and interesting ways to dumb down an AI will be a cool area for research and new ideas in the future, but we need strong AI before we can figure out fun ways to weaken it.
Not sure if the AI can toggle the debug menu, which shows the Y level, or if it can see the Y level during training.
Is this a job well suited to AI, vs a canned algorithm?
Observations: https://reddit.com/r/dataisbeautiful/comments/efvgve/_/fc2p6...
> The submission must train a machine learning model without relying heavily on human domain knowledge. A manually specified policy may not be used as a component of this model. Likewise, the reward function may not be changed (shaped) based on manually engineered, hard-coded functions of the state. For example, though a learned hierarchical controller is permitted, meta-controllers may not choose between two policies based on a manually specified function of the state, such as whether the agent has a certain item in its inventory. Similarly, additional rewards for approaching tree-like objects are not permitted, but rewards for encountering novel states (“curiosity rewards”) are permitted.
[1] https://gitlab.aicrowd.com/minerl/minerl-resources/blob/mast...
Sounds like a simple state machine. Don't they program objectives into their agents? Or are they relying purely on training?
I was working on something that would automate MC a while back, too, and I programmed it as a state machine that would loop through various tasks, which consist of subtasks. When you have a goal (like gathering diamonds), that's just a list of tasks -- and then each subtask consists of a list of tasks too, until you break it down to the "point at block", "walk to block", "swing hammer" level, etc.
It was fun but I am pathologically bad at finishing projects that I start.
[1] http://minerl.io/docs/tutorials/first_agent.html#taking-acti...