The NetHack Challenge: Dungeons, Dragons, and Tourists (2021)
nethackchallenge.com
nethackchallenge.com
Even spoiled, when it's achieved, an ability to complete a task like ascending in NetHack seems more world changing to me than the LaMDA conversations.
The Challenge highlights an important limitation of modern machine learning systems, that they can only learn from their training data and have only a very, very limited ability to incorporate "background knowledge".
"Background knowledge" in this case is the "spoiling" you're refering to. There is virtually no way to teach a neural net system how to play Nethack, other than to let it try, and fail, and fail again, at ascending. You can't explain to it for example, how writing "ELBERETH" in the dirt with your finger will scare monsters away (which is important background knowledge about the game that is very hard to learn just by trial and error).
https://groups.google.com/g/rec.games.roguelike.nethack/c/wc...
If I had the GPUs, it would be an interesting thing to experiment with.
However, having comprehended the guidebook, it then needs to also comprehend the nethack terminal client. The pile of obstacles does not cease to grow.
I have ascended once. (I was 25 at the time)
I got stuck on the vibrating square many times, the castle many times, and have died on the elemental planes many times. Because I had no idea to do what was expected to me. The first I got to Medusa I died instantly.
Yes, I think some of the final puzzles are a bit too obscure... But I think you can beat it without spoilers.
But I would spend entire days to learn mechanics... IE: day of kicking sinks is a day I remember fondly in my childhood.
I am sure there is tons I don't know. But I have never used explore mode, and I have never looked up spoilers before ascending.
Note, I had some word of mouth friends, and it took me 3 years alone to figure out altars.
This feels akin to Classic Game AI vs Modern AI debates that happen all the time. And even in 2022, with desktop GPU capabilities nearing supercomputing levels, it feels like rules-based, goal-oriented planning still dominates NPC & Enemy AI in games. The question really becomes, why did symbolic ai research die when it's so effective at specialization? Rather, the research obsession is solving generalization. And the problems always seem to stem around the black box itself, the lack of "human-legible representations" of its latent spaces ;)
I think this is largely because the goal of most game AI is not to win per se, but rather to provide an engaging experience to the player. To that end, a game designer wants to be able to purposefully craft the experience, and explicit rule/goal based systems allow that.
I think we want "AI" that scales faster than humans can implement. Game "AI" looks nothing like machine learning and looks a lot more like some programmer just smashing out some code and trying it out for a while. Some autotuning can be done, of course, but it's largely a program being written. Who has time for that?
Moravec's paradox again: humans find perception easy but things like Sokoban hard, while GOFAI approaches find Sokoban so trivial that it's common to use it (or Sudoku) as a toy problem introduction to constraint solving. Nethack literally has levels which are just Sokoban and which are important to solve; GOFAI can push the boulders around perfectly, even though it would be unable to recognize a photograph of 'a boulder' or use the word 'boulder' in a story.
Then you have the extensive hand-engineering of expert knowledge which goes into ascension agents and the symbolic winners, above and beyond merely plugging in a Sokoban solver. There are increasing experiments in making DRL agents exploit or initialize from pretrained language models (https://arxiv.org/abs/2005.07648#google https://arxiv.org/abs/2201.12122 https://arxiv.org/abs/2204.01691#google https://ai.stanford.edu/blog/DrRepair/ https://arxiv.org/abs/2005.07648#google https://arxiv.org/abs/2009.03393 https://arxiv.org/abs/2204.00598 come to mind) or reading manuals (https://arxiv.org/abs/1401.5390 all the way back in 2012!), and of course, a DRL agent can learn a tremendous amount without actually doing any playing by offline and off-policy and imitation learning, but while it is exciting and things like Gato look like the future, there is a long way to go from feeding in a dump of the Nethack wiki which mentions offhandedly "you can do X" to an agent recognizing an opportunity for X in the wild and executing it. (Which is something that symbolic approaches also don't come anywhere close to doing, because they just cheat by the capability being given to them by hand-engineering rather than having to autonomously read, understand, and apply.)
And especially with those logged games, language models with retrieval will be useful: https://arxiv.org/abs/2206.05314#deepmind
This year, I don't think there are any obvious bots running. The only time a bot clan was participating AFAIK was 2015.
https://junethack.net/archive/2015/scoreboard.html
And it didn't do badly (it did score 2 of the 5 clan trophies) but also not particularly well. That's mostly due to the tournament including lots of NetHack forks and the clan trophies being special side achievements that aren't necessarily needed for winning the game. So a generic bot that is trained to win the game isn't best at getting those.
And for today, I think there is no bot that can win the game for the newer versions or any forks. The bots from the NetHack Challenge also can't participate as they need a custom NetHack binary that outputs the game data in a machine parsable way whereas the tournament requires you to play on existing public servers.
Maybe that's related to the nature of NetHack itself, compared to other games where ML has seen (often overwhelming) success?
Will rules-based systems continue to succeed for NetHack, or is this a temporary condition until the state-of-the-art in machine learning improves?
(I did find BotHack[1] and a video of its' first ascension[2], both of which are cool)
So it's not really surprising that hard-coded rules based bots incorporating spoiler knowledge beat learning bots trying to learn the game directly from the interface.
qw: A fully automated lua bot written by elliptic, with some code borrowed from parabolic and xw. The first DCSS bot to ever achieve an uninterrupted and unassisted win (see '!lg qw won 2'). Now maintained at https://github.com/crawl/qw by the DCSS devteam. See qw[2] for a summary of results.
As of 0.29-a, qw has a 0.36% winrate with GrBe with 1 win in 276 attempts. See https://crawl.dcss.io/crawl/morgue/qwqw/morgue-qwqw-20220613... and games are sometimes played on cdi. Historically its best 3-rune winrate was 15% DDFi^Makhleb, and, for 15 runes, about 1% with GrFi^TSO before the 0.28 hell rework.
Branch order: D -> Lair:8 -> Orc:3 -> D:15 -> S:5 -> Vaults:4 -> Depths:5 -> S:5 -> Vaults:5 -> Zot
On the online servers, qw plays with an extra added delay so that it doesn't use too much server CPU. Playing locally without this delay, qw is much faster.
You need a lot of knowledge to win DCSS, too, but it's more about wise use of limited resources, playing the hand you're dealt, and the dangers of various locations. It's much harder to codify as spoilers, but you will learn by trial and error.
I think there is a lot of opportunity to make NetHack even more amazing by having smarter monsters that come up with new strategies, often as a group, to beat the player.
Firaxis iirc has never shipped an iteration of civilization where the AI could compete on a level playing field with even quite mediocre humans.
https://alt.org/nethack/top60d.html
5000 points is still a relatively low score, it corresponds to reach Dungeon level between 5 and 7, which is quite early in the game.
What's more interesting is a low score ascension. That is insanely hard to pull off as it means you're taking a very fragile character into the endgame, and yet somehow pulling through.
It's definitely a hard nut to crack, and I can see why the machine learning attempts would've struggled so much, but honestly I'm surprised that the best symbolic logic bots only achieved a few thousand points on average. That's really not even getting very far into the game.
I guess NetHack is a harder game than StarCraft II, as with SC, the best bots are competitive with the best humans (albeit they have advantages in inputs/attention and disadvantages in grand strategy). NetHack though is literally turn-based (like Chess), so there's no advantages to be had here, and clearly they're struggling for it.
> But once you put that much actual effort into getting good at it, you stay good
I had the same experience, but I don't think it's a matter of being "good" at the game so much as having memorized all of the stupid ways to die (many of them insta-deaths ) plus how to avoid them, and knowing the mechanics of how to win (I think you're supposed to figure out the invocation ritual by giving tons of money to the Oracle, but it's been decades and I can't recall if I learned it that way or just read spoilers). edit: just remembered the whole vibrating square thing and that there are a few ways of making a run unwinnable, which is a kick in the teeth when you have a promising @ going.
Maybe that's just splitting hairs (is that just becoming "good"?) but in retrospect, I don't think that getting good at nethack was worth the investment in time and frustration. I will say that playing nethack improved my vocabulary.
A large part of becoming good at NetHack ends up being learning to be patient, playing conservatively, and taking lots of time to think before acting. And there are plenty of mini-games within the game (such as Sokoban, potion/scroll IDing, pudding farming setup) that you need to learn well too.
What helps most of all is being resourceful. There are so many different systems in the game that can all be used in the right situations to give you an advantage (e.g. Elbereth, wielding cockatrice corpses), so you just need to play slowly and cautiously and continuously think about every single possible action of the dozens available at any one point that you can use to maximal effect in any specific hairy situation.
I'm with you on the ID minigame, but Sokoban is very skippable: winning it gives you one of two valuable items, but it's not the only chance to get them. And I've ascended every class without ever once pudding farming. Your overall point stands, though.
Nethack is not a game where the dice are fair. Nethack is a game where the phase of the moon has real influence.
Nethack is, though not originally designed for that purpose, a game which encourages players to learn to read the source code and familiarize themselves with text files, editing, and searching.
You could build a nice first-year CS class around Nethack, taking people from "this is a fun game" to "I have a tenuous grasp on programming" in a semester.
Hard disagree. No human player would remotely detect a difference if the exploitable pseudorandom seed were replaced with a truly random source.
> Nethack is not a game where the dice are fair. Nethack is a game where the phase of the moon has real influence.
The phase of the Moon _is not random_! It is deterministic! You're arguing against yourself here.
> Nethack is, though not originally designed for that purpose, a game which encourages players to learn to read the source code and familiarize themselves with text files, editing, and searching.
Entirely orthogonal to the issue of using a better random number source. Indeed, a good class project might be improving the source of randomness used by NetHack.
- is necessary for technical reasons
- is easily distinguishable from a different one by players
- can (and since it is Nethack, should) include the phase of the moon as a seed> Fountains have a 1/30 chance of summoning a Water Demon when quaffed, this demon has a (80+DL)/100 chance of being hostile. If not hostile, it grants a wish and vanishes.
> every time the character attempts to walk into a wall, it calls random() without wasting any in-game time
> So it is advancing the RNG millions of times (or whatever is necessary, maybe not that many times) for each "real turn" until it reaches a favorable point in the RNG as tested in parallel offline (e.g., find a spot in the RNG such that a quaff 1) gets a demon that 2) grants a wish and 3) the fountain stays... advance the RNG as many times as it takes to repeat that... again and again, for 90 wishes)?
This is not solving nethack, this is just save scumming / cheating.
The tournament in the NetHack Challenge is a bunch of bots playing the game in an ordinary manner. There is no comparison between the two things.
Yes, the bot is taking advantage of hidden state with the RNG... but there's nothing stopping a human from doing the same. It's possible for a human to take enough RNG-based actions and observe the outcomes to derive the current state of the RNG seed. Maybe that number is in the thousands or millions of actions, but a human could do it the same as a bot, we just don't choose to try.
There's an argument that would call this cheating, but I don't think so. There's plenty of hidden state in Nethack (unidentified items, anything out of your field of view at the moment) and playing the game is largely an exercise in deducing that. The so-called RNG is no different, it's just another piece of hidden state, as far as the program is concerned.