Then I read the article and saw that this project went to absolutely insane lengths to work around the "nondeterministic" and "hidden information" parts.
Then I read the article and saw that this project went to absolutely insane lengths to work around the "nondeterministic" and "hidden information" parts.
The bumping into a wall to advance the RNG state without changing the game time is a really neat trick that I would otherwise have been completely unaware of, and I have played Nethack for many years and done some source diving myself.
Truly a great feat of engineering here.
Initially thought just to probably be some numerical instability when the next actions are all equally likely (I.e. there’s nothing for it to do) but what was actually happening is the pseudo RNG was partially fed by controller input, and certain patterns of controller input would make random rewards more likely to spawn.
I always thought it was cool how these exploits get discovered by the deep nets
I was under the impression that machine learning is only good at learning things that can be approximated by continuous functions.
A RNG is almost the complete opposite of that. It's everywhere discontinuous!
Some games were worse, far worse. Doom used a list of 256 random numbers that it looped through.
Pokemon is another game where the RNG function is extensively mapped, and even used in combination with bugs to generate Mew events, something that should never ever happen.
You see a lot of TAS of old games, especially popular ones, that abuse the RNG generator where feasible.
Here's the link for anyone interested: http://www.roguelikeradio.com/2016/06/episode-122-nethack-to...
A very niche topic on an already niche podcast, but probably will get some decent overlap on HN.
(Also I'm assuming this is the same team? It's hard to tell.)
I'm not sure about the exact terminology. Is an ascension by a bot (like https://www.youtube.com/watch?v=unCQHAbGsAA) also as a TAS?
Although you are right that the group that tried to do a TAS for NetHack (not sure what's their status) chose the MS-DOS port of 3.4.3 for ease of manipulating the RNG in memory and so getting rid of the "hidden information" part.
Edit: also in a game like NetHack, there is potential for an automated tool to help the player. InterHack (https://taeb.github.io/interhack/) was an interface layer for 3.4.3 that added lots of useful stuff, e.g. automatic price identification or wand id from engraving. Most of that has since been incorporated in vanilla NetHack or at least forks.
They seem straightforward to implement even for nethack.
The problem with it is that it's less efficient than exploring on your own.
I usually don't care about that. You learn fast when to autoexplore and when not to, to not miss particular places you value higher than the program. When playing on a tablet, it's tremendously useful and a real time saver.
[0] gamesdonequick.com