What kind of bug would make machine learning suddenly 40% worse at NetHack?
arstechnica.com
arstechnica.com
It is absolutely an anti-feature. Most things in life function on some level of predictability. Computers are pretty high on predictability scale, do the least surprising thing and all that.
I sincerely hope there is some kind of feedback mechanism for the "not a technical manager" to either learn or get less decision making authority.
This is pretty common. They wanted to test v2 in the regular environment, so they can see how it reacts to real user traffic. If it reacted well, they would replace v1 with v2 completely under the hood! See, they didn’t want v1 and v2, they wanted the existing endpoint updated while maintaining backwards compatibility!! Their system was too old or had poor test coverage, so in order to minimize risk of disruption to the service, they came up with the idea to try it out for 2 off-peak hours.
Why do that vs better test coverage? They were confident enough, or the risk of breaking things for a little but was understood and justified, maybe they had some super old machines or configs they cannot accurately model and really wanted to get proper user traffic in.
This is how many a/b tests happen too. Turn on feature for a segment, and then turn it back off. Or rolling deployments - have code run on 5% of servers then increase the percentage. Rollback if problem.
This is that just except time based. Kinda hacky but everything in the story sounds reasonable to me
It gets a pass but my top 1 remains "I can't send emails further than 500 miles"
Edit: it was the other way around.
As someone who doesnt understand ML - I have always assumed the whole point of ML is to try different things in the game, almost randomly, and over (long) periods of time the AI gets better and better at the game.
If having a single unexpected event causes such a large swing in outcome, and the AI cant "explain" what is different to cause the swing, then what exactly is the ML doing for it to fail on such a seemingly simple change? Doesnt that defeat the whole purpose of this?
I'm obviously missing something obvious - because I would assume the real goal of ML is that it can teach itself the game, even if that involves unexpected situations, as a human does?
If nobody includes the full moon message as input to the ML model, and tries to operate the ML model with the training it has achieved running in non-full-moon mode, its operating score in full-moon-mode may be lower.
Even if it had proportional training time against full-moon-mode to incorporate that into the model, if you don't tell it when full-moon-mode is active wouldn't the optimal behavior be to optimize the score for 27/28 days vs 1/28 days of the month?
If full-moon-mode is an input to the model, then it can trained to optimize for both scenarios.
Some of those assumptions were different, and since it's not learning/training it couldn't adjust for those new assumptions, so it didn't do as well
If you a human, were forced to follow a set script/assumptions the same would happen to you.
If you train an ML model on thousands of attempts at going around some racetracks where touching the walls slows you down, and the score is achieved by executing a fast lap, and the inputs to the model include where the car is and where the walls are, it should optimize towards avoiding touching the wall.
This behavior would likely still work even on new procedurally generated tracks that the model had not previously seen, as long as the relationship of inputs (car, walls) to desired behavior (fast lap) still applied.
If every N number of runs for a large value of N the game changes so that the walls are actually speed boosts and the center of the track slows you down, and there is no input to the ML model to tell it that the situation is different, it will initially try the previous strategy and perform worse, and it will be difficult to train it to handle both versions of the game without some discriminating input value to train on.
I predict the next "annoying non-bug" will be Friday, June 13th of 2025.
When you try to fight a werecreature in animal form it can summon large numbers of animals of its kind to attack you. This can be extremely deadly for a player who is unaware of this ability. An experienced player knows to attack werecreatures only at range or avoid fighting them altogether. However, encountering the werecreature in its human form is much less dangerous unless it's carrying a powerful weapon.
This tells me the algo is trying to hard to predict the game or learn a decent static strategy, rather than make situational decisions.
Wouldn’t be very surprising if the agent hyper-optimised farming those critters for points. It would not be able to change strategy if the cost/benefit of that farming changed massively, so would now be performing significantly worse.
The advantage of human common sense over machine learning models — at least when it comes to role-playing games — is that we carry around a ton of this cultural information. A model trained only on NetHack — not on broader culture or folklore/fairytales/mythology/fantasy — is simply not going to be aware of this link between full moons and specific monsters becoming more dangerous. So if it’s developed a fairly naïve strategy of just fighting or avoiding everything in its path based on a model of relative strength then it’s going to be tripped up when an outside event (the phase of the moon) upends that model.
bet it comes down to how much memory the algorithm has, since the transformation might occur way later than being bitten, while most poison kills are fairly quick. The problem is NetHack requires you to have at minimum 1000 turns of memory to know when to pray. Even more if you want to keep track of where stuff was.
Or, something that has been going on regularly for decades.
This isn't even close to the quality of the 500 mile email IMO, yet they seem to be doing everything possible to ride those coattails
My memory has not lasted well
The bug here is not in nethack but just in the training, which meant some of the special cases (full moons, fri 13ths etc) weren't in the training data. They should have been running the training in VMs with the clock set to include these cases.
Honestly a lot of the reporting of this "bug" seems wildly overblown.
I don't agree that the creators would be watching the game play either. Usually during such training phases you'd run as many copies of the game as the available hardware can manage. I wouldn't be surprised if they had at least hundreds of runs going in parallel and the researcher is definitely not going to be watching them all. If anything, they are going to bed and train the model overnight as much as possible.
Then that’s a poor model then. Significant anomalies should be flagged for manual review, otherwise you corrupt datasets unintentionally.
Assuming you don't adapt your playstyle in any way, how would that lead to worse overall performance?
>edit: aha, I think this is it - attacked werecreatures are much more likely to summon help on full moons. Poor bot probably got overrun.
It may have better outcomes in most situations, but if you're depending heavily on a particular strategy and that changes then you're in trouble.
By the way, this story is several weeks old, ars is late in covering it.
NetHack does warn you when you (re)-start a session that it's full moon (or new moon, which also has effects).
> Maximizing the score means that you will just farm monsters. Finding items required for ascention or even Just doing a quest is too much for pure RL agent.
So what is the stop condition then? Elapsed time? Does it run out of monsters sooner because the fullmoon makes "werecreatures mostly kept to their animal forms" and there are simply less easily farmable points in early levels?
There is a LOT of knowledge and strategy that is VERY FAR from obvious in this game. "Unspoiled" players who haven't read the wiki only have a very faint chance of winning the game.
If you sit around in early levels without trying to make progress, you eventually run out of food, your equipment will not improve and may even degrade, and worst of all you level up, which means monsters start scaling faster than you. You have to rely on prayers to survive, but prayers have a random cooldown, and if you pray too early, your god will make sure you regret it.
If score is not tied to progress in the game, I'd say the agent's scoring system is, by definition, incorrect.
> our whole environment is in a single, self-contained file
Except the parts of the environment which aren't in that file, like the current date.
Sometimes it's hard, sometimes you need to come up with novel visualizations, sometimes it can only give partial insight or just be noise, but I always strive, where I can, to have some type of method to look at data, in as close of a form as it was generated as possible.
I think in this instance, look at an actual run might have caught this issue earlier. They might have seen the "it's a full moon" message and might think that's odd, or see were-creatures keeping their form or seeing agents being extra lucky, or whatever, but running it headless and just looking at the score means they're cutting off a huge signal vector (and, admittedly, a huge noise vector).
Just started a game, I got:
"Be careful! New moon tonight."
I wonder if that will affect this learning test ?
Nethack, by design, changes the gameplay (slightly for some characters, greatly for others) based on the phase of the moon.
It tells you this, the same way the game tells you everything else.
There is no sane justification to call this a bug on either side, it's just a poorly trained model responding poorly to a feature it hadn't seen in training
https://www.talkrl.com/episodes/pierluca-doro-and-martin-kli...
They blew it: It’s “What a horrible night…”
Of course, "score" is not a real metric for success in NetHack, as Cupiał himself noted. Ask a model to get the best score, and it will farm the heck out of early-stage monsters because it never gets bored. "Finding items required for [ascension] or even [just] doing a quest is too much for pure RL agent," Cupiał wrote. Another neural network, AutoAscend, does a better job of progressing through the game, but "even it can only solve sokoban and reach mines end," Cupiał notes.
The NN seems to be good at grinding. They should make some for those free to play games.