The bug here is not in nethack but just in the training, which meant some of the special cases (full moons, fri 13ths etc) weren't in the training data. They should have been running the training in VMs with the clock set to include these cases.
Honestly a lot of the reporting of this "bug" seems wildly overblown.
My memory has not lasted well
I don't agree that the creators would be watching the game play either. Usually during such training phases you'd run as many copies of the game as the available hardware can manage. I wouldn't be surprised if they had at least hundreds of runs going in parallel and the researcher is definitely not going to be watching them all. If anything, they are going to bed and train the model overnight as much as possible.
Then that’s a poor model then. Significant anomalies should be flagged for manual review, otherwise you corrupt datasets unintentionally.
Assuming you don't adapt your playstyle in any way, how would that lead to worse overall performance?
>edit: aha, I think this is it - attacked werecreatures are much more likely to summon help on full moons. Poor bot probably got overrun.
It may have better outcomes in most situations, but if you're depending heavily on a particular strategy and that changes then you're in trouble.