The article mentions that “It simply doesn't have data about full moon variables in its training data, so a branching series of decisions likely leads to lesser outcomes, or just confusion.”
My memory has not lasted well
The bug here is not in nethack but just in the training, which meant some of the special cases (full moons, fri 13ths etc) weren't in the training data. They should have been running the training in VMs with the clock set to include these cases.
Honestly a lot of the reporting of this "bug" seems wildly overblown.