I don't agree that the creators would be watching the game play either. Usually during such training phases you'd run as many copies of the game as the available hardware can manage. I wouldn't be surprised if they had at least hundreds of runs going in parallel and the researcher is definitely not going to be watching them all. If anything, they are going to bed and train the model overnight as much as possible.
Then that’s a poor model then. Significant anomalies should be flagged for manual review, otherwise you corrupt datasets unintentionally.
Assuming you don't adapt your playstyle in any way, how would that lead to worse overall performance?
>edit: aha, I think this is it - attacked werecreatures are much more likely to summon help on full moons. Poor bot probably got overrun.
It may have better outcomes in most situations, but if you're depending heavily on a particular strategy and that changes then you're in trouble.