I think this looks like a bigger problem specifically because you are in AutoML.
Suppose you are training a GAN. There's notoriously a certain amount of luck involved in traditional GAN training, because you need the adversary and the generator to balance each other just right. So people try many times until they succeed. Probably they were not even recording each attempt, so they do not report how many times they had to run before getting good results.
From an AutoML point of view, this is BS work - the training procedure cannot be automated, and (apart from using the actual seeds) the work cannot be reproduced.
But from the point of view of everyone else, maybe it is fine. They get a generator model at the end, it works, other people can run it.