The original MNIST training data made the assumption that the digits are not flipped. But you could solve it by creating more training data by flipping the original digits. But then you suddenly end up with an awful lot of data, and then the training process would take days, literally.