One of the main points of Llemma is being an open-source reproduction of Minerva, so basically that's the only comparison they "need" to make. Besides, isn't WizardMath trained on GPT-4 output? They may want to be comparing only among models that can be used commercially